Instructions to use GeekedOutAi/Geeked-Out-Quantization-Software with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use GeekedOutAi/Geeked-Out-Quantization-Software with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf GeekedOutAi/Geeked-Out-Quantization-Software:IQ2_M # Run inference directly in the terminal: llama cli -hf GeekedOutAi/Geeked-Out-Quantization-Software:IQ2_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf GeekedOutAi/Geeked-Out-Quantization-Software:IQ2_M # Run inference directly in the terminal: llama cli -hf GeekedOutAi/Geeked-Out-Quantization-Software:IQ2_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf GeekedOutAi/Geeked-Out-Quantization-Software:IQ2_M # Run inference directly in the terminal: ./llama-cli -hf GeekedOutAi/Geeked-Out-Quantization-Software:IQ2_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf GeekedOutAi/Geeked-Out-Quantization-Software:IQ2_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf GeekedOutAi/Geeked-Out-Quantization-Software:IQ2_M
Use Docker
docker model run hf.co/GeekedOutAi/Geeked-Out-Quantization-Software:IQ2_M
- LM Studio
- Jan
- Ollama
How to use GeekedOutAi/Geeked-Out-Quantization-Software with Ollama:
ollama run hf.co/GeekedOutAi/Geeked-Out-Quantization-Software:IQ2_M
- Unsloth Desktop
- Pi
How to use GeekedOutAi/Geeked-Out-Quantization-Software with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf GeekedOutAi/Geeked-Out-Quantization-Software:IQ2_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "GeekedOutAi/Geeked-Out-Quantization-Software:IQ2_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use GeekedOutAi/Geeked-Out-Quantization-Software with Docker Model Runner:
docker model run hf.co/GeekedOutAi/Geeked-Out-Quantization-Software:IQ2_M
- Lemonade
How to use GeekedOutAi/Geeked-Out-Quantization-Software with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull GeekedOutAi/Geeked-Out-Quantization-Software:IQ2_M
Run and chat with the model
lemonade run user.Geeked-Out-Quantization-Software-IQ2_M
List all available models
lemonade list
- Hermes Agent
How to use GeekedOutAi/Geeked-Out-Quantization-Software with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf GeekedOutAi/Geeked-Out-Quantization-Software:IQ2_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default GeekedOutAi/Geeked-Out-Quantization-Software:IQ2_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use GeekedOutAi/Geeked-Out-Quantization-Software with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf GeekedOutAi/Geeked-Out-Quantization-Software:IQ2_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "GeekedOutAi/Geeked-Out-Quantization-Software:IQ2_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
|
Download GEEKED_OUT_INFO.md from GeekedOutAi/Geeked-Out-Quantization-Software: direct link, hf CLI and curl.
- Browser
- Download file 5.12 kB
-
https://huggingface.co/GeekedOutAi/Geeked-Out-Quantization-Software/resolve/main/GEEKED_OUT_INFO.md
- Command line
-
hf download hf://GeekedOutAi/Geeked-Out-Quantization-Software/GEEKED_OUT_INFO.md
-
curl -L -o GEEKED_OUT_INFO.md https://huggingface.co/GeekedOutAi/Geeked-Out-Quantization-Software/resolve/main/GEEKED_OUT_INFO.md
5.12 kB
| # The Geeked Out Quantizer | |
| ## What Is It? | |
| **The Geeked Out Quantizer** is a production-ready quantization environment built for Windows systems. It specializes in extreme model compression using importance-aware quantization techniques, particularly the IQ2_M format which achieves 16x compression with minimal quality loss. | |
| ## The Mission | |
| Traditional model quantization forces a choice: small file size or good quality. The Geeked Out Quantizer breaks this trade-off by using **importance matrices** β statistical analysis that identifies which weights matter most, allowing intelligent bit allocation. | |
| ## Core Capabilities | |
| ### π― Importance-Aware Quantization | |
| - Generates importance matrices automatically using calibration data | |
| - Allocates precision where it matters most | |
| - Achieves 2-bit quantization with only 3-8% quality loss | |
| ### β‘ Hardware Optimization | |
| - Auto-detects CPU, memory type (DDR4/DDR5), and GPU capabilities | |
| - Optimizes thread counts and processing parameters | |
| - GPU acceleration for 5-10x speedup on imatrix generation | |
| - CUDA 12.4+ support with dynamic GPU layer offloading | |
| ### π§ Intelligent Memory Management | |
| - Reserves system RAM to keep Windows responsive during conversion | |
| - Monitors memory pressure and auto-pauses when needed | |
| - Configurable retry logic for transient resource constraints | |
| ### π¦ Complete Workflow Support | |
| - Scans directories for valid source models | |
| - Selects optimal source format (BF16 > F16 > F32) | |
| - Handles sharded models while preserving structure | |
| - Batch processing for multiple models | |
| - Desktop GUI for interactive use | |
| ## Quantization Pipeline | |
| ``` | |
| Source Model (BF16/F16) | |
| β | |
| Calibration Data Analysis | |
| β | |
| Importance Matrix Generation | |
| β | |
| Smart Bit Allocation | |
| β | |
| IQ2_M Quantization | |
| β | |
| Quality Verification | |
| β | |
| Production-Ready Model (16x smaller) | |
| ``` | |
| ## Supported Formats | |
| ### Importance-Aware (IMatrix Required) | |
| | Format | Bits/Weight | Best For | | |
| |--------|-------------|----------| | |
| | IQ1_M | 1.0 | Ultra-compact mobile/edge | | |
| | IQ2_XXS | 2.0 | Maximum compression | | |
| | IQ2_XS | 2.0 | Balanced compression | | |
| | **IQ2_M** | **2.0** | **Best quality 2-bit** β | | |
| | IQ2_S | 2.0 | Higher quality, slower | | |
| | IQ3_M | 3.0 | Near-Q4 quality | | |
| | IQ4_XS | 4.0 | Importance-aware 4-bit | | |
| ### Standard K-Quant Formats | |
| Q2_K, Q3_K variants, Q4 variants, Q5 variants, Q6_K, Q8_0 | |
| ### Ternary Formats | |
| TQ2_0, TQ1_0 β experimental 3-value quantization | |
| ## Why IQ2_M? | |
| IQ2_M represents the sweet spot for extreme quantization: | |
| - **16x smaller** than FP32 models | |
| - **2-3x faster** inference | |
| - **VRAM usage** reduced to ~1/16th | |
| - **Quality** approaches Q4_K with proper imatrix | |
| - **Compatible** with llama.cpp inference stack | |
| ## Use Cases | |
| - π€ **Edge AI** β Run large models on limited hardware | |
| - π **Browser-Based Inference** β Smaller models for WebGPU/WebGL | |
| - π± **Mobile Deployment** β Fit large models on phones/tablets | |
| - π **High-Throughput APIs** β Serve more requests with less VRAM | |
| - πΎ **Archive Storage** β Preserve models at minimal storage cost | |
| ## Technical Philosophy | |
| The Geeked Out Quantizer focuses on: | |
| 1. **Quality Preservation** β Never sacrifice more quality than necessary | |
| 2. **Automation** β Minimize manual tuning through intelligent defaults | |
| 3. **Hardware Awareness** β Adapt to the system's capabilities | |
| 4. **Production Ready** β Robust error handling and retry logic | |
| 5. **Calibration Quality** β Emphasize representative data selection | |
| ## Model Curation | |
| Not all models are equal candidates. The quantizer evaluates: | |
| - Source format quality (BF16 preferred) | |
| - Model architecture compatibility | |
| - Existing quantization state | |
| - Expected use case alignment | |
| ## Calibration Best Practices | |
| The quality of your quantized model depends heavily on calibration data: | |
| β **DO:** | |
| - Use domain-relevant text (code for code models, medical for medical models) | |
| - Include diverse topics and writing styles | |
| - Provide 100-500 chunks of typical document length | |
| - Ensure natural token distribution | |
| β **DON'T:** | |
| - Use repetitive or overly simple text | |
| - Include corrupted or random data | |
| - Rely on single-domain text for general-purpose models | |
| ## Collaboration & Research | |
| The Geeked Out Quantizer methodology is available for: | |
| - Research collaborations on quantization techniques | |
| - Edge deployment optimization projects | |
| - Custom calibration strategies for specialized domains | |
| - Hardware-specific optimization studies | |
| ## Community | |
| All models in this Hugging Face profile are quantized using this toolchain. Each model card includes: | |
| - Quantization specifications | |
| - Calibration methodology | |
| - Quality metrics | |
| - Use case recommendations | |
| ## Future Directions | |
| - Expanded format support (new GGML quantization types) | |
| - Domain-specific calibration datasets | |
| - Hardware-specific optimization profiles | |
| - Batch processing automation | |
| --- | |
| *The Geeked Out Quantizer: Making extreme compression intelligent.* | |
| For questions about quantization methodology, collaboration opportunities, or technical discussions, please open an issue or discussion on any model in this profile. | |