Instructions to use baobabtech/evalexplorer-classify-gguf with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use baobabtech/evalexplorer-classify-gguf with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf baobabtech/evalexplorer-classify-gguf:Q4_K_M # Run inference directly in the terminal: llama cli -hf baobabtech/evalexplorer-classify-gguf:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf baobabtech/evalexplorer-classify-gguf:Q4_K_M # Run inference directly in the terminal: llama cli -hf baobabtech/evalexplorer-classify-gguf:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf baobabtech/evalexplorer-classify-gguf:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf baobabtech/evalexplorer-classify-gguf:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf baobabtech/evalexplorer-classify-gguf:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf baobabtech/evalexplorer-classify-gguf:Q4_K_M
Use Docker
docker model run hf.co/baobabtech/evalexplorer-classify-gguf:Q4_K_M
- LM Studio
- Jan
- Ollama
How to use baobabtech/evalexplorer-classify-gguf with Ollama:
ollama run hf.co/baobabtech/evalexplorer-classify-gguf:Q4_K_M
- Unsloth Desktop
- Pi
How to use baobabtech/evalexplorer-classify-gguf with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf baobabtech/evalexplorer-classify-gguf:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "baobabtech/evalexplorer-classify-gguf:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use baobabtech/evalexplorer-classify-gguf with Docker Model Runner:
docker model run hf.co/baobabtech/evalexplorer-classify-gguf:Q4_K_M
- Lemonade
How to use baobabtech/evalexplorer-classify-gguf with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull baobabtech/evalexplorer-classify-gguf:Q4_K_M
Run and chat with the model
lemonade run user.evalexplorer-classify-gguf-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use baobabtech/evalexplorer-classify-gguf with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf baobabtech/evalexplorer-classify-gguf:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default baobabtech/evalexplorer-classify-gguf:Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use baobabtech/evalexplorer-classify-gguf with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf baobabtech/evalexplorer-classify-gguf:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "baobabtech/evalexplorer-classify-gguf:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
evalexplorer-classify-gguf
GGUF exports of the EvalExplorer document classifier adapters, for llama.cpp. Each folder holds Q8_0, Q5_K_M and
Q4_K_M files of the base model with the LoRA merged in, the LoRA itself as GGUF (*-lora-f16.gguf, to apply at load
time with --lora on an unmerged base) and the importance matrix used for the K-quants. Every file carries its base
model's licence.
The models classify an international development evaluation report from its first pages into evaluation_approach,
evaluation_type, temporality, themes and countries. They were trained on the EvalExplorer ingestion
pipeline's LLM labels as a quick exploration of what small models can do; the intended next version is trained on
the GLM-5.3-Flash relabelling.
Scores
Test split of baobabtech/evalexplorer-data (config classify_codes, 134
documents). Score is the mean field score against the pipeline labels the models learned; vs GLM scores the same
answers against the GLM-5.3-Flash relabelling, which the models never saw. Both label sets are unreviewed LLM output.
"With schema" constrains the response to a JSON schema of the allowed codes. Speed: llama-server on one A100,
8 parallel slots. With 134 documents, differences below about 0.03 are within sampling noise.
qwen3.5-2b-grpo-countries
Base unsloth/Qwen3.5-2B, adapter baobabtech/evalexplorer-classify-adapters/qwen3.5-2b-grpo-countries.
| Variant | File | Size | Score | vs GLM | Score with schema | s/doc |
|---|---|---|---|---|---|---|
| PyTorch bf16 + LoRA (reference) | adapter | 0.847 | 0.761 | 0.67 | ||
| Q8_0 base + LoRA | qwen3.5-2b-grpo-countries-lora-f16.gguf on the base's Q8_0 |
0.03 GB | 0.850 | 0.764 | 0.846 | 0.57 |
| Q8_0 | qwen3.5-2b-grpo-countries-Q8_0.gguf |
2.08 GB | 0.848 | 0.763 | 0.848 | 0.52 |
| Q5_K_M | qwen3.5-2b-grpo-countries-Q5_K_M.gguf |
1.45 GB | 0.841 | 0.764 | 0.839 | 0.52 |
| Q4_K_M | qwen3.5-2b-grpo-countries-Q4_K_M.gguf |
1.31 GB | 0.828 | 0.741 | 0.827 | 0.45 |
qwen3.5-4b-sft
Base unsloth/Qwen3.5-4B, adapter baobabtech/evalexplorer-classify-qwen3.5-4b-sft.
| Variant | File | Size | Score | vs GLM | Score with schema | s/doc |
|---|---|---|---|---|---|---|
| PyTorch bf16 + LoRA (reference) | adapter | 0.847 | 0.778 | 1.22 | ||
| Q8_0 base + LoRA | qwen3.5-4b-sft-lora-f16.gguf on the base's Q8_0 |
0.06 GB | 0.845 | 0.778 | 0.843 | 1.05 |
| Q8_0 | qwen3.5-4b-sft-Q8_0.gguf |
4.61 GB | 0.843 | 0.778 | 0.843 | 0.97 |
| Q5_K_M | qwen3.5-4b-sft-Q5_K_M.gguf |
3.16 GB | 0.842 | 0.771 | 0.841 | 0.95 |
| Q4_K_M | qwen3.5-4b-sft-Q4_K_M.gguf |
2.78 GB | 0.841 | 0.779 | 0.838 | 0.87 |
gemma-4-e2b-grpo-lr5e6
Base unsloth/gemma-4-E2B-it, adapter baobabtech/evalexplorer-classify-adapters/gemma-4-e2b-grpo-lr5e6.
| Variant | File | Size | Score | vs GLM | Score with schema | s/doc |
|---|---|---|---|---|---|---|
| PyTorch bf16 + LoRA (reference) | adapter | 0.827 | 0.747 | 1.22 | ||
| Q8_0 base + LoRA | gemma-4-e2b-grpo-lr5e6-lora-f16.gguf on the base's Q8_0 |
0.05 GB | 0.824 | 0.750 | 0.826 | 0.55 |
| Q8_0 | gemma-4-e2b-grpo-lr5e6-Q8_0.gguf |
4.97 GB | 0.821 | 0.751 | 0.821 | 0.82 |
| Q5_K_M | gemma-4-e2b-grpo-lr5e6-Q5_K_M.gguf |
3.63 GB | 0.817 | 0.748 | 0.817 | 0.56 |
| Q4_K_M | gemma-4-e2b-grpo-lr5e6-Q4_K_M.gguf |
3.43 GB | 0.808 | 0.755 | 0.806 | 0.48 |
gemma-4-26b-a4b-sft
Base unsloth/gemma-4-26B-A4B-it, adapter baobabtech/evalexplorer-classify-gemma-4-26b-a4b-sft.
| Variant | File | Size | Score | vs GLM | Score with schema | s/doc |
|---|---|---|---|---|---|---|
| PyTorch bf16 + LoRA (reference) | adapter | 0.844 | 0.803 | 1.46 | ||
| Q8_0 | gemma-4-26b-a4b-sft-Q8_0.gguf |
26.86 GB | 0.790 | 0.783 | 0.786 | 2.03 |
| Q5_K_M | gemma-4-26b-a4b-sft-Q5_K_M.gguf |
19.13 GB | 0.790 | 0.777 | 0.791 | 1.86 |
| Q4_K_M | gemma-4-26b-a4b-sft-Q4_K_M.gguf |
16.80 GB | 0.815 | 0.793 | 0.804 | 1.76 |
Every row links to a full report in baobabtech/evalexplorer-classify-experiments
(runs named <adapter>--gguf-<quant>[-lora][-schema]--test).
Findings
- Q8_0 matches the PyTorch scores within 0.006 for the dense models (Qwen3.5 2B and 4B, Gemma 4 E2B).
- Q4_K_M costs about 0.02 for the 2B-class models and less than 0.01 for Qwen3.5-4B.
- The JSON schema changes the scores by less than 0.005 for the dense models: they already return valid JSON.
- Gemma 4 26B-A4B (MoE) scores about 0.05 below its PyTorch run, mostly on evaluation approach. The GGUF is
faithful to the adapter as plain transformers applies it: there, the same adapter scores 0.796 unmerged and 0.794
merged (
jobs/merge_check.py), against 0.790 for Q8_0. The 0.844 PyTorch score was measured under Unsloth, whose Gemma 4 MoE path applies the expert LoRA differently. Its LoRA sits on the expert tensors, which llama.cpp's LoRA converter cannot map, so the adapter was merged with PEFT first and there is no LoRA reference row.
Usage
llama-server -m gemma-4-e2b-grpo-lr5e6/gemma-4-e2b-grpo-lr5e6-Q4_K_M.gguf -c 8192
Render the prompt from the dataset's prompt column with the base model's chat template and thinking off
(enable_thinking=False), send it to /completion with temperature: 0, and parse the JSON answer. To constrain
the output, pass json_schema with the allowed codes per field; jobs/gguf.py in
code/ builds it (json_schema) and has the full export and scoring pipeline.
Built with llama.cpp b11361: convert_hf_to_gguf.py and convert_lora_to_gguf.py, llama-export-lora,
llama-imatrix on 200 training documents, llama-quantize.
- Downloads last month
- 472