evalexplorer-classify-gguf

GGUF exports of the EvalExplorer document classifier adapters, for llama.cpp. Each folder holds Q8_0, Q5_K_M and Q4_K_M files of the base model with the LoRA merged in, the LoRA itself as GGUF (*-lora-f16.gguf, to apply at load time with --lora on an unmerged base) and the importance matrix used for the K-quants. Every file carries its base model's licence.

The models classify an international development evaluation report from its first pages into evaluation_approach, evaluation_type, temporality, themes and countries. They were trained on the EvalExplorer ingestion pipeline's LLM labels as a quick exploration of what small models can do; the intended next version is trained on the GLM-5.3-Flash relabelling.

Scores

Test split of baobabtech/evalexplorer-data (config classify_codes, 134 documents). Score is the mean field score against the pipeline labels the models learned; vs GLM scores the same answers against the GLM-5.3-Flash relabelling, which the models never saw. Both label sets are unreviewed LLM output. "With schema" constrains the response to a JSON schema of the allowed codes. Speed: llama-server on one A100, 8 parallel slots. With 134 documents, differences below about 0.03 are within sampling noise.

qwen3.5-2b-grpo-countries

Base unsloth/Qwen3.5-2B, adapter baobabtech/evalexplorer-classify-adapters/qwen3.5-2b-grpo-countries.

Variant File Size Score vs GLM Score with schema s/doc
PyTorch bf16 + LoRA (reference) adapter 0.847 0.761 0.67
Q8_0 base + LoRA qwen3.5-2b-grpo-countries-lora-f16.gguf on the base's Q8_0 0.03 GB 0.850 0.764 0.846 0.57
Q8_0 qwen3.5-2b-grpo-countries-Q8_0.gguf 2.08 GB 0.848 0.763 0.848 0.52
Q5_K_M qwen3.5-2b-grpo-countries-Q5_K_M.gguf 1.45 GB 0.841 0.764 0.839 0.52
Q4_K_M qwen3.5-2b-grpo-countries-Q4_K_M.gguf 1.31 GB 0.828 0.741 0.827 0.45

qwen3.5-4b-sft

Base unsloth/Qwen3.5-4B, adapter baobabtech/evalexplorer-classify-qwen3.5-4b-sft.

Variant File Size Score vs GLM Score with schema s/doc
PyTorch bf16 + LoRA (reference) adapter 0.847 0.778 1.22
Q8_0 base + LoRA qwen3.5-4b-sft-lora-f16.gguf on the base's Q8_0 0.06 GB 0.845 0.778 0.843 1.05
Q8_0 qwen3.5-4b-sft-Q8_0.gguf 4.61 GB 0.843 0.778 0.843 0.97
Q5_K_M qwen3.5-4b-sft-Q5_K_M.gguf 3.16 GB 0.842 0.771 0.841 0.95
Q4_K_M qwen3.5-4b-sft-Q4_K_M.gguf 2.78 GB 0.841 0.779 0.838 0.87

gemma-4-e2b-grpo-lr5e6

Base unsloth/gemma-4-E2B-it, adapter baobabtech/evalexplorer-classify-adapters/gemma-4-e2b-grpo-lr5e6.

Variant File Size Score vs GLM Score with schema s/doc
PyTorch bf16 + LoRA (reference) adapter 0.827 0.747 1.22
Q8_0 base + LoRA gemma-4-e2b-grpo-lr5e6-lora-f16.gguf on the base's Q8_0 0.05 GB 0.824 0.750 0.826 0.55
Q8_0 gemma-4-e2b-grpo-lr5e6-Q8_0.gguf 4.97 GB 0.821 0.751 0.821 0.82
Q5_K_M gemma-4-e2b-grpo-lr5e6-Q5_K_M.gguf 3.63 GB 0.817 0.748 0.817 0.56
Q4_K_M gemma-4-e2b-grpo-lr5e6-Q4_K_M.gguf 3.43 GB 0.808 0.755 0.806 0.48

gemma-4-26b-a4b-sft

Base unsloth/gemma-4-26B-A4B-it, adapter baobabtech/evalexplorer-classify-gemma-4-26b-a4b-sft.

Variant File Size Score vs GLM Score with schema s/doc
PyTorch bf16 + LoRA (reference) adapter 0.844 0.803 1.46
Q8_0 gemma-4-26b-a4b-sft-Q8_0.gguf 26.86 GB 0.790 0.783 0.786 2.03
Q5_K_M gemma-4-26b-a4b-sft-Q5_K_M.gguf 19.13 GB 0.790 0.777 0.791 1.86
Q4_K_M gemma-4-26b-a4b-sft-Q4_K_M.gguf 16.80 GB 0.815 0.793 0.804 1.76

Every row links to a full report in baobabtech/evalexplorer-classify-experiments (runs named <adapter>--gguf-<quant>[-lora][-schema]--test).

Findings

  • Q8_0 matches the PyTorch scores within 0.006 for the dense models (Qwen3.5 2B and 4B, Gemma 4 E2B).
  • Q4_K_M costs about 0.02 for the 2B-class models and less than 0.01 for Qwen3.5-4B.
  • The JSON schema changes the scores by less than 0.005 for the dense models: they already return valid JSON.
  • Gemma 4 26B-A4B (MoE) scores about 0.05 below its PyTorch run, mostly on evaluation approach. The GGUF is faithful to the adapter as plain transformers applies it: there, the same adapter scores 0.796 unmerged and 0.794 merged (jobs/merge_check.py), against 0.790 for Q8_0. The 0.844 PyTorch score was measured under Unsloth, whose Gemma 4 MoE path applies the expert LoRA differently. Its LoRA sits on the expert tensors, which llama.cpp's LoRA converter cannot map, so the adapter was merged with PEFT first and there is no LoRA reference row.

Usage

llama-server -m gemma-4-e2b-grpo-lr5e6/gemma-4-e2b-grpo-lr5e6-Q4_K_M.gguf -c 8192

Render the prompt from the dataset's prompt column with the base model's chat template and thinking off (enable_thinking=False), send it to /completion with temperature: 0, and parse the JSON answer. To constrain the output, pass json_schema with the allowed codes per field; jobs/gguf.py in code/ builds it (json_schema) and has the full export and scoring pipeline.

Built with llama.cpp b11361: convert_hf_to_gguf.py and convert_lora_to_gguf.py, llama-export-lora, llama-imatrix on 200 training documents, llama-quantize.

Downloads last month
472
GGUF
Model size
25B params
Architecture
gemma4
Hardware compatibility
Log In to add your hardware

4-bit

5-bit

8-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Space using baobabtech/evalexplorer-classify-gguf 1

Collection including baobabtech/evalexplorer-classify-gguf