Instructions to use ceferra/qwen2.5-7b-adele-annotator with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use ceferra/qwen2.5-7b-adele-annotator with PEFT:
from peft import PeftModel from transformers import AutoModelForCausalLM base_model = AutoModelForCausalLM.from_pretrained("Qwen/Qwen2.5-7B-Instruct") model = PeftModel.from_pretrained(base_model, "ceferra/qwen2.5-7b-adele-annotator") - Notebooks
- Google Colab
- Kaggle
Distilled Qwen2.5-7B ADeLe Demand-Level Annotator (QLoRA adapter)
A QLoRA-fine-tuned adapter for Qwen/Qwen2.5-7B-Instruct that predicts
ADeLe v1.0 / DeLeAn demand-level rubric scores (18 rubrics, 0-5 scale),
distilled from three large annotator LLMs (Qwen2.5-72B-Instruct-AWQ,
Llama-3.3-70B-Instruct-AWQ, Mistral-Large-Instruct-2411-AWQ).
Full project, training/eval scripts, and results: https://github.com/ceferra/adele-annotation-distillation
Based on the official ADeLe project: https://github.com/Kinds-of-Intelligence-CFI/ADeLe-AIEvaluation
Training data
Official ADeLe v1.0 battery (16,108 items, 63 tasks). Tasks were split
51 (train) / 6 (validation) / 6 (test), by task name (seed 42), so that
validation/test tasks are never seen during training — see
docs/split_manifest.json in the main repo for the exact task lists.
Training hyperparameters
- Base model:
Qwen/Qwen2.5-7B-Instruct - Method: QLoRA, 4-bit NF4 quantization
- LoRA: r=16, alpha=32, dropout=0.05
- Target modules:
q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj - Max sequence length: 4096
- Epochs: 1.0
- Learning rate: 2e-4
- Effective batch size: 4 GPUs x 1 (per-device) x 16 (gradient accumulation)
Evaluation summary
Against the actual GPT-4o annotations present in the project's held-out benchmark data (5 external benchmarks: MBPP, TheoremQA, Omega, RelBench, SWE-bench-Verified; rWG within-group agreement index, mean over 18 rubrics):
| Annotator | rWG vs. GPT-4o |
|---|---|
| Qwen2.5-72B-Instruct-AWQ | 0.892 |
| Llama-3.3-70B-Instruct-AWQ | 0.908 |
| Mistral-Large-Instruct-2411-AWQ | 0.918 |
| Distilled Qwen2.5-7B (this adapter) | 0.934 |
| Median ensemble of all 4 | 0.944 |
The distilled 7B model matches or exceeds each individual large model's agreement with GPT-4o, at a small fraction of the inference cost, on 17 of 18 rubrics (the exception being rubric KNa, where all candidates, including the full ensemble, score below 0.78).
See the main repository's results/rwg_gpt4o_analysis/ and
reports/DeLeAn_Full_Project_Report.pdf for full details and all 18
per-rubric numbers.
How to use
from peft import PeftModel
from transformers import AutoModelForCausalLM, AutoTokenizer
base = AutoModelForCausalLM.from_pretrained("Qwen/Qwen2.5-7B-Instruct")
model = PeftModel.from_pretrained(base, "path/to/this/adapter")
tok = AutoTokenizer.from_pretrained("path/to/this/adapter")
For fast batch inference, merge the adapter into the base model first
(see scripts/merge_lora.py in the main repository) and serve with vLLM.
License
MIT. Please cite the ADeLe paper (Zhou et al., Nature 2026; arXiv:2503.06378) and the official ADeLe repository (https://github.com/Kinds-of-Intelligence-CFI/ADeLe-AIEvaluation) if you use this adapter, since it is trained entirely on ADeLe's official task battery.
Framework versions
- PEFT 0.13.2
- Downloads last month
- 13