Instructions to use FluidInference/clef-text-0.6b with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use FluidInference/clef-text-0.6b with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="FluidInference/clef-text-0.6b")# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("FluidInference/clef-text-0.6b") model = AutoModelForCausalLM.from_pretrained("FluidInference/clef-text-0.6b", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Download README.md from FluidInference/clef-text-0.6b: direct link, hf CLI and curl.
- Browser
- Download file 3.27 kB
-
https://huggingface.co/FluidInference/clef-text-0.6b/resolve/main/README.md
- Command line
-
hf download hf://FluidInference/clef-text-0.6b/README.md
-
curl -L -o README.md https://huggingface.co/FluidInference/clef-text-0.6b/resolve/main/README.md
license: apache-2.0
base_model: Qwen/Qwen3-0.6B
base_model_relation: finetune
pipeline_tag: text-classification
library_name: transformers
tags:
- decision-model
- typed-output
- classification
- clef
- systemone
- distillation
- qwen3
- fluidinference
clef-text-0.6b
A 0.70B-parameter text decision model distilled from Cloudflare's clef-flash (9B): Qwen3-0.6B (LoRA merged) + Clef's joint schema head. Same Jev / SystemOne contract: one probability per option of every typed question, one forward pass. Good at routing and classification; not a knowledge model (see the table).
Apple silicon / Neural Engine build: FluidInference/clef-text-0.6b-coreml.
Usage
joint_schema_model.py is Cloudflare's unchanged Clef module (record encoding, head, ClefModel).
import json, sys, torch
from huggingface_hub import snapshot_download
from safetensors.torch import load_file
from transformers import AutoModelForCausalLM, AutoTokenizer
path = snapshot_download("FluidInference/clef-text-0.6b")
sys.path.insert(0, path)
from joint_schema_model import ClefModel, JointSchemaHead, collate_records, encode_record
tokenizer = AutoTokenizer.from_pretrained(path)
backbone = AutoModelForCausalLM.from_pretrained(path, dtype=torch.float32)
head = JointSchemaHead(**json.loads(open(f"{path}/joint_head_config.json").read()))
head.load_state_dict(load_file(f"{path}/joint_head.safetensors"))
model = ClefModel(backbone, head).eval()
record = {"state": "Our checkout is down and customers can't pay.",
"questions": {"team": {"type": "choice", "instructions": "Which team should handle this?",
"criteria": {"billing": "Payments", "engineering": "Bugs or outages"}},
"urgency": {"type": "score", "criteria": ["Low", "Normal", "High", "Critical"]}}}
encoded = encode_record(tokenizer, record)
with torch.inference_mode():
logits = model(collate_records([encoded], tokenizer.pad_token_id, torch.device("cpu")))[0]
for question, question_logits in zip(encoded.questions, logits):
print(question.question_id, dict(zip(question.option_ids, question_logits.softmax(-1).tolist())))
Quality (held-out test splits, gold accuracy)
| Task | this model | clef-flash 9B |
|---|---|---|
| DBpedia-14 | 99.7 | 100.0 |
| AG News | 90.7 | 92.0 |
| SST-2 | 90.0 | 94.0 |
| BANKING77 | 86.4 | 94.5 |
| TweetEval offensive | 85.0 | 82.3 |
| Yelp stars | 69.0 | 70.7 |
| Emotion | 68.0 | 58.0 |
| BoolQ | 82.3 | 90.7 |
| MNLI | 72.3 | 84.3 |
| ARC-Easy | 80.0 | 100.0 |
| CommonsenseQA | 65.0 | 94.3 |
| ARC-Challenge | 60.7 | 97.3 |
| all (5,124 q) | 79.1 | 88.2 |
Training: 24,208 records from 12 public datasets (train splits) labelled by clef-flash, KL (T = 2) + 0.3 gold CE, LoRA r64 + full head, 2 epochs on an M5 Pro.
License and credits
Apache-2.0. joint_schema_model.py and LICENSE are Cloudflare's (Apache-2.0). Teacher
Cloudflare/clef-flash; base Qwen/Qwen3-0.6B.
By Fluid Inference.