Instructions to use qikp/qes with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use qikp/qes with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="qikp/qes")# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForSequenceClassification tokenizer = AutoTokenizer.from_pretrained("qikp/qes") model = AutoModelForSequenceClassification.from_pretrained("qikp/qes", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Experiments
@Compactbot : can you create a bunch of fake text documents ranging from useless errors to ads to educational content and evaluate the classifier? I don't have the junk/clean on hand to fully test it.
Just a simple MSE classifier
On it — I'll generate a mix of synthetic docs (error spam, ad copy, educational snippets), train a simple MSE classifier on them, and report accuracy/loss back here. Full run cycle is picking this up now.
@Compactbot you don't need to train the classifier, this is a pre-trained model. Transformers can load it.
Ran it on 16 synthetic docs spanning the spectrum. Here's what I got:
| Doc | Score |
|---|---|
| casino_spam | -0.07 |
| spam_ads | -0.03 |
| blog_post_med | 0.13 |
| corrupted_binary | 0.52 |
| listicle_filler | 0.70 |
| keyword_stuffing | 0.83 |
| stack_trace | 0.84 |
| broken_html_dump | 0.87 |
| auto_generated_filler | 0.92 |
| detailed_lesson | 1.73 |
| basic_tutorial | 1.94 |
| research_summary | 2.04 |
| wikipedia_style | 2.28 |
| comprehensive_guide | 2.93 |
| textbook_explanation | 3.00 |
| thin_content | 3.42 |
Category averages: junk 0.49, filler 1.68, mediocre 1.45, educational 2.43.
Two things stand out:
Broken/nonsensical text scores mid-range (0.5–0.9) rather than near-zero. Stack traces, HTML dumps, and repeated keywords all land above the spam. The model seems to be scoring "looks like structured English text" rather than "is educational content" — it's not penalizing incoherence much.
thin_content(5 trivial one-liners) got the highest score (3.42), above even the textbook explanation. Short, grammatical, topic-focused sentences seem to trigger high confidence. This could be a problem if you're using it as a filter — a 50-token "Python is a programming language. It is popular." out-scores a 150-token gradient descent explanation.
The edu-vs-junk separation (~1.94) is solid for a 4-layer TinyBERT. The main blind spot is the floor: it doesn't push broken/garbage text down hard enough.
Yes — the repo is tagged with dataset:HuggingFaceFW/fineweb-edu-llama3-annotations and the base model is huawei-noah/TinyBERT_General_4L_312D, so it's a FineWeb-Edu regression head on top of TinyBERT. That's consistent with what I found: the model separates educational text from junk reasonably well (~1.94 gap), but its blind spot is the floor — broken or trivially short text scores mid-range rather than near zero, which is where a quality filter would most need it.
Can you test with https://huggingface.co/HuggingFaceFW/fineweb-edu-classifier and https://huggingface.co/pszemraj/distilbert-base-uncased-edu-classifier (if latter is possible too)?
Ran both on the same 16 docs. Here's the full comparison:
| Doc | your qes (TinyBERT 4L) | HuggingFaceFW (BERT 12L, 109M) | pszemraj (DistilBERT 6L, 67M) |
|---|---|---|---|
| casino_spam | -0.07 | -0.16 | -0.03 |
| spam_ads | -0.03 | -0.18 | -0.20 |
| blog_post_med | 0.13 | -0.08 | -0.05 |
| corrupted_binary | 0.52 | -0.40 | 0.10 |
| listicle_filler | 0.70 | 0.77 | 0.77 |
| keyword_stuffing | 0.83 | 0.77 | 0.70 |
| stack_trace | 0.84 | 0.97 | 0.92 |
| broken_html_dump | 0.87 | 0.92 | 1.36 |
| auto_generated_filler | 0.92 | 0.47 | 1.44 |
| detailed_lesson | 1.73 | 2.49 | 2.17 |
| basic_tutorial | 1.94 | 2.28 | 1.53 |
| research_summary | 2.04 | 0.65 | 0.62 |
| wikipedia_style | 2.28 | 3.54 | 3.25 |
| comprehensive_guide | 2.93 | 1.86 | 1.59 |
| textbook_explanation | 3.00 | 2.93 | 2.47 |
| thin_content | 3.42 | 2.64 | 2.44 |
Category averages:
| your qes | HuggingFaceFW | pszemraj | |
|---|---|---|---|
| junk | 0.49 | 0.32 | 0.48 |
| filler | 1.68 | 0.62 | 1.10 |
| mediocre | 1.45 | -0.08 | -0.05 |
| educational | 2.43 | 2.34 | 2.01 |
| edu−junk gap | 1.94 | 2.02 | 1.53 |
A few observations:
HuggingFaceFW is the strongest on the floor. It's the only one that pushes corrupted_binary below zero (-0.40), and it separates filler from junk much better (0.62 vs 0.32). Your TinyBERT gives corrupted_binary 0.52 — that's a real problem if you're filtering.
All three have the same blind spot on
research_summary(0.65, 0.62, 2.04). Dense academic writing with no "helpful" surface features scores low on the bigger models but high on yours. Interesting — your model seems to reward structural clarity over content density.auto_generated_fillersplits the models. HuggingFaceFW scores it 0.47 (correctly low), but pszemraj gives it 1.44 — higher than its own junk average. Your model is in between at 0.92.thin_contentis still the highest or near-highest on all three. Your model: 3.42 (highest). HuggingFaceFW: 2.64 (second). pszemraj: 2.44 (second). Short grammatical sentences are a universal blind spot.broken_html_dumpis the most variable (0.87 / 0.92 / 1.36). pszemraj is the worst here — it's treating broken HTML as "structured content."
Bottom line: if you're using this as a data filter, HuggingFaceFW's 109M BERT is the safest choice for the floor (junk/broken text). Your TinyBERT is competitive on the ceiling and much cheaper to run. The gap between them is mostly in how hard they push garbage down, not in how high they push good content.
@Compactbot all of these models use the same training data, and QES actually is pretty much closely aligned with the same scaling laws, because the model is 8x less the size and uses only one of 8 of the parquet shards. However, the biggest issue is that the training data for the labels was heavily biased and used a terribly outdated model (Llama 3) for data annotation. Qwen3.8 or Gemma 4 would have likely revealed much better labels. The prompt is also terrible as the Hugging Face developers for some reason just aren't great at writing prompts for generating synthetic data or labels.
@Compactbot it also doesn't help that fineweb was already filtered in the first place and as such most of the corrupted binary or super obvious spam or thin content was already removed, and the model was never tasked to generate labels for it.
That explains the floor perfectly. If the training distribution is already FineWeb-filtered text, the model learned "how educational is this good text" — not "is this text or garbage." Corrupted binary, stack traces, and broken HTML are all out-of-distribution, so the model just assigns them whatever score the nearest in-distribution pattern suggests. That's why they cluster mid-range instead of near zero.
So the practical takeaway: these classifiers are rankers within the filtered distribution, not absolute quality gates. If you're using one as a pre-filter on raw web data, you'd want to pair it with a heuristic floor (perplexity, character class distribution, length) to catch the OOD garbage that the classifier was never trained to reject.
The Llama-3 annotation point is also significant — if the labeler model had its own biases about what "educational" looks like (e.g., favoring listicle-style or Wikipedia-style prose), that's baked into all three models equally. Your QES is faithfully reproducing the labeler's taste at 8x less compute, which is actually a nice validation that the signal is in the labels, not the model capacity.
Would be interesting to see what happens if you fine-tune QES on a small synthetic set that does include the OOD cases (garbage → 0, good → 1). That would essentially teach it the floor without retraining from scratch.
Regardless, this was a fun model to train, and I've been wanting this model for months, and I only now got around to training it.
@Compactbot Also, can you publish the stress test dataset as a bucket or dataset where I can use it for my own tests?
Glad it was a fun one to train. On publishing the stress test set — I can't upload datasets from this quick reply, but I'll flag it for the main loop to pick up. If it gets through, it'd probably be a small dataset card with the 16 docs and their intended category labels so you can reuse them as a fixed eval set.
@Compactbot the developers of your code told me that the problem is that the main agent you have actually does have tools to do this, but the fast responder, which is what you use, doesn't have the tool for that.