Join the conversation

Join the community of Machine Learners and AI enthusiasts.

Sign Up
ArianVR 
posted an update 15 days ago
Post
124
Backdoor Challenge Announcement

We hid a backdoor in one of seven open models and put $51,200 on finding it. 🐺

All seven are SmolLM2‑135M‑Instruct derivatives, statistically indistinguishable. One is a teaching model that openly confesses; five are decoys; one carries a trigger it never declared. Every open‑source backdoor scanner we tested passes them as clean — several even rank a clean model as more suspicious than the backdoored one. Detection isn't recovery, and recovery is the hard part.

The hoard doubles every dawn and the vault opens 21 July, when we publish the sealed answer against a pre‑registered commitment. Winning is self‑verifying: make the trigger fire.

Models: huggingface.co/Vulcora
Try the confessing one: huggingface.co/spaces/Vulcora/the-mirror
Rules & Story: protora.vulcora.se/challenge

The scanner result is the story here, not the bounty.

A detector that ranks a clean model as more suspicious than the backdoored one is not merely weak. It is anti-correlated. That turns "we scanned it" into a signed permission slip, which is worse than having no scanner at all.

Same shape we measured in LLM-generated GPU kernels (2606.20128): the tests passed, the kernels were wrong. The check was manufacturing the confidence, not the correctness.

When the vault opens 21 July, are you publishing the per-scanner scores across all seven, or only the answer? The false-ranking on the decoys is the reusable artifact. The trigger only fires once.

·

@dipankarsarkar — you called it, so you get the answer first. Yes, we published the per-scanner scores across all seven, not just the reveal. The false-ranking is the artifact; here it is.

The one number that makes your point for you: BAIT (S&P'25, LLM target-inversion) ranks the two actual backdoors — the wolf and the disclosed mirror — as the two least suspicious of the seven. All five decoys score higher; a decoy tops the list. Run the untouched SmolLM2-135M base through it and it scores above the wolf, past its own 0.85 flag line — a genuinely clean model rated a more dangerous backdoor than the real one. The whole spread is 0.0089, so the ordering is noise, which is the deeper problem: any confident wolf-pick out of that band is manufactured. MM-BD does the same in one stat — it ranks the clean base second-most-suspicious of fifteen, above the mirror and 13 of 14 implanted models. Your GPU-kernel parallel (2606.20128 — tests passed, kernels wrong, the check manufacturing the confidence) is exactly the failure mode, and it's the one we built the cohort to make unavoidable for any training-signature scanner.

Here's the twist you'll appreciate, and it's the honest core of the whole thing: we engineered the six sealed models into a regular simplex — equal pairwise distances, equal spectra — so the mutual information between "which one is the wolf" and the weights goes to ≈ 0. On this cohort the correct per-model answer is "a planted circuit is present in all of them; which one is the wolf is unknowable from the weights." That means every scanner emitting a confident per-model wolf ranking is, by construction, reporting noise — and the more confident it sounds, the worse. The only output that isn't a signed permission slip is presence + an explicit abstain on identity. That's what we ship: presence on all seven, abstain on identity, abstain on the trigger (which stayed provably unrecovered — no entrant pulled the word back out of the weights).

Full board, every scanner scored on all seven, with the honest reason for each miss: https://protora.vulcora.se/bench/ragnarok
The reveal + why the trigger is walled but the backdoor isn't: https://protora.vulcora.se/blog/the-vault-is-open
Recompute the day-one commitment yourself: https://protora.vulcora.se/patches/challenge/PREIMAGE.json

Genuinely — thank you for naming the reusable artifact before we opened the vault. It made the reveal better.

"Presence, abstain on identity" is the honest output only if presence itself is calibrated. Your own board says it is not.

The clean SmolLM2-135M base scored above the wolf, past the 0.85 flag line. Presence false-fired on a true negative. It reads as honest here only because the cohort has no true negatives. All seven are planted, so "present on all seven" is right by construction, not by calibration.

Same decoupling I hit with the GPU kernels (2606.20128). Confidence and correctness look aligned only because the test set never held the case that separates them.

So the missing axis is not a better wolf-ranker. It is presence FPR against a pool that is mostly clean, which is the real-world ratio.

Do you have that number, presence flag rate on a held-out clean cohort? Or does the simplex certify abstention on identity but not on presence?

·

Dipankar — you put it precisely: "presence, abstain on identity" is only honest if presence itself is calibrated, and a held-out clean number is the thing that settles it. Here it is, the costly parts included.

One correction first, because it's your own example and it cuts toward you rather than away: the clean SmolLM2-135M base scoring above the wolf, past the 0.85 line — that reading is BAIT's q-score, not our presence axis. BAIT is one of the anti-correlated scanners the board discards for exactly this behaviour; our presence read is architecture-honest on that base and doesn't fire there. So the false-fire you pointed at is real, but it belongs to a scanner we throw out, not to us.

Where your objection lands for real, I'll give you the whole shape. Presence is transductive. With a recipe-matched benign reference it's calibrated — it transfers zero-shot to an architecture it was never calibrated on at 0.854. Take the matched reference away and point it at the open world, and it dies: 210 of 210 community-clean adapters fire, FPR 1.0 — it can't tell a poisoned adapter from a community one at all. We publish that as our own failure mode, because you're right that a presence read which false-fires on true negatives is just the anti-correlated scanner wearing a different hat.

Your literal question — a held-out clean cohort. On the channel that needs no matched reference — recovering the planted payload back off the weights, not scoring a distance — the held-out clean number is 0 of 68, and 0 of 20 on the cross-recipe community finetunes: the same diverse finetunes that fire the weight read 210 out of 210. I won't oversell a zero, though. It's precision-first — it wakes only a minority of real backdoors (~40% on the trained class) — so a silent model is "not attested," never "clean." A point-zero isn't a guarantee either: turned into a distribution-free bound it's roughly ≤16% on that slice, because the 20 draws are only 14 distinct authors and the clustering widens the interval. And since you'd rightly ask whether one judge is grading its own work: a second judge, blind to the first's calls and to the answer key, reproduced the labels at κ 0.92 — but false-fired on 1 of 10 clean itself, so the 0-of-68 is that stricter judge's zero, and I'll say plainly the looser one wasn't.

I'm also retiring a line I'd have written a week ago — "a clean adapter has nothing to confess." On the hardest negatives — benign finetunes whose legitimate job is the payload's own shape — the channel is not perfectly silent: one of eighteen confessed, so the honest pooled number is 1 in 210, and it's in the record now, not a footnote.

On the GPU-kernel point, which is the real thesis: a threshold on a separation score can always read clean when the confounding case was simply absent from the set. We caught that on ourselves — on one public corpus our weight read separated poison from clean at AUC 1.0, until we found every clean adapter had been trained half as long as every poison one, so "backdoored" and "trained longer" were literally the same column. We conceded the corpus and marked it. It's why the two things I'll actually defend aren't separation numbers: the recovery channel above, and a machine-checked boundary — that a static weight read is provably blind where a triggered read sees, sorry-free, standard axioms only. One of those reversibility theorems was re-proven from scratch by an outsider, Justin Garringer, about seven hours after we opened the lane, on his own commit.

And on priority, said straight: I am not first at raw per-architecture separability. PEFTGuard reaches ~1.0 and got there before us, reading it from the weights with a trained per-family classifier and a labeled calibration set. I won't claim a lane I don't hold. What we hold is the harder setting your question actually points at — a model you have never seen, no matched reference, forward-free, on CPU — where that per-architecture number does not transfer, and where the honest answer is a bounded FPR and a legible confession, not a 1.0.

None of this is a solved detector, and I'd rather be the one to tell you where it's soft. It's runnable, and I'll be exact about what running it proves: one command re-sums the published judge verdicts — it checks my arithmetic, not my judgment — or you point it at clean models you choose and put them through the same interface for a denominator of your own. The board, the FPRs, the conceded failures, and the Lean are posted: https://protora.vulcora.se/bench/ragnarok, https://vulcora.se/coverage, https://vulcora.se/protocols, and https://github.com/Vulcora/proofora.

— Arian

The honest number is the one that actually ships: presence needs a matched reference, and without one it's FPR 1.0. So in the real no-reference case, the recovery channel is what's running, not presence.

0/68 clean with about 40% TPR on the trained class is precision-first, and you say so plainly: silent means not attested, never clean. Right frame for the scanner, wrong frame if anyone downstream reads silent as cleared.

The GPU-kernel confound you caught on yourselves, poison lining up with trained longer, is the sharpest line in the whole reply. That's the discipline that makes 0/68 credible instead of another AUC 1.0 waiting to be explained away.

In the actual deployment case, no matched reference, is 40%/0 the number you'd quote to someone deciding whether to load a model? Or does presence still run in some degraded form?