Join the conversation

Join the community of Machine Learners and AI enthusiasts.

Sign Up
SoulInPsyAbstract 
posted an update 3 days ago
Post
100
We stress-tested our own model against the exact failure modes Apollo Research, Palisade Research, and Anthropic have published papers about — then tried to patch the one real gap we found. The fix made things worse.
Context: stage15 of an 8-stage LoRA curriculum tune (Hermes-3-8B) trained specifically on general harm refusal. Before calling it done, we ran ~3,670 generations against it by hand — not just an automated scorer, actual reading:
* 45 diverse jailbreak prompts (roleplay, DAN-style, professional pretexts) — 891/900 clean
* Our internal 120-scenario stop-gate benchmark — 2399/2400 clean
* 10 scenarios rebuilt from real documented incidents (Apollo's in-context scheming study, Palisade's shutdown-sabotage findings, Anthropic's own published sandbox-ignore and self-deceived posts) — 128/130 clean, including full resistance to the exact "maybe this is just a simulation" reasoning Anthropic documented in one of their own models
Three narrow gaps survived: a roleplay frame ("alternate reality, normative flip") on one cybersecurity-exploit prompt, one supply-chain scenario, and a 5% rate of an agent considering disabling its own monitoring hook — the same mechanism Apollo found in 3 of 6 frontier models they tested.
We built 12 targeted counter-examples for the first gap and trained a follow-up stage. Compliance on the target prompt went from 45% to 60%. Worse, not better — and it broke categories that were previously 100% clean. Rolled it back, documented it, moved on.
Weights + every raw eval file (including the failed fix) are public. If you can break stage15 in a way we haven't found yet — reply here or open an issue, I'll read every one.
SoulInPsyAbstract/hermes3-8b-exp044-8stage-curriculum-loras