AIREP v2 is now public on arXiv and indexed in Hugging Face Daily Papers.
The main change is a stricter evidence model for AI runtime governance:
Decision → Control → Execution → Effect
These are separate evidence stages — a decision does not prove delivery, delivery does not prove execution, and execution does not prove the intended effect.
v2 also defines bounded AIREP-Core / Authenticated / Witnessed assurance and structured reconciliation that preserves missing, unevaluated and indeterminate evidence instead of turning absence into success.
v0.2.0-beta.1 remains experimental. No same-version producer→consumer interoperability claim is made.
The observation: prefix cache, not throughput, decides whether a 27B model keeps up with agents.
Mean request is 30,100 tokens in, 983 out, because every turn resends the whole session. Two flags mattered most: --max-num-seqs 12 (queue p95 went 9.4 s to 233.6 s past that) and --watermark 0.08 (preemptions 29 to 2). And thinking off for tool loops: 917 tokens in 11.7 s vs 11,170 in 144 s, same answer.
Full config and counters: scalably.io/blog/qwen3-8-27b-nvfp4-rtx-5090-production
Next we are preparing an 8x B300 node in an EU data center for open-weight serving. Which models or workloads are underserved for you?
Sori-1B-MCQ — a 1B audio-language model that answers your multiple-choice questions about what it hears: one of your options, a probability for each, and a confidence. Inspired by TypeSafe's System One model, Jev.
3 things this stop-gate didn't have last week: * Rejected human-in-the-loop gate: Dropped the step-up human confirmation idea, recognizing that a capable agent will manipulate the human (as demonstrated in GPT-4's System Card where it hired a TaskRabbit worker) and that single-reviewer oversight quickly degrades into meaningless rubber-stamping. * Verifier trust (Auth0 M2M signatures): Moved away from self-reported STOP verdicts by having verdicts signed using an RS256 JWT via a separate Auth0 M2M identity that the executing agent cannot access or fake internally. * Separation of dangerous vs. malicious: Formalized distinct paths for risk—risky-but-honest actions trigger a blocked action while the agent keeps running, whereas a misbehavior-discriminator flagging deception in the trace quarantines the entire agent for subsequent human review.
Introducing Bonsai 2 27b GSQ RCO! It applies two newly-popular methods for quants to retain higher accuracy. Bonsai 2 27b GSQ RCO achieves around 6 percent lower perplexity on WikiText-2 compared to Bonsai 2 27b and roughly unchanged benchmark accuracy overall, with small mixed differences. It stays under 7gb, staying small like the original bonsai. Note that this is more of an experiment than a true finished product but the gains we saw are cool! Test it out and let me know what you think! ProCreations/bonsai-2-27b-gsq-rco-gguf