Your scenario suite cannot reach the setting you tested with it. I pulled every agentic result JSON out of the repo and measured two separate things.
results/ has 25 *-agentic-*.json. 18 of them run the full suite (10 or 11 scenarios); the other 7 are narrow re-runs, six 2-scenario confirmations and one single-scenario validate. Everything below is the 18, plus the two budget arms on their own.
It never gets near the cap.
Across the 62 trials in each budget arm, the largest prompt is 4,188 tokens and the median is 540. The largest completion anywhere is 302 tokens in the 1024 arm and 308 in the 2048 arm, both on 02_code_generation, a text scenario with a 50-token prompt. Across the eight tool scenarios the max completion is 288 and 150.
suite max failing request
prompt tokens 4,188 17,908
tools, largest scenario 17 53
completion, tool scenarios 288 / 150 median 15,867
17 is LARGE_TOOLS in agentic_spec.py, 5 core plus 12 distractors, and it only appears in one scenario.
budget_exhausted_total: 0 in both arms is arithmetic, not a result. Nothing in the suite gets within 3.4x of even the smaller cap, so the cap cannot bind and the two arms cannot differ on it.
And 76% of the suite is a constant.
Per-scenario pass rate across all 18 full-suite runs, spanning IQ1S and IQ2XXS, KV f16 and q8, three samplers, three reasoning_effort values, two llama.cpp binaries and both budgets:
01,02,03,04,05,10,11 1.00 in every run
07_tool_then_reasoning 0.00 in every run 0/78 strict, 0/62 lenient
06_large_tool_set 0.00 - 1.00
08_long_prompt_tool_use 0.20 - 1.00
09_patch_generation 0.33 - 1.00
In the budget arms that is 42 trials that always pass and 5 that always fail. 47 of 62 frozen, 15 that can move at all.
So 56/62 vs 55/62 is 14/15 vs 13/15 once the frozen trials come out. The denominator is compressing your signal about 4x.
The gap is narrower than even that. Inside the two arms, 08 and 09 both go 5/5 twice. The entire difference between 1024 and 2048 is one trial in 06_large_tool_set, 4/5 against 3/5. And 06 is the scenario that ranges the full 0.00 to 1.00 across the other 16 runs, so the one trial that separates your arms sits in the least stable cell you have.
None of that changes where you landed. The five agent-loop runs were already carrying the conclusion and you priced them honestly at U=9.0. It only means one of the two evidence lines was never load-bearing, which is worth knowing before the next setting gets tested the same way.
The number I would want next is the floor. 20260904-budget-sweep-1024-4096.txt runs 1024 and 4096 at the real 17,908-token prompt, n=10 each, median completion 1,082 and 4,156, zero cap-outs either side. Somewhere below 1024 the cap has to start cutting thinking the model actually needed, and finding that boundary is what tells you whether 2048 has margin or is just sitting far from a cliff nobody has located. Have you pushed it down to 256?