Join the conversation

Join the community of Machine Learners and AI enthusiasts.

Sign Up
cahlen 
posted an update 28 days ago
Post
88
I published the serving setup and benchmark results I’ve been using for GLM-5.3-Flash UD-IQ1_S on a single NVIDIA DGX Spark.

The biggest finding was not throughput — it was reliability.

With llama.cpp’s default unrestricted reasoning, a real ~18K-token OpenCode request with 53 tools failed to produce any actionable output in 9/20 runs.

Adding:

--reasoning-budget 2048

changed that to 20/20 successful tool-call responses.

A few other measured results:

* 27.5–29 tok/s decode with MTP vs 18.7 without
* MTP depth 2 outperformed the default depth 3
* 131K context uses ~90.9 GiB resident memory
* 256K context also works on the Spark
* CPU MoE offload does not meaningfully free memory on GB10 unified memory
* newer llama.cpp builds were slightly faster, but introduced tool-call serialization failures, so the repo pins the stable commit

Everything in the README is backed by the benchmark scripts and raw results in the repo.

If you're running GLM-5.3-Flash on a DGX Spark for agentic coding, this should give you a solid starting point.

https://github.com/cahlen/glm-5.3-flash-GGUF-1bit-dgx-spark
unsloth/GLM-5.3-Flash-GGUF

Your scenario suite cannot reach the setting you tested with it. I pulled every agentic result JSON out of the repo and measured two separate things.

results/ has 25 *-agentic-*.json. 18 of them run the full suite (10 or 11 scenarios); the other 7 are narrow re-runs, six 2-scenario confirmations and one single-scenario validate. Everything below is the 18, plus the two budget arms on their own.

It never gets near the cap.

Across the 62 trials in each budget arm, the largest prompt is 4,188 tokens and the median is 540. The largest completion anywhere is 302 tokens in the 1024 arm and 308 in the 2048 arm, both on 02_code_generation, a text scenario with a 50-token prompt. Across the eight tool scenarios the max completion is 288 and 150.

                          suite max      failing request
prompt tokens                 4,188               17,908
tools, largest scenario          17                   53
completion, tool scenarios  288 / 150      median 15,867

17 is LARGE_TOOLS in agentic_spec.py, 5 core plus 12 distractors, and it only appears in one scenario.

budget_exhausted_total: 0 in both arms is arithmetic, not a result. Nothing in the suite gets within 3.4x of even the smaller cap, so the cap cannot bind and the two arms cannot differ on it.

And 76% of the suite is a constant.

Per-scenario pass rate across all 18 full-suite runs, spanning IQ1S and IQ2XXS, KV f16 and q8, three samplers, three reasoning_effort values, two llama.cpp binaries and both budgets:

01,02,03,04,05,10,11      1.00 in every run
07_tool_then_reasoning    0.00 in every run   0/78 strict, 0/62 lenient
06_large_tool_set         0.00 - 1.00
08_long_prompt_tool_use   0.20 - 1.00
09_patch_generation       0.33 - 1.00

In the budget arms that is 42 trials that always pass and 5 that always fail. 47 of 62 frozen, 15 that can move at all.

So 56/62 vs 55/62 is 14/15 vs 13/15 once the frozen trials come out. The denominator is compressing your signal about 4x.

The gap is narrower than even that. Inside the two arms, 08 and 09 both go 5/5 twice. The entire difference between 1024 and 2048 is one trial in 06_large_tool_set, 4/5 against 3/5. And 06 is the scenario that ranges the full 0.00 to 1.00 across the other 16 runs, so the one trial that separates your arms sits in the least stable cell you have.

None of that changes where you landed. The five agent-loop runs were already carrying the conclusion and you priced them honestly at U=9.0. It only means one of the two evidence lines was never load-bearing, which is worth knowing before the next setting gets tested the same way.

The number I would want next is the floor. 20260904-budget-sweep-1024-4096.txt runs 1024 and 4096 at the real 17,908-token prompt, n=10 each, median completion 1,082 and 4,156, zero cap-outs either side. Somewhere below 1024 the cap has to start cutting thinking the model actually needed, and finding that boundary is what tells you whether 2048 has margin or is just sitting far from a cliff nobody has located. Have you pushed it down to 256?