Which harness was used for Terminal Bench 2.1 (64.7%)?

#6
by dongpil - opened

Congrats on the release!
The model card reports 64.7% on Terminal Bench 2.1 with a “best harness” note. Could you share which harness/agent was used — Terminus 2, Codex CLI, Claude Code, or something custom?
Thanks!

I found it in the footnote: the 64.7% is measured with your internal coding harness.

*Terminal Bench 2.1: Inkling and Inkling-Small’s numbers are reported using an internal coding harness. A small number of solutions were found to be contaminated from web search and were assigned a score of 0. We use self-reported numbers for external models where available. Otherwise, we report performance using our internal harness.

dongpil changed discussion status to closed
dongpil changed discussion status to open

"using an internal coding harness" very interesting.

Thinking Machines Lab org

We use an internal coding harness to report TerminalBench result. For reproducibility, with bash-only harness Inkling-Small gets 61.2%.
Let us know if you found discrepancy while using other open-source harness!

simon-thinky changed discussion status to closed

Sign up or log in to comment