Why no comparision to your backbone Qwen3.6-27B?

#20
by weisunding - opened

Sorry, my bad, your backbone is Qwen3.6-35B-A3B.

Benchmark Qwen3.6-27B KAT-Coder-V2.5-Dev VeriLoop Coder-E1
SWE-bench Verified 77.2 69.40 85.20
SWE-bench Multilingual 71.3 63.00 -
SWE-bench Pro 53.5 45.96 62.38

Take benchmarks just as reference. Often, 3A models, perform incredibly good on agentic pipelines, which is much harder to measure, but it's probably what you are gonna use in real life. Probably not worth the x3 slower 27B version, unless you are aiming for something particularly difficult and need every extra bit of compute.

Sign up or log in to comment