Its better than a lot of models present currently

#30
by testbasis - opened

In my tests Q8 it came out better than the following models with tool use:

Qwen3 Coder Next 8bit

Qwen 3.5 122b A10B HauHau Q5 KM

Qwen 3.6 27B Fable Fus 711 Heretic GGUF Q8_0

Qwen 3.6 27b Abinterated Heretic Uncensored LMX 8 bit

Intern Science org

@testbasis

Thank you for sharing your test results โ€” we truly appreciate the feedback and will continue contributing to the community by improving agentic model capabilities across more real-world use cases.

it seems even better than a lot of frontier models for tasks that really require a lot of information to be connected

The thread highlights aggregate quality, but agent-runtime behavior often diverges from raw benchmark scores. Specifically, how does this model handle recovery loops after a tool call fails (e.g., returning a 404 or invalid JSON)? Does it degrade into infinite retry cycles, or does it gracefully pivot to alternative tools? Also, does the 'better' performance hold up when the required information span exceeds the context window of standard LLMs, forcing reliance on external retrieval rather than parametric memory?

nice ai slop bro. the whole thing we are saying here is the model was one of the first ones that did a huge number of searches to understand the problem before answering.

Sign up or log in to comment