OK, and? Prove it!

#1
by darkmatter2222 - opened

"This model (both 4 bit and 8 bit) exceeds the base Qwen 3.6 27B in 6 out of 7 benchmarks, and matches it on the 7th AND exceeds all 7 benchmarks for Qwen3.6-35B-A3B"

Show us the evidence!

Would love some benchmarks!
do you plan on extending the context to 1m?

Q: "Would you mind to share the real benchmark results please?"
A: "Trust me bro, it's the best."

πŸ˜†

Arent these the benchmarks?
arc/c arc/e boolq hswag obkqa piqa wino

Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-NEO-MAX-MTP-GGUF [instruct mode]
mxfp8 0.711,0.879,0.910,0.790,0.514,0.823,0.763
mxfp4 0.701,0.873,0.909,0.786,0.488,0.813,0.759

Qwen3.6-27B-Instruct: [base, non heretic]
mxfp8 0.647,0.803,0.910,0.773,0.450,0.806,0.742

Qwen3.6-35B-A3B-Instruct [base, non heretic]
mxfp8 0.581,0.757,0.892,0.751,0.428,0.803,0.688

Qwen3.5-27B-Instruct: [base, non heretic]
mxfp8 0.557,0.711,0.868,0.533,0.452,0.706,0.695

I've had great success with DavidAU's models, they handle the same complex skills/prompts I use with Claude, while even the stock versions break down with subagents and enter tool loops.

for Claude Code use the froggeric/Qwen-Fixed-Chat-Templates template. Otherwise it doesnt work for some reason

Hey All;

RE org posters question and answered by @Hunterx

These represent the final benchmarks, not the benchmarks during the process to reach these.

Each stage was benched to check if the process[es] were going in the right direction or not.
To be blunt, there was a lot of "nots" during this journey.

In addition, human checks were performed using the org model and new trained model for direct one on one comparison.
The reason for this: Benchmarks can be wrong.

In tuning, and benching models cases have occurred where the model tests "perfectly" yet the model is (IMO) broken / did not meet quality standards.
So the final step is always human testing.

@Hinks

RE: 1 million ;
Not at this time.

Reason: Because of odd issues during training (with changing context), all the training of the model must be at 1 million context settings.
Although it is possible to "switch" it to 1 million with the source , the results this way (vs training from start) are usually poor.

SWE-bench Verified score at q8 and q4?

I'm testing it, it's good, thank you! :)

@ahinks

Update:
We have a version running at 1 million context, with ARC-C at 700 (just over 700..) for both 8 bit and 4 bit.
The 256k version is running at ARC-C of 712, 706 [8/4 bit] ;

Here is the bar chart of the benchmarks.
grafik

Sign up or log in to comment