40B : "There is a BIGGER storm coming... and it will take no prisoners."
1290 Tensors, 96 layers : 50% larger than 27B Qwen 3.6 it is based on.
ABOUT:
40B version(s) based on the wildly powerful "711" which operates in/near "OpenAI, Claude and Gemini" closed source intelligence levels:
(Benchmarks of the "711" at the above repo that surpass the Qwen 3.6 27B in all respects)
TESTING:
- Prototype 1, light tune for testing of TWO 700+ ARC-C (OpenAi,Claude, and Gemini level intelligence) Qwen 3.6B super tunes joined "at the hip" to make a 40B monster.
- Prototype 4, Model "Fable-Fusion-711" joined "at the hip" with Deckard 27B (Qwen 3.6) to make a 40B monster in pre-tune testing.
- Testing / benching and post expansion tuning in progress.
IMPORTANT:
- THIS IS A WORK in PROGRESS, and will highlight some parts of this process (40B version of "711") as we proceed.
- NAME of this repo will CHANGE as the project proceeds.
- RUNNING Benchmarks [subject to change] below.
PROJECT NOTES:
- If you want to join the waitlist, you will be notified by email when the FINAL, fully tested and optimized version releases.
- During expansion is normal to lose some benchmarks levels, which are then restored during the post expansion tuning.
- It will (likely) take a number of tuning rounds/steps/stages to bring the 40B up to "700" club status.
RUNNING NOTES [reverse order]:
- Alpha 2 and 3, 3b in testing // additional "non trained" expanded also in testing.
- Selecting model(s) for Beta staging in progress.
- Prelim benchmarks for "Qwen3.6-40B-Grand-Intelligence-One-IQ-tune-alpha" (tuned expansion) posted, moving on to next stage(s). Other Alphas are pending too.
- Benchmark below for "Qwen3.6-40B-Grand-Intelligence-One" (the root model for "Qwen3.6-40B-Grand-Intelligence-One-IQ-tune-alpha") BEFORE post expansion (27B to 40B) training.
BENCHMARKS by Nightmedia
arc/c arc/e boolq hswag obkqa piqa wino
Qwen3.6-40B-Grand-Intelligence-One-IQ-tune-alpha3b [mid-light repair, deeper tune only]
mxfp8 0.695,0.864,0.902,0.819,0.494,0.814,0.772
REMARKS: Slightly lower, however 100% stable. SOTA IQ/power at 40B.
Qwen3.6-40B-Grand-Intelligence-One-IQ-tune-alpha3 [mid-light repair tune only]
mxfp8 0.701,0.862,0.903,...
REMARKS: Excellent, but unstable with some prompts.
Qwen3.6-40B-Grand-Intelligence-One-IQ-tune-alpha2 [light repair tune only]
mxfp8 0.690,0.864,0.908,...
Qwen3.6-40B-Grand-Intelligence-One-IQ-tune-alpha [light repair tune only]
mxfp8 0.689,0.859,0.903,...
------------------------------------------------------------
RAW EXPANDED MODEL(s) - pre stage to be tuned/adjusted.
------------------------------------------------------------
Qwen3.6-40B-Grand-Intelligence-One
[NOT TRAINED YET, expansion only]
mxfp8 0.675,0.860,0.900,0.795,0.478,0.801,0.752
Qwen3.6-40B-Grand-Intelligence-Four-raw
[NOT TRAINED YET, expansion only]
mxfp8 0.679,0.853,0.905
------------------------------------------------------------
ORG MODELS FROM QWEN, no tuning, non heretic.
------------------------------------------------------------
Qwen3.6-27B-Instruct: [base, non heretic]
mxfp8 0.647,0.803,0.910,0.773,0.450,0.806,0.742
Qwen3.6-35B-A3B-Instruct [base, non heretic]
mxfp8 0.581,0.757,0.892,0.751,0.428,0.803,0.688
Qwen3.5-27B-Instruct: [base, non heretic]
mxfp8 0.557,0.711,0.868,0.533,0.452,0.706,0.695
NOTES:
- Models are tested in "Instruct" mode because this generally works better with the testing harness.
- Testing via "thinking" mode also shows the metrics (and changes) but not the true extent.
- In actual fact when the model IS in thinking mode, it will exceed INSTRUCT benchmark scores in most cases.
- Downloads last month
- 3