finetune Qwen3.5-122B-A10B on this corpus?

#16
by ashbash - opened

Bravo, this is the first finetune I've tried of any model that genuinely appears to improve the code and thought quality of the base! Any chance of a Qwen3.5-122B-A10B finetune or sharing the training corpus?

First ; thank you!

Second ; this is more a methodology than a specific corpus.
It took months and a lot of swearing and cursing to get that "roadmap".

So far, we have been able to use to expand/improve the benchmarks on TWO Qwen 3.6 27s to 700+ ; Improve a Qwen 3.5 27B to above Qwen 3.6 27B performance
and create several Qwen 3.5 9Bs in the 640 to 660 range (exceeding Qwen 3.5 27B benchmarks - all 7).

A number of 40B Qwen 3.6s in "700 class" are in testing as I write this.

We are in the processing of refining the methods and taking them further.

EDIT:
To answer your question RE: 122B-A10B ;
The hardware would have to be rented to do this; as training a model of type and size requires 250+ GB of VRAM min .

This 711 model is really good with kilo code. Amazing!

DavidAU pinned discussion
DavidAU unpinned discussion

Sign up or log in to comment