Request for oQ6e and oQ4e variants

#1
by TheOlRazzleDazzle - opened

Hi @gittyeric ,

Thanks for your contribution to the oQe family of models!

I share a similar experience in that Minimax-M2.7 is still exceptional in this class.

I am running a smaller EXO cluster so a bit tighter on room for KV cache with this current precision - any chance you would be willing to create the oQ6e and/or oQ4e variants?

Much appreciated and happy generating🫶

I suppose I can get around to this once some batch stuff I'm running is down, but I mostly added the oQ8 quant because there were no oQ's at all, there were however oQ4's and oQ6's already if you search "https://huggingface.co/models?sort=trending&search=minimax+oq". To be honest, I'm not really noticing any benefit from the -e imatrix bit but just feelsies. In any case I can generate the biggest one you can fit and you can use those to get a check on what fits for you.

Here they are! oQ4e and oQ6e are now under my account. Pretty soon I'll run terminal and SWE benchmarks against all 3 quants since I'm a bit curious how they all hold up and append to the READMEs.

Shoot me a reply on how the different quants work for ya since the benchmarks are sorta made up anyway, curious how the oQ bit + the e bit works for you since on paper it should top any other quants on HF.

Sign up or log in to comment