It is very good, first local model that answered most of the simple bench questions correctly
I only tested the first 5 so far, and it got 5/5 right on first attempt. I have never seen any local model answer anything past the second question.
Running the q4_k_m , - NON MTP - (for some reason I cant ever get it to work, and if I do its Worst and 3x slower lol rather than faster) on llama.cpp windows 11.
5800x3d + rtx4090 + 32gb ddr4 = 40-50tps which is very usable, great work man
simple-bench-git
Thank You!