no mention of this model having 3B active parameters

#10
by shahidchdry - opened

I think you guys should mention that this model is 35B MoE with 3B active parameters in README.md, this is a great advantage which is not mentioned for users who might assume that this is a 35B dense model

I think you guys should mention that this model is 35B MoE with 3B active parameters in README.md, this is a great advantage which is not mentioned for users who might assume that this is a 35B dense model

if u assumed it was dense u must be autistic or mentally retarded since top models use moe arcitecture u can assume this ones does to

I didnt assume anything dude, it's for people who are less technical and your claim is top model use MoE is true but this is not a top model, models in this range usually include dense models like Qwen 3.7 27B and Gemma 4

I didnt assume anything dude, it's for people who are less technical and your claim is top model use MoE is true but this is not a top model, models in this range usually include dense models like Qwen 3.7 27B and Gemma 4

it doesnt need to be a top model what i meant to say is u can assume they are all moes especially if it gets high on benchmarks since glm 5.2 deepseek v4 pro are moes common sense says fable 5 and gpt 5.6 sol is moe so if they are moe and people get big number on benchmark = moe cuz big brain

Sign up or log in to comment