We're announcing our BananaMind 2.1 model series! The models will include: - BananaMind 2.1 Nano: 10M parameters with 60B tokens. - BananaMind 2.1 Lite: 25M parameters with 40B tokens. - BananaMind 2.1 Flash: 50M parameters with 55B tokens. - BananaMind 2.1 Pro: 135M-145M parameters (still deciding) with 100B tokens. These model will use a multi tower architecture (like BananaMind/BananaMind-2.1-Unified) with some more architectural changes.
BananaMind 2.1 Pro will probrably use 2 no output towers, instead of one!
We're currently training some experimental models based on this architecture to see its scaling!
There has been a leaked memo (now struck down) from the founder of DeepSeek. I'm not here to circulate it, but comment on the minimum-effort evolutionary path he proposed.
This makes sense to me: even at the agent stage I learn world models much faster than when I learned LLM at the LLM stage.
But this means humans are still needed beyond the digital singularity, until robots can close their own loop: eval, manufacturing, self improvement, i.e. physical singularity.