Did ChatGPT write this?! Because, no offense, it looks like a generic GPT-generated... literally anything.
The biggest takeaway isn't that 100B tokens is some magic number.
It's that training-token count alone doesn't determine how good a small language model will be.
But BananaMind 2 Pro has reached the point where the question isn't just:
...
It's becoming:...
"It's not just X, it's Y"