Can we get a base model?

#3
by nraxl1 - opened

First of all, thank you for the work you put into progressing the frontier for models that can truly run on actual consumer hardware, as opposed to the "enthusiast hardware" that models around the 30B mark require. The sparsity at this size is also very helpful for people on smaller budgets (eg. students like myself) experimenting with RL as the rollouts are so much faster compared to similarly sized dense model with aggressive KV growth.

A model of this size and architecture seems very suitable for post-training on a humble budget for specific usecases. I believe the community (and I) would greatly appreciate a base model that would be more responsive to post-training than an already trained one. It would be a great service to the community and a gesture of goodwill if you could publish the base model for Ling 3.0 tiny.

Thank you for your work once again.

inclusionAI org

Hey @nraxl1 thanks for your kind words and we really appreciate the feedback! It exactly reflected why we build the "tiny" model.

Please stay tuned!

Moreover, if you are into post training, please check out Areno. https://github.com/inclusionAI/AReno
It is a local LLM post-training toolkit for RL, SFT/DPO-style training, serving, and agentic RL, originally designed and developed by engineers from the ASystem Team at InclusionAI.

The tiny and Areno combo would go well together, like bread and butter, fish and chips, etc

Please also hit our Discord channel if you are interested to learn more. We would provide news, updates, tutorials, FAQs there~ https://t.co/6LEkFlo2cq
Cheers!

Sign up or log in to comment