damn i was in the process of doing something very similar

#2
by nraxl1 - opened

is the embedding concat thing deepseek engram or longcat's n-gram thing with extended vocab? did you end up writing a custom kernel for the mHC depth attention?

Nanbeige LLM Lab org

Great to hear that we’re exploring similar ideas! 🤝
Our concatenated n-gram embeddings build on LongCat’s extended-vocabulary n-gram approach, with some further modifications.
For mHC with depth attention, we have already implemented a custom kernel for training. An inference kernel will be provided when Nanbeige4.5 is released.

Sign up or log in to comment