Separating Representation from Reconstruction Enables Scalable Text Encoders Paper • 2607.04011 • Published 23 days ago • 1
news-crawler-LM: A Small Long-Context Model For High-Quality News Crawling Paper • 2607.21284 • Published 4 days ago • 1
view article Article Be Ready Before the Attack: A Practical Guide to Self-Hosting an Open Model for Cyber Defense jeffboudier • 6 days ago • 14
A Sovereign, Open-Source Foundation Model for German and English Paper • 2607.09424 • Published 17 days ago • 15
view article Article Welcome Inkling by Thinking Machines +2 burtenshaw, merve, pcuenq, ariG23498 • 12 days ago • 122
view article Article Native-speed vLLM transformers modeling backend hmellor, lysandre • 19 days ago • 59
KVpop -- Key-Value Cache Compression with Predictive Online Pruning Paper • 2607.05061 • Published 21 days ago • 24
view article Article Hugging Face and Cerebras bring Gemma 4 to real-time voice AI +2 A-Mahla, andito, lvwerra, vyassaurabh • 26 days ago • 86
Apertus Mini Collection Distillations and Quantizations of our models into more compact formats (<8B parameters) • 17 items • Updated Jun 24 • 11
Do We Still Need Fine Tuning? Turkish Sentiment Analysis in the Era of Large Language Model Paper • 2606.29614 • Published 29 days ago • 1
A Study of Temporal Fusion Strategies for Named Entity Recognition in Historical Texts Paper • 2606.27881 • Published about 1 month ago • 1
MultiHashFormer: Hash-based Generative Language Models Paper • 2606.28057 • Published about 1 month ago • 21
MÖVE: A Holistic LLM Benchmark for the German Public Sector Paper • 2606.13111 • Published Jun 11 • 3
On Subquadratic Architectures: From Applications to Principles Paper • 2606.12364 • Published Jun 10 • 24