Synthetic Pre-pretraining Survives Scale, but Not as a Grammatical Prior Paper • 2609.39827 • Published 6 days ago • 12
Synthetic Pre-pretraining Survives Scale, but Not as a Grammatical Prior Paper • 2609.39827 • Published 6 days ago • 12
When Tokenizers Fail: Byte-Level Chunking for Zero-Shot Transfer to Low-Resource Languages Paper • 2608.27658 • Published Aug 27
How Can We Synthesize High-Quality Pretraining Data? A Systematic Study of Prompt Design, Generator Model, and Source Data Paper • 2604.13977 • Published Apr 15 • 3
On the Utility and Factual Reliability of Pruned Mixture-of-Experts Models in the Biomedical Domain Paper • 2607.01444 • Published Jul 1 • 2
moe-pruning-analysis-project/Qwen3-30B-A3B-Instruct-2507-medinst-frequency-2 16B • Updated Mar 12 • 3