Macaron-V1: Towards Open Continual Learning with Self-Improvement and Mixture-of-LoRA Paper • 2608.09819 • Published 3 days ago • 320
trl-internal-testing/tiny-Qwen2ForCausalLM-2.5 Text Generation • 2.43M • Updated Dec 19, 2025 • 14M • 20
Quickest Detection of Hallucination Onset: Delay Bounds and Learned CUSUM Statistics Paper • 2606.12476 • Published Jun 10 • 2
Value-Aware Stochastic KV Cache Eviction for Reasoning Models Paper • 2606.03928 • Published Jun 2 • 8
DelTA: Discriminative Token Credit Assignment for Reinforcement Learning from Verifiable Rewards Paper • 2605.21467 • Published May 20 • 207