LoGRA: Scaling LLM Reinforcement Learning with Low-Rank Gradient Sketches Paper • 2610.06647 • Published 4 days ago • 121
autotrust/GLM5.3-Flash-E224-DGX-Spark Image-Text-to-Text • 128B • Updated about 10 hours ago • 5.42k • 399
STEPQuant: When and Where Errors Matter in Delta-Rule Recurrent State Quantization Paper • 2609.38169 • Published 10 days ago • 101
johnsondaniel/notes-visual-question-answering Visual Question Answering • 49.6k • Updated 17 days ago • 22
johnsondaniel/notes-visual-question-answering Visual Question Answering • 49.6k • Updated 17 days ago • 22
LimiX-2: A Contextual Mechanism Network Towards General Structured-Data Intelligence Paper • 2609.17488 • Published 24 days ago • 817
trl-internal-testing/tiny-Qwen2ForCausalLM-2.5 Text Generation • 2.44M • Updated about 1 month ago • 7.5M • 56
NCP-ArchPreview Technical Report: Moving towards Latent Space Language Models through Next Concept Prediction Paper • 2609.10715 • Published about 1 month ago • 331