view article Article Transformers now runs llama.cpp quants +1 marcsun13, ArthurZ, lysandre • 3 days ago • 60
view article Article Pruning LLMs Like a Physicist: Block Removal as an Ising Optimization Problem MultiverseComputingCAI • 4 days ago • 30
view article Article tokenizers v1: encode, decode and scaling, measured +2 ArthurZ, sbrandeis, mcpotato, lysandre • 4 days ago • 74
view article Article Async GRPO with LoRA across HF Jobs: a bucket, a proxy, and no NCCL +2 aminediroHF, qgallouedec, kashif, sergiopaniego • 15 days ago • 52
view article Article IBM releases SOTA Granite Time Series PatchTST-FM-r2 model with commercial-friendly license ibm-research • 16 days ago • 60
ibm-granite/granite-timeseries-patchtst-fm-r2 Time Series Forecasting • 0.4B • Updated 16 days ago • 170k • 16
view article Article Training a coding model to paint watercolours with TRL and OpenEnv sergiopaniego • 22 days ago • 72
view article Article NeoMME: an efficient Multimodal-native and Multilingual Encoder Hcompany • 22 days ago • 110
VoiceMem: Streaming Dual-Brain Memory for Real-Time Interaction Paper • 2608.26005 • Published about 1 month ago • 163
Self-OPD: On-Policy Distillation for Flow Matching Models without Teacher Paper • 2608.26872 • Published 29 days ago • 65
JIT-Agent: Scaling Harness Intelligence via Just-in-Time Harness Evolution Paper • 2608.25593 • Published about 1 month ago • 69