-
Warp-as-History: Generalizable Camera-Controlled Video Generation from One Training Video
Paper • 2605.15182 • Published • 40 -
STALE: Can LLM Agents Know When Their Memories Are No Longer Valid?
Paper • 2605.06527 • Published • 47 -
Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis
Paper • 2605.14392 • Published • 9 -
World Action Models: The Next Frontier in Embodied AI
Paper • 2605.12090 • Published • 73
🏗️ Building on HF
Madalin Tatarciuc
cetusian
AI & ML interests
instinctively gravitating toward the frontier where intelligence meets constraints.
Recent Activity
liked a Space about 2 hours ago
multimodalart/jev-decision-index liked a Space about 7 hours ago
lerobot/visualize_dataset posted an update about 9 hours ago
Surogate Speech is out 🐇 small open speech models for agents, the speech side of Surogate.
First language shipped is Romanian, and the numbers came out better than I expected: 5.69% WER on FLEURS with a 116M model. Canary 1B gets 5.95% on the same clips, Whisper large-v3 8.42%. Leaderboard runner, one RTX 5090.
What's in it:
🎧 jackrabbit-110m-ro: speech recognition, 2,500× real time on one GPU
⚡ jackrabbit-110m-ro-streaming: live, final text about 0.7 s after you stop talking
🗣️ amami-357m-ro: TTS with three voices, runs on a CPU
Serving is one line with our engine: surogate serve --stt surogate/jackrabbit-110m-ro
Transcripts for every clip are public if you want to rescore it.
https://huggingface.co/collections/surogate/surogate-speech-6ab680eb84c7ff75fb73ad5a
https://github.com/invergent-ai/surogate-speech