432 GB of ultra-fast HBM4 and up to 23.3 TB/s of memory bandwidth on a single GPU š¤Æ.
Two weeks ago, we got early access to AMD's new Instinct MI455X, and our first goal was simple: make sure š¤ Transformers works on day one.
Over the past few weeks, we worked closely with the AMD team to validate the platform, enable Flash Attention, add torchcodec support for multimodal models, and resolve issues uncovered during testing.
The result: ā 99.5% success rate across our 24 core Transformers model architectures - already on par with our daily CI on previous AMD and NVIDIA platforms.
The hardware is just as exciting. With 432 GB of HBM per GPU, our early capacity experiments showed more than 3Ć the concurrent long-context requests compared to MI300, thanks to the much larger KV cache capacity.
A huge thanks to the AMD team for the early access and the great collaboration!
- Esper 4, our flagship agentic coder: specialist in coding, architecture, DevOps, and MLOps! - Tachibana-Agent, trained only on code for dedicated, predictable deployment! - Guardpoint, our structured medical reasoning model: medical diagnosis, management, knowledge, and understanding in structured, concise form!
We'll be expanding Esper 4 to more models and releasing new models as funding allows - donate for more, faster, better models and datasets: sequelbox/SupportOpenSource
š³ The RoboCasa Kitchen Leaderboard What does it take for a robot to handle kitchen chores the way a person does? It has to see (Vision), understand instructions (Language), and actually act (Action) ā and VLA (Vision-Language-Action) models are emerging as the answer. They're the bridge between large multimodal models and real-world embodied control.
RoboCasa Kitchen is a leading robot-learning benchmark in which a single-arm robot (Franka Panda) performs 24 atomic manipulation tasks ā picking up cups and bowls, opening drawers and doors, turning faucets, pressing buttons, and more ā inside a photorealistic simulated kitchen. Because the layout and object placement are randomized every episode, it tests genuine generalization rather than memorized motions. The score (success rate, SR) is the average fraction of the 24 tasks completed as instructed, measured over multiple seeds so results aren't down to luck.
The catch: this benchmark has no official leaderboard, and protocols (number of demonstrations, evaluation setup) differ from paper to paper, leaving scores scattered. Lining the numbers up naively quickly turns into an apples-to-oranges comparison.
This leaderboard fixes that by collecting published scores with their sources and comparing only what is genuinely comparable. It's split into three tables:
š Kitchen 24-task (matched) ā head-to-head under identical conditions (per the RLDX-1 Technical Report). This is the core ranking you can actually trust. ā Other protocols ā self-reported under different setups (e.g. fewer demos). Not directly comparable, so kept separate. š¤ GR1-Tabletop ā a different, humanoid-based variant suite, separated to avoid confusion.
Any researcher can submit their own model's score directly, and submissions are reviewed before they appear on the board. Every number links to its source paper, so you can verify it yourself.
𧬠Darwin-27B-Opus: 86.9% on GPQA Diamond ā World #5, Zero Training We are excited to share Darwin-27B-Opus, a 27B model that achieved 86.9% on GPQA Diamond ā ranking #5 globally on the HuggingFace leaderboard ā without a single gradient update.
How? Darwin breeds pretrained models through evolutionary FFN crossbreeding. The father (Qwen3.5-27B) provides the reasoning architecture; the mother (Claude 4.6 Opus Reasoning Distilled) contributes structured chain-of-thought knowledge. CMA-ES automatically discovers optimal per-layer blending ratios ā no human tuning required.
The result surpasses the original Qwen3.5-27B (85.5%), GLM-5.1 (744B, 86.2%), and Qwen3.5-122B (86.6%). A 27B model outperforming 744B ā with zero training, zero data, one GPU, ~2 hours.
We also confirmed hybrid vigor on Korean benchmarks: Darwin-27B-KR (2nd generation offspring) surpassed both parents on CLIcK, winning 7 out of 11 categories. The evolutionary optimizer independently assigned 93% of FFN from the Korean-specialized mother while preserving 93% of attention from the reasoning-specialized father ā autonomously validating our core principle: FFN carries knowledge, Attention carries reasoning.
š Public release: 10 days ā 300+ community derivatives, 120K+ downloads.
Gemma 3 (1b, 4b, 12b and 27b) - Uncensored full Reasoning/Thinking models fine tuned using top distill datasets.
20 Gemma 3 models 1B, 4B, 12B and 27B with full reasoning using GLM 4.7 Flash, GPT, Claude and Gemini datasets and more fully fine tuned using Unsloth.
Most models are Heretic'ed (uncensored) first, and tuned second. This vastly improves the model.
Models are also bench marked and in almost all cases exceed org model metrics - and in some cases by a lot.
Enjoy the freedom and more powerful THINKING/REASONING and UNCENSORED Gemma 3s !