How to use from the
Use from the
Transformers library
# pip install -U transformers accelerate
# Load model directly
from transformers import AutoModel
model = AutoModel.from_pretrained("ruhai-lin/MixerLoop", device_map="auto")
Quick Links

MixerLoop

Canonical GDN, MixerLoop T4 and FullLoop T4 checkpoints, organized by architecture, dataset, processed-token budget and training seed. Directories containing only .gitkeep are planned experiments, not available models. Historical releases and their original evaluation results are preserved unchanged under legacy/.

Loading

Directories containing model.safetensors provide HF weights; the architectures require FLA and the MixerLoop model registration code. Install the dependencies from MixerLoop commit 54a37b8 and run from that checkout (or install it as a package). Plain Transformers without these registrations does not recognize these model types.

import fla.models
import custom_models
from transformers import AutoModelForCausalLM, AutoTokenizer

repo = "ruhai-lin/MixerLoop"
run = "mixerloop-13m/climbmix-10B-s1337"
tokenizer = AutoTokenizer.from_pretrained(repo, subfolder=run)
model = AutoModelForCausalLM.from_pretrained(repo, subfolder=run)

Tokenizer and generation files are the actual save_pretrained() outputs. Native training checkpoints, optimizer states and logs are not published here.

Evaluation and budgets

eval/core_eval.csv contains all 22 Karpathy CORE tasks and the Core_v2 aggregate, stored as fractions, not percentages. Evaluation uses the canonical evaluator at the commit above: official Karpathy bundle, in-memory v2 baseline corrections (CommonsenseQA 40.3%, LSAT AR 25%, language identification 25%), FP32 weights and BF16 autocast. Historical scores in legacy/ retain their original protocols and should not be assumed to be CORE v2.

First-batch runs use context 1024 and global batch 128. The 1B runs contain 7,630 optimizer steps (1,000,079,360 processed tokens); the 10B run contains 76,294 steps (10,000,007,168 processed tokens), all seed 1337. Actual parameters: MixerLoop-13m 12,896,380; MixerLoop-100m 100,447,266; GDN-100m 100,444,194. Other directories are reserved experiment labels, not claims about completed training or measured parameter counts.

CORE v2 comparison

Scores are centered CORE v2 scores, stored as fractions. The 13M and 100M models were trained on ClimbMix 10B; the 100M columns use seed 1337.

Because 13M models show greater run-to-run noise, we ran three seeds (42, 1337, 2026) and report their mean.

Task 13M GDN 13M MixerLoop 13M FullLoop 100M GDN 100M MixerLoop 100M FullLoop
hellaswag_zeroshot 0.034898 0.033526 0.029675 0.132842 0.133771 0.167231
jeopardy 0.000315 0.000157 0.000630 0.007085 0.006613 0.010392
bigbench_qa_wikidata 0.040106 0.031396 0.071863 0.214950 0.257074 0.282811
arc_easy 0.155069 0.152076 0.151702 0.317621 0.341190 0.385522
arc_challenge -0.037922 -0.054228 -0.037543 -0.001138 0.018203 0.045506
copa -0.113333 -0.093333 -0.120000 0.000000 0.020000 0.020000
commonsense_qa -0.180258 -0.234217 -0.246564 -0.244278 -0.015177 -0.306011
piqa 0.162858 0.167211 0.166123 0.316649 0.334059 0.355822
openbook_qa 0.035556 0.041778 0.042667 0.090667 0.088000 0.098667
lambada_openai 0.105764 0.117407 0.116825 0.253639 0.246264 0.263536
hellaswag 0.028923 0.029277 0.023479 0.125938 0.126734 0.163513
winograd 0.072039 0.030525 0.050061 0.172161 0.113553 0.172161
winogrande 0.027624 0.020258 0.031834 0.027624 0.000789 0.029203
bigbench_dyck_languages 0.001333 0.008333 0.005667 0.128000 0.172000 0.093000
agi_eval_lsat_ar -0.051208 -0.004831 -0.016425 -0.014493 0.072464 0.049275
bigbench_cs_algorithms 0.239394 0.306566 0.221717 0.351515 0.351515 0.352273
bigbench_operators 0.134921 0.139683 0.139682 0.095238 0.090476 0.090476
bigbench_repeat_copy_logic 0.000000 0.000000 0.000000 0.000000 0.000000 0.000000
squad 0.013844 0.008514 0.011542 0.070766 0.047114 0.049669
coqa 0.044135 0.031484 0.036077 0.113366 0.116623 0.123888
boolq -0.408874 -0.361661 -0.354418 -0.187832 -0.059070 -0.054241
bigbench_language_identification 0.001689 0.007733 0.002755 -0.003467 0.010000 0.002533
Core 0.013949 0.017166 0.014880 0.089402 0.112373 0.108874

Download the comparison CSV.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support