Dense Split H/L · 336B tokens
Checkpoint of the dense_split_hl recipe at the 336B-token endpoint (optimizer step 80,000),
from Towards Looped Models Done Right. Part I: Topology, Input Injection, Recurrent-State Organization.
dense_split_hl applies each six-layer H/L module once per state update; dense_split_hl_x2 applies it twice. Both use eight state updates, with input injection and cross-state addition once per update, and share the same parameter count and checkpoint tensor layout.
Loading this checkpoint requires the xLLM code. The weights are stored in xLLM's native format (BF16 Safetensors with an xLLM
config.jsonandartifact_manifest.json). This is not a Hugging Facetransformerscheckpoint:AutoModel.from_pretrainedcannot load it. Get the code at https://github.com/ifm-ai/xllm-loop and load the directory withxllm.paper_part1.native_inference.load_native_model.
Model details
| Field | Value |
|---|---|
| Recipe | dense_split_hl |
| Family | Dense H/L state variant |
Architecture (model.arch) |
huginn |
| Endpoint | 336b: optimizer step 80,000, 335,544,320,000 training tokens |
| Parameters (resident) | 1,499,030,016 |
| Parameters (activated per token) | 1,499,030,016 |
| Training seed | 1 |
| Optimizer | AdamW; weight decay 0.1 |
| LR schedule | cosine; 200-step warmup; 119,210-step horizon (the 500B endpoint) |
| Global batch | 512 sequences × 8,192 tokens = 4,194,304 tokens/step |
| Weights | BF16 Safetensors, 1 shard of at most 5 GiB |
| Tokenizer | tokenizer/, derived from IFM/K2-Horizon-3.7B (see NOTICE) |
All Part I endpoints share one 119,210-step learning-rate schedule, so this 336B-token checkpoint is taken mid-schedule, before the cosine decay completes at the 500B-token endpoint.
Download
hf download IFM/LoopedLM-P1-dense-split-hl-336b --local-dir models/dense_split_hl/336b
The xLLM loader checks the directory against artifact_manifest.json: it rejects symbolic links
and files the manifest does not list, apart from the .gitattributes file and the
.cache/huggingface/ folder that hf download --local-dir adds. Download into a directory as
above, not into the Hub cache (~/.cache/huggingface/hub), whose files are symbolic links.
Use
Native inference requires a CUDA GPU and FlashAttention 3; the xLLM repository documents installation.
export ENABLE_FLASH_ATTENTION_3=true
from xllm.paper_part1.native_inference import generate_native, load_native_model
model, tokenizer, config = load_native_model("models/dense_split_hl/336b")
tokens = generate_native(model, tokenizer, ["The capital of France is"],
max_gen_len=32, use_sampling=False)
print(tokenizer.decode(tokens[0]))
The loader verifies the directory against artifact_manifest.json, then strictly loads tensor
names, shapes and dtypes. Evaluate with the paper protocol:
ENABLE_FLASH_ATTENTION_3=true python eval_paper_part1.py --artifact models/dense_split_hl/336b \
--tasks-root /path/to/composite-eval-root --output /path/to/evaluation-output \
--protocol paper-part1
Retrain from scratch with the same recipe (the base config holds the data, tokenizer, logging and checkpoint-retention settings):
ENABLE_FLASH_ATTENTION_3=true python train_paper_part1.py --recipe dense_split_hl --target 336b \
--base-config /path/to/base.json --dump-dir /path/to/new-run
Files
config.json: the recipe'smodelfields with the training checkpoint's values, and thetokenizersettings withtokenizer.pathset totokenizer/.model.safetensors.index.jsonandmodel-*-of-*.safetensors: BF16 weights; no tensor is split across shards.tokenizer/:tokenizer.json,tokenizer_config.json,special_tokens_map.json.artifact_manifest.json: size and SHA-256 of every file.LICENSE,NOTICE.
Optimizer, dataloader and training state are not included; resuming a run mid-training needs the original training checkpoint.
Paper and citation
Towards Looped Models Done Right. Part I: Topology, Input Injection, Recurrent-State Organization: https://huskydoge.github.io/husky-blog/posts/recursive_models/towards-looped-models-done-right/
@misc{huang2026loopedmodels,
title = {Towards Looped Models Done Right. Part I: Topology, Input Injection, Recurrent-State Organization},
author = {Benhao Huang and Chufan Shi and Junlin Chen and Shicheng Wen and Zhengzhong Liu and Eric Xing and Xuezhe Ma},
year = {2026},
url = {https://huskydoge.github.io/husky-blog/posts/recursive_models/towards-looped-models-done-right/}
}
License
The weights and tokenizer are released under the Apache License 2.0 (LICENSE). NOTICE
records the tokenizer's attribution to
IFM/K2-Horizon-3.7B
and its modifications. The xLLM code is distributed under its own license.
- Downloads last month
- 132