Dense Split H/L · 336B tokens

Checkpoint of the dense_split_hl recipe at the 336B-token endpoint (optimizer step 80,000), from Towards Looped Models Done Right. Part I: Topology, Input Injection, Recurrent-State Organization.

dense_split_hl applies each six-layer H/L module once per state update; dense_split_hl_x2 applies it twice. Both use eight state updates, with input injection and cross-state addition once per update, and share the same parameter count and checkpoint tensor layout.

Loading this checkpoint requires the xLLM code. The weights are stored in xLLM's native format (BF16 Safetensors with an xLLM config.json and artifact_manifest.json). This is not a Hugging Face transformers checkpoint: AutoModel.from_pretrained cannot load it. Get the code at https://github.com/ifm-ai/xllm-loop and load the directory with xllm.paper_part1.native_inference.load_native_model.

Model details

Field Value
Recipe dense_split_hl
Family Dense H/L state variant
Architecture (model.arch) huginn
Endpoint 336b: optimizer step 80,000, 335,544,320,000 training tokens
Parameters (resident) 1,499,030,016
Parameters (activated per token) 1,499,030,016
Training seed 1
Optimizer AdamW; weight decay 0.1
LR schedule cosine; 200-step warmup; 119,210-step horizon (the 500B endpoint)
Global batch 512 sequences × 8,192 tokens = 4,194,304 tokens/step
Weights BF16 Safetensors, 1 shard of at most 5 GiB
Tokenizer tokenizer/, derived from IFM/K2-Horizon-3.7B (see NOTICE)

All Part I endpoints share one 119,210-step learning-rate schedule, so this 336B-token checkpoint is taken mid-schedule, before the cosine decay completes at the 500B-token endpoint.

Download

hf download IFM/LoopedLM-P1-dense-split-hl-336b --local-dir models/dense_split_hl/336b

The xLLM loader checks the directory against artifact_manifest.json: it rejects symbolic links and files the manifest does not list, apart from the .gitattributes file and the .cache/huggingface/ folder that hf download --local-dir adds. Download into a directory as above, not into the Hub cache (~/.cache/huggingface/hub), whose files are symbolic links.

Use

Native inference requires a CUDA GPU and FlashAttention 3; the xLLM repository documents installation.

export ENABLE_FLASH_ATTENTION_3=true
from xllm.paper_part1.native_inference import generate_native, load_native_model

model, tokenizer, config = load_native_model("models/dense_split_hl/336b")
tokens = generate_native(model, tokenizer, ["The capital of France is"],
                         max_gen_len=32, use_sampling=False)
print(tokenizer.decode(tokens[0]))

The loader verifies the directory against artifact_manifest.json, then strictly loads tensor names, shapes and dtypes. Evaluate with the paper protocol:

ENABLE_FLASH_ATTENTION_3=true python eval_paper_part1.py --artifact models/dense_split_hl/336b \
    --tasks-root /path/to/composite-eval-root --output /path/to/evaluation-output \
    --protocol paper-part1

Retrain from scratch with the same recipe (the base config holds the data, tokenizer, logging and checkpoint-retention settings):

ENABLE_FLASH_ATTENTION_3=true python train_paper_part1.py --recipe dense_split_hl --target 336b \
    --base-config /path/to/base.json --dump-dir /path/to/new-run

Files

  • config.json: the recipe's model fields with the training checkpoint's values, and the tokenizer settings with tokenizer.path set to tokenizer/.
  • model.safetensors.index.json and model-*-of-*.safetensors: BF16 weights; no tensor is split across shards.
  • tokenizer/: tokenizer.json, tokenizer_config.json, special_tokens_map.json.
  • artifact_manifest.json: size and SHA-256 of every file.
  • LICENSE, NOTICE.

Optimizer, dataloader and training state are not included; resuming a run mid-training needs the original training checkpoint.

Paper and citation

Towards Looped Models Done Right. Part I: Topology, Input Injection, Recurrent-State Organization: https://huskydoge.github.io/husky-blog/posts/recursive_models/towards-looped-models-done-right/

@misc{huang2026loopedmodels,
  title  = {Towards Looped Models Done Right. Part I: Topology, Input Injection, Recurrent-State Organization},
  author = {Benhao Huang and Chufan Shi and Junlin Chen and Shicheng Wen and Zhengzhong Liu and Eric Xing and Xuezhe Ma},
  year   = {2026},
  url    = {https://huskydoge.github.io/husky-blog/posts/recursive_models/towards-looped-models-done-right/}
}

License

The weights and tokenizer are released under the Apache License 2.0 (LICENSE). NOTICE records the tokenizer's attribution to IFM/K2-Horizon-3.7B and its modifications. The xLLM code is distributed under its own license.

Downloads last month
132
Safetensors
Model size
1B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Collection including IFM/LoopedLM-P1-dense-split-hl-336b