LoopedLM-P2-d4-m / README.md
HuskyDoge's picture
Add the arXiv link
c4f01ef verified
|
Raw History Blame Contribute Delete
2.89 kB
metadata
license: apache-2.0
library_name: xllm
pipeline_tag: text-generation
tags:
  - xllm
  - looped-language-model
  - paper-part2
  - dense

d4-m

Table 4, M, D4 of Towards Looped Models Done Right. Part II: Rethinking at Fixed Points.

Loading this checkpoint requires the xLLM code. The weights are stored in xLLM's native format (BF16 Safetensors with an xLLM config.json and artifact_manifest.json). This is not a Hugging Face transformers checkpoint: AutoModel.from_pretrained cannot load it. Get the code at https://github.com/ifm-ai/xllm-loop and load the directory with xllm.paper_part2.artifacts.load_artifact.

Model details

Model dense Transformer of 4 blocks, untied output layer; width 3,072, 48 heads (12 KV heads), SwiGLU FFN 8,704
Parameters 810,052,608, including the embedding and output layers
Evaluation full KV cache
Training recipe m_d4: 20,480 updates of 512 x 8,192 tokens (85.90B), AdamW, warmup-stable-decay
Weights BF16 Safetensors: the training checkpoint's FP32 weights rounded to BF16
Tokenizer tokenizer/, the Jais64k tokenizer of training

Download

hf download IFM/LoopedLM-P2-d4-m --local-dir d4-m

The xLLM loader checks the directory against artifact_manifest.json: it rejects symbolic links and files the manifest does not list, apart from the .gitattributes file and the .cache/huggingface/ folder that hf download --local-dir adds. Download into a directory as above, not into the Hub cache (~/.cache/huggingface/hub), whose files are symbolic links.

Use

Evaluate with eval_paper_part2.py from the xLLM repository:

ENABLE_FLASH_ATTENTION_3=true python eval_paper_part2.py --artifact d4-m \
    --data /path/to/eval-data/data.json --out out ppl

data.json and the evaluation inputs come from release/paper-part2/prepare-eval-data.py --output /path/to/eval-data.

xllm.paper_part2.artifacts.load_artifact("d4-m") verifies the directory and returns the model config fields and the BF16 state dict on CPU.

artifact_manifest.json records the size and SHA-256 of every file in this repository.

Paper and citation

Towards Looped Models Done Right. Part II: Rethinking at Fixed Points: https://arxiv.org/abs/2610.06833

@article{huang2026fixedpoints,
  title   = {Towards Looped Models Done Right, Part II: Rethinking at Fixed Points},
  author  = {Benhao Huang and Chufan Shi and Junlin Chen and Shicheng Wen and Zhengzhong Liu and Eric Xing and Xuezhe Ma},
  journal = {arXiv preprint arXiv:2610.06833},
  year    = {2026}
}

License

The weights are released under the Apache License 2.0 (LICENSE); NOTICE records the tokenizer's attribution and how the artifact was prepared. The xLLM code is distributed under its own license.