--- license: apache-2.0 library_name: xllm pipeline_tag: text-generation tags: - xllm - looped-language-model - paper-part2 - dense --- # d4-m Table 4, M, D4 of *Towards Looped Models Done Right. Part II: Rethinking at Fixed Points*. > [!IMPORTANT] > **Loading this checkpoint requires the xLLM code.** The weights are stored in xLLM's native > format (BF16 Safetensors with an xLLM `config.json` and `artifact_manifest.json`). This is not a > Hugging Face `transformers` checkpoint: `AutoModel.from_pretrained` cannot load it. > Get the code at [https://github.com/ifm-ai/xllm-loop](https://github.com/ifm-ai/xllm-loop) and load the directory with `xllm.paper_part2.artifacts.load_artifact`. ## Model details | | | | --- | --- | | Model | dense Transformer of 4 blocks, untied output layer; width 3,072, 48 heads (12 KV heads), SwiGLU FFN 8,704 | | Parameters | 810,052,608, including the embedding and output layers | | Evaluation | full KV cache | | Training | recipe `m_d4`: 20,480 updates of 512 x 8,192 tokens (85.90B), AdamW, warmup-stable-decay | | Weights | BF16 Safetensors: the training checkpoint's FP32 weights rounded to BF16 | | Tokenizer | `tokenizer/`, the Jais64k tokenizer of training | ## Download ```bash hf download IFM/LoopedLM-P2-d4-m --local-dir d4-m ``` The xLLM loader checks the directory against `artifact_manifest.json`: it rejects symbolic links and files the manifest does not list, apart from the `.gitattributes` file and the `.cache/huggingface/` folder that `hf download --local-dir` adds. Download into a directory as above, not into the Hub cache (`~/.cache/huggingface/hub`), whose files are symbolic links. ## Use Evaluate with `eval_paper_part2.py` from the xLLM repository: ```bash ENABLE_FLASH_ATTENTION_3=true python eval_paper_part2.py --artifact d4-m \ --data /path/to/eval-data/data.json --out out ppl ``` `data.json` and the evaluation inputs come from `release/paper-part2/prepare-eval-data.py --output /path/to/eval-data`. `xllm.paper_part2.artifacts.load_artifact("d4-m")` verifies the directory and returns the model config fields and the BF16 state dict on CPU. `artifact_manifest.json` records the size and SHA-256 of every file in this repository. ## Paper and citation *Towards Looped Models Done Right. Part II: Rethinking at Fixed Points*: https://arxiv.org/abs/2610.06833 ```bibtex @article{huang2026fixedpoints, title = {Towards Looped Models Done Right, Part II: Rethinking at Fixed Points}, author = {Benhao Huang and Chufan Shi and Junlin Chen and Shicheng Wen and Zhengzhong Liu and Eric Xing and Xuezhe Ma}, journal = {arXiv preprint arXiv:2610.06833}, year = {2026} } ``` ## License The weights are released under the Apache License 2.0 (`LICENSE`); `NOTICE` records the tokenizer's attribution and how the artifact was prepared. The xLLM code is distributed under its own license.