distilled-m-random-init
Table 4, M, Distilled, random init of Towards Looped Models Done Right. Part II: Rethinking at Fixed Points: a distilled prefill path for huginn-m-learned-entropy0p01.
Loading this checkpoint requires the xLLM code. The weights are stored in xLLM's native format (BF16 Safetensors with an xLLM
config.jsonandartifact_manifest.json). This is not a Hugging Facetransformerscheckpoint:AutoModel.from_pretrainedcannot load it. Get the code at https://github.com/ifm-ai/xllm-loop and load it with its teacher throughxllm.paper_part2.distill.load_student.
Model details
| Model | a d x d input map and 2 Transformer blocks of the teacher's width |
| Parameters | 217,067,520 |
| Teacher | huginn-m-learned-entropy0p01 |
| Initialization | random |
| Target | the teacher's pre-coda state after 5 recurrences, from its prelude output |
| Loss | hidden-state MSE over the target energy, plus KL from the teacher's next-token distribution |
| Training | recipe m_distill_random_init: 10,240 updates of 256 x 8,192 tokens, AdamW (0.9, 0.95), cosine schedule |
| Prefill | the student's state, one teacher recurrence and the teacher's coda write the four KV banks; the teacher decodes at R = 5 |
| Weights | BF16 Safetensors: the trained FP32 weights rounded to BF16 |
Download
hf download IFM/LoopedLM-P2-huginn-m-learned-entropy0p01 --local-dir huginn-m-learned-entropy0p01
hf download IFM/LoopedLM-P2-distilled-m-random-init --local-dir distilled-m-random-init
The xLLM loader checks the directory against artifact_manifest.json: it rejects symbolic links
and files the manifest does not list, apart from the .gitattributes file and the
.cache/huggingface/ folder that hf download --local-dir adds. Download into a directory as
above, not into the Hub cache (~/.cache/huggingface/hub), whose files are symbolic links.
Use
A distilled student is evaluated together with its teacher, which supplies the prelude, one
recurrence and the coda. Download both, then run eval_paper_part2.py from the xLLM
repository:
ENABLE_FLASH_ATTENTION_3=true python eval_paper_part2.py --artifact huginn-m-learned-entropy0p01 --student distilled-m-random-init \
--data /path/to/eval-data/data.json --out out ppl
data.json and the evaluation inputs come from release/paper-part2/prepare-eval-data.py --output /path/to/eval-data.
The student's config.json pins its teacher's manifest_sha256, the digest recorded in the
teacher's artifact_manifest.json (see Provenance); xllm.paper_part2.distill.load_student
refuses any other teacher artifact, so use the teacher repository at the matching revision.
Provenance
- Teacher manifest_sha256:
bf7c7107c020596a3c7e57c9ead09052e6f8ca65c949e90ead2e1bd763668f2a
artifact_manifest.json records the size and SHA-256 of every file in this repository.
Paper and citation
Towards Looped Models Done Right. Part II: Rethinking at Fixed Points: https://github.com/ifm-ai/xllm-loop/blob/main/papers/part2.pdf
@misc{huang2026fixedpoints,
title = {Towards Looped Models Done Right. Part II: Rethinking at Fixed Points},
author = {Benhao Huang and Chufan Shi and Junlin Chen and Shicheng Wen and Zhengzhong Liu and Eric Xing and Xuezhe Ma},
year = {2026},
url = {https://github.com/ifm-ai/xllm-loop/blob/main/papers/part2.pdf}
}
License
The weights are released under the Apache License 2.0 (LICENSE); NOTICE records the
tokenizer's attribution and how the artifact was prepared. The xLLM code is distributed under its
own license.
- Downloads last month
- 26