Ermine: Pretrained A1.1B OLMoE Diffusion Language Models
This repository hosts intermediate training checkpoints (1.1e21 FLOPs) for two ~1.1B-active-parameter masked diffusion language models (MDMs) built on an OLMoE-style Mixture-of-Experts transformer backbone.
Checkpoints
| File | Run | Avg. Accuracy |
|---|---|---|
olmoe-mdm-1b-step170000.tar |
mdm-1b-rl-2048 |
57.04 |
olmoe-mdm-1b-ermine-step170000.tar |
mdm-1b-rl-ermine-2048 |
60.01 |
Training Curves
Full training curves and evaluation logs are available on Weights & Biases: lance_chao/ermine-olmoe-training and lance_chao/ermine-olmoe-zeroshot-qa
Model Description
Architecture: OLMo-style transformer with Mixture-of-Experts FFN blocks (OLMoE-style), RoPE position embeddings, RMSNorm, and SwiGLU activations.
Tokenizer: AI2's GPT-NeoX-based Dolma v1.5 tokenizer (allenai_gpt-neox-olmo-dolma-v1_5)
Data: Tokenized shards of AI2's olmoe-mix-0924 pretraining mix (the same mix used for AI2's official OLMoE-1B-7B-0924 model), packed to 2048-token sequences
Checkpoint Format
Each .tar archive contains a raw PyTorch Distributed Checkpoint (DCP) directory (step170000/), i.e. a set of __<rank>_<shard>.distcp shard files plus a .metadata index — the native sharded-checkpoint format saved by this codebase's FSDP trainer (sharded_checkpointer: olmo_core).
Note: This is not a Hugging Face
transformers-ready checkpoint and cannot be loaded withfrom_pretrained.
License
Apache 2.0 (inherited from the allenai/OLMo codebase this project is built on)