Ermine: Pretrained A1.1B OLMoE Diffusion Language Models

This repository hosts intermediate training checkpoints (1.1e21 FLOPs) for two ~1.1B-active-parameter masked diffusion language models (MDMs) built on an OLMoE-style Mixture-of-Experts transformer backbone.

Checkpoints

File Run Avg. Accuracy
olmoe-mdm-1b-step170000.tar mdm-1b-rl-2048 57.04
olmoe-mdm-1b-ermine-step170000.tar mdm-1b-rl-ermine-2048 60.01

Training Curves

Full training curves and evaluation logs are available on Weights & Biases: lance_chao/ermine-olmoe-training and lance_chao/ermine-olmoe-zeroshot-qa

Model Description

Architecture: OLMo-style transformer with Mixture-of-Experts FFN blocks (OLMoE-style), RoPE position embeddings, RMSNorm, and SwiGLU activations.

Tokenizer: AI2's GPT-NeoX-based Dolma v1.5 tokenizer (allenai_gpt-neox-olmo-dolma-v1_5)

Data: Tokenized shards of AI2's olmoe-mix-0924 pretraining mix (the same mix used for AI2's official OLMoE-1B-7B-0924 model), packed to 2048-token sequences

Checkpoint Format

Each .tar archive contains a raw PyTorch Distributed Checkpoint (DCP) directory (step170000/), i.e. a set of __<rank>_<shard>.distcp shard files plus a .metadata index — the native sharded-checkpoint format saved by this codebase's FSDP trainer (sharded_checkpointer: olmo_core).

Note: This is not a Hugging Face transformers-ready checkpoint and cannot be loaded with from_pretrained.

License

Apache 2.0 (inherited from the allenai/OLMo codebase this project is built on)

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Collection including chen-hao-chao/ermine-olmoe