Miowtion

MiniMax-H3 T2VA · Veda predictor · 8 NFE · step 600 (preview)

The Veda tile-score predictor for MiniMax-H3 text-to-audio-video at eight denoising steps. It selects, per layer and per head, which 128-token key tiles each query tile attends to, so attention runs block-sparse at a 10% keep ratio instead of dense.

Paper · Project page · Code · Deployment guide

Preview checkpoint: trained for 600 updates on 5.17 s clips only. It generalizes to the 10.1 s and 14.4 s geometries below, but was not trained on them.

Files

File
minimax_h3_t2va_veda_8nfe_600step_preview_fp8.safetensors Predictor weights (fp8 e4m3, 275M params) and the tile plans, in __metadata__
config.json Architecture, keep ratio, and a summary of the packed plans
AGENTS.md Step-by-step deployment for Ada (SM89), Hopper (SM90) and Blackwell (SM120)

The predictor is useless without the tile plan its scores are indexed against: a plan fixes the latent grid, the per-head tile shape, and the block ordering. Both live in one file so they cannot be mispaired — a wrong pairing is silent, the predictor still emits scores, just for a tiling it never saw.

Precision: why fp8

The weights are stored as float8_e4m3fn with one fp32 amax scale per head, kept beside each tensor as <key>.__scale. miowtion.veda.bundle.load() dequantizes them to bf16 at load time (torch.bmm has no e4m3 path), and scoring upcasts to fp32 internally, so fp8 halves the download and the load, not the resident copy.

The reason this is safe is that top-k selection does not care about the absolute value of a score: it is invariant to a per-row constant and to any monotone per-row rescaling. Only the ordering near the budget boundary can flip. Measured on a sibling checkpoint (the latent_t 102 run at update 200, same architecture and the same export path), over 8 denoising steps × 50 layers at 16:9 / latent_t 102, 104603 tokens, keep 0.1:

storage file recall vs. oracle attention mass kept blocks identical to bf16
bf16 525 MiB 0.6017 0.6153 —
fp8 e4m3 263 MiB 0.6017 0.6153 0.9944

That is: halving the file changes recall by 2e-5 and the kept attention mass by 6e-6. About one selected block in 180 changes, and those blocks sit on the budget boundary where they carry no measurable attention mass. The relative error on the raw scores is 5.2e-3 and does not propagate to the selection.

A bf16 export of the same checkpoint can be produced from the training checkpoint with scripts/export_predictor.py --dtype bfloat16; the released artifact is fp8 because nothing measurable is lost.

Geometries

Twelve plans are packed in the file: aspect ratios 16:9, 9:16, 4:3 and 1:1 at latent_t 37 / 72 / 102 (5.17 / 10.1 / 14.4 s). A plan assigns one tile shape per head, chosen by search; every shape holds exactly 128 tokens.

Aspect latent_t Latent grid (T × H × W) Tile shapes in use (T×H×W)
16:9 37 37 × 24 × 42 8x4x4, 2x8x8, 4x8x4, 8x8x2
16:9 72 72 × 24 × 42 8x4x4, 2x8x8, 4x4x8, 4x8x4, 8x8x2
16:9 102 102 × 24 × 42 8x4x4, 2x8x8, 4x8x4, 8x8x2
9:16 37 37 × 42 × 24 8x4x4, 2x8x8, 4x4x8, 8x2x8
9:16 72 72 × 42 × 24 8x4x4, 2x8x8, 4x8x4, 4x4x8, 8x2x8
9:16 102 102 × 42 × 24 8x4x4, 2x8x8, 4x4x8, 8x2x8
4:3 37 37 × 24 × 32 8x4x4, 2x8x8, 4x4x8, 4x8x4, 8x8x2, 8x2x8
4:3 72 72 × 24 × 32 8x4x4, 2x8x8, 4x4x8, 4x8x4, 8x8x2, 8x2x8
4:3 102 102 × 24 × 32 8x4x4, 2x8x8, 4x4x8, 4x8x4, 8x8x2, 8x2x8
1:1 37 37 × 24 × 24 8x4x4, 2x8x8, 4x4x8, 8x8x2, 8x2x8
1:1 72 72 × 24 × 24 8x4x4, 2x8x8, 4x4x8, 4x8x4, 8x8x2, 8x2x8
1:1 102 102 × 24 × 24 8x4x4, 2x8x8, 4x4x8, 4x8x4, 8x8x2, 8x2x8

Asking for a geometry that is not in the table raises; the predictor is not interpolated across grids.

Quick start

Full instructions, including what to install per GPU architecture, are in AGENTS.md. The short version:

git clone https://github.com/veda-sparse/Miowtion.git
cd Miowtion
pip install -e '.[gpu,encode]'

hf download MiniMaxAI/MiniMax-H3 --local-dir weights/MiniMax-H3
hf download Veda-Sparse/Minimax-H3-T2VA-Veda-8NFE-600Step-Preview \
  --local-dir weights/veda/h3-t2va-8nfe-600

Prompts are encoded once into a sample cache (the text tower is large; the generator does not want to hold it):

echo '{"id": "demo", "task": "t2va", "prompt": "<structured T2VA prompt>"}' \
  > prompts.jsonl
python scripts/encode_samples.py --root weights/MiniMax-H3 \
  --manifest prompts.jsonl --out artifacts/samples/demo

Then generate:

python scripts/generate.py \
  --root weights/MiniMax-H3 --variant FL2VA \
  --schedule turbo --num-steps 8 \
  --adapter weights/turbo_lora/<8-step-lora>.safetensors \
  --sample-cache artifacts/samples/demo --sample-id demo \
  --geometry 16:9@37 --attention veda \
  --predictor weights/veda/h3-t2va-8nfe-600/minimax_h3_t2va_veda_8nfe_600step_preview_fp8.safetensors \
  --out-dir artifacts/generate/demo

--predictor brings its own plans and keep ratio, so --plan-dir and --checkpoint are not needed. Use --attention dense veda to also render the dense reference and a side-by-side video with per-step timings in summary.json; step 0 includes kernel compilation and is excluded from the speedups.

Options worth knowing:

Flag
--geometry <aspect>@<latent_t>, one of the twelve above; repeatable, one per --sample-id
--keep-ratio overrides the bundle's 0.1; the predictor was trained at 0.1
--dense-steps 0 keep the listed denoising steps dense (step 0 is the noisiest and the hardest to select for)
--offload-blocks transformer blocks streamed from host memory; 50 = all of them, which is what a 24 GB card needs at latent_t 102
--mlp-chunk-rows rows per MLP chunk; lower it if the allocator runs out at long geometries
--seed same seed gives the same noise for every attention mode, which is what makes dense and Veda comparable
--decode-dtype video VAE weights, bf16 (default) or fp32

One process drives every visible GPU; pick them with CUDA_VISIBLE_DEVICES. Each sample runs on one GPU, so several GPUs mean several samples at once, not one sample faster.

To load the weights directly:

from miowtion.veda import bundle

loaded = bundle.load(
    'minimax_h3_t2va_veda_8nfe_600step_preview_fp8.safetensors', device='cuda')
loaded.predictor          # TileScorePredictor, eval mode, bf16
loaded.plans.select(geo)  # the plan for a geometry
loaded.keep_ratio         # 0.1

Measured speedups

3×RTX 4090 (SM89, weights offloaded to host memory), FL2VA + 8-step Turbo LoRA, seed 0, keep 0.1, 20 held-out prompts, step 0 excluded:

Geometry Clips Dense Veda End-to-end Attention
latent_t 37 (5.17 s) 5 553 s 345 s 1.57× 4.55×
latent_t 72 (10.1 s) 8 3043 s 1342 s 2.21× 6.14×
latent_t 102 (14.4 s) 7 5187 s 1852 s 2.76× 6.66×
total 20 146.4 min 59.0 min 2.24× 5.92×

The end-to-end gain grows with sequence length because attention is O(n²) and everything else is O(n): attention is 42% of the dense step at 1:1 / latent_t 37 but 71% at 16:9 / latent_t 102, so the same ~6× attention speedup buys very different totals. Best single clip: 16:9 / latent_t 102, 3.08× end-to-end and 6.87× attention.

Scope

Text-to-audio-video only. The predictor replaces attention block selection; it does not change the denoiser, the schedule, or the VAE, and it is orthogonal to the few-step LoRA it runs under. Sparse and dense outputs differ — block-sparse attention is not bit-exact against dense. At a 10% keep ratio the attainable attention mass is bounded: in our latent_t 37 measurements even an oracle mask recovers only about 0.63 of it.

This checkpoint inherits the MiniMax H3 Community License from its base model.

Citation

@inproceedings{han2026veda,
  title={Veda: Scalable Video Diffusion via Distilled Sparse Attention},
  author={Han, Shihao and Yang, Hao and Hu, Xinting and Mei, Xiaofeng
          and Jiang, Yi and Qi, Xiaojuan},
  booktitle={International Conference on Machine Learning (ICML)},
  year={2026}
}
Downloads last month
179
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Veda-Sparse/Minimax-H3-T2VA-Veda-8NFE-600Step-Preview

Finetuned
(146)
this model

Paper for Veda-Sparse/Minimax-H3-T2VA-Veda-8NFE-600Step-Preview