YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

HunyuanImage-3.0-Instruct-Distil · MXFP8 (auto_round)

本目录是 tencent/HunyuanImage-3.0-Instruct-Distil 的 MXFP8 量化产物, 用 AutoRound 的 **--format auto_round**(vLLM INC 路径)导出。 在 vLLM + vLLM-Omni 上,AR(自回归语言模型)与 DiT(扩散 Transformer)两个 stage 都已实测跑通。

AR+DiT 全流程实测

上图为 seed=42、8 步、prompt A cute cat 的 AR+DiT 全流程输出(AR 先产出 CoT + 比例 token,KV 复用给 DiT)。

概览

项 值
基座模型 tencent/HunyuanImage-3.0-Instruct-Distil(Distil 版,cfg_distilled=true / use_meanflow=true)
量化方案 AutoRound MXFP8(--scheme MXFP8),分块 32 元素、8 bit、对称、共享指数(E8M0 scale)
导出格式 --format auto_round → quantization_config.quant_method = "auto-round"、data_type = "mx_fp"(vLLM INC 路径)
量化工具 auto-round 0.15.0(--model_free,不需要校准数据)
磁盘占用 86 GB(基座 BF16 为 158 GB,约 0.54×)
权重 dtype 打包为 BF16 存储的 FP8-E4M3 值 + uint8 MX 分块 scale
已量化 DiT/AR 的 Linear(含 MoE 专家 gate_and_up_proj / down_proj),作用域 block_name_to_quantize = "model.layers"
未量化(保持 BF16) vision(ViT)、wte/lm_head、guidance_emb、timestep_emb、timestep_r_emb、final_layer,以及 extra_config 里显式标了 bits=16 的层 —— 含每层 MoE 的 router(…mlp.gate.wg,共 32 层)

quantization_config 关键字段(见 config.json):

{
  "quant_method": "auto-round",
  "data_type": "mx_fp",
  "bits": 8,
  "group_size": 32,
  "sym": true,
  "packing_format": "auto_round:llm_compressor",
  "block_name_to_quantize": "model.layers",
  "extra_config": { ...220 条,其中 32 条是 "model.layers.N.mlp.gate.wg": {"bits": 16, ...} ... }
}

与姊妹产物 …-MXFP8-ct 的关系

两者的 on-disk 权重逐字节相同,只有 quantization_config 的元数据写法不同。 (已实测核对:两个目录的 32 个 model-*.safetensors 文件名、大小一致,且 sha256 全部相同。)

元数据差异如下:

本产物(auto_round) …-MXFP8-ct(llm_compressor)
quant_method auto-round(vLLM INC 路径) compressed-tensors
作用域/忽略的表达 block_name_to_quantize + extra_config ignore(301 条)
加载器 vLLM INCConfig → INCMxfp8MoEMethod 等 vLLM CompressedTensorsConfig
实测(同环境、同配置) DiT-only PSNR vs BF16 31.17 dB DiT-only PSNR vs BF16 31.59 dB

选哪个? 两者都能跑通、精度同级。llm_compressor(compressed-tensors)是 vLLM 的一等公民格式, 元数据语义更标准;auto_round 这条走的 INC 路径在 vLLM 上需要 ≥ 0.29.0(0.29 才新增 INCMxfp8MoEMethod, 更早版本会在 MoE 上报缺 w13_weight_scale)。若不确定,优先用 …-MXFP8-ct。

量化命令

# 需要:auto-round >= 0.15.0(本产物用 0.15.0),以及一份基座 BF16 权重
auto-round \
  --model_name tencent/HunyuanImage-3.0-Instruct-Distil \
  --model_free \
  --scheme MXFP8 \
  --ignore_layers "vision,guidance_emb,timestep_emb,timestep_r_emb,final_layer,wte" \
  --format auto_round \
  --device cuda:0 \
  --output_dir ./HunyuanImage-3.0-Instruct-Distil-MXFP8
  • 与 llm_compressor 版唯一差别就是 --format:权重完全一样,只是元数据换一种写法。
  • **--ignore_layers 必须排除 vision**:ViT 的 mlp.fc2 输入维度为 4304,不是 32 的整数倍, MXFP8 分块量化无法表示(否则报 requires input_size_per_partition (2152) to be divisible by 32)。
  • MoE 的 router 由 AutoRound 自动写进 extra_config("model.layers.N.mlp.gate.wg": {"bits": 16})。

推理环境

组件 版本
vLLM 0.29.0(auto_round/INC 路径需要 ≥ 0.29.0,见上)
vLLM-Omni main 最新版(本产物验证于 1c7476ec,版本串 0.29.0rc2.dev161+g1c7476ec1,editable 安装)
PyTorch 2.13.0+cu132
FlashInfer 0.6.18
GPU NVIDIA H200(141 GB/卡)× 2 或 × 4

⚠️ 需要 vLLM-Omni 包含 HunyuanImage-3.0 MXFP8 加载相关修复的版本(修复正在向上游提 PR)。 更早的 vLLM-Omni 版本下,extra_config 里的高精度标记不会按运行时模块名重写, 会导致每层 MoE router 被误量化、weight_scale 未初始化,输出退化为接近纯色(PSNR ≈ 12 dB)。

# 安装(与本产物验证时一致)
pip install vllm==0.29.0
pip install -e /path/to/vllm-omni        # main 分支

# 自检:确认加载的是你期望的那份 vllm-omni
python -c "import vllm, vllm_omni, os; print(vllm.__version__); print(os.path.dirname(vllm_omni.__file__))"

环境变量 / 运行目录

# 本机(H200 + 本仓库验证环境)实测需要的两个开关;如果你的机器没有对应问题可省略:
export NCCL_NVLS_ENABLE=0              # 本机 NVLS fabric 不可用,TP>=2 建组会 NCCL error
export VLLM_USE_FLASHINFER_SAMPLER=0   # 本机 CUDA 头文件与 FlashInfer 采样器 JIT 不匹配

# vLLM-Omni 会以子进程 import 模型代码;cwd 必须是中立目录,
# 否则子进程可能 import 到你本地的 vllm-omni 源码副本
cd /tmp

不需要设置 VLLM_ALLREDUCE_USE_FLASHINFER:vLLM ≥ 0.29 默认开启 FlashInfer all-reduce, vLLM-Omni main 已修好 Diffusion worker 与 vLLM 分布式全局状态的对齐(上游 PR #7676)。

推理示例:AR + DiT 全流程(text2img)

走 vLLM-Omni main 的官方离线入口 examples/offline_inference/text_to_image/text_to_image.py, 配一份两 stage 的 deploy YAML(AR = stage 0,DiT = stage 1,KV 通过共享内存传给 DiT)。

0. 准备(把下面这块存为 hunyuan_image_3_moe.yaml)

# AR (stage 0) + DiT (stage 1),4 张卡:AR 用 2 卡、DiT 用 2 卡
pipeline: hunyuan_image_3_moe
async_chunk: false
trust_remote_code: true

connectors:
  shared_memory_connector:
    name: SharedMemoryConnector

stages:
  - stage_id: 0
    is_comprehension: true
    final_output: true
    final_output_type: text
    max_num_seqs: 1
    gpu_memory_utilization: 0.9
    enforce_eager: true
    max_num_batched_tokens: 32768
    devices: "0,1"
    tensor_parallel_size: 2
    hf_overrides:
      rope_parameters:
        mrope_section: [0, 32, 32]
        rope_type: default
    omni_kv_config:
      need_send_cache: true
    output_connectors:
      to_stage_1: shared_memory_connector
    default_sampling_params:
      temperature: 0.0
      top_p: 1
      top_k: -1
      max_tokens: 8192
      detokenize: true
      skip_special_tokens: false
      include_stop_str_in_output: true

  - stage_id: 1
    max_num_seqs: 1
    gpu_memory_utilization: 0.9
    enforce_eager: true
    devices: "2,3"
    distributed_executor_backend: "mp"
    omni_kv_config:
      need_recv_cache: true
    parallel_config:
      tensor_parallel_size: 2
      enable_expert_parallel: true
    input_connectors:
      from_stage_0: shared_memory_connector
    default_sampling_params:
      num_inference_steps: 8
      guidance_scale: 0

edges:
  - from: 0
    to: 1
    window_size: -1
    max_inflight: 1
  • devices 用本地序号(配合 CUDA_VISIBLE_DEVICES 使用)。
  • 只有 2 张卡时,把两个 stage 的 tensor_parallel_size 改成 1、devices 改成 "0" / "1" (TP=1 时每个 stage 要把整份权重放单卡,实测 ~85 GiB;141 GB 单卡装得下,见「显存占用」)。
  • 用 shared_memory_connector(单机)替代官方 YAML 里的 RDMA/Mooncake 传输,免掉外部依赖。
  • DiT stage 的 guidance_scale: 0 只是兜底,实际用命令行传入的值(见下)。

显存占用(实测,H200 141 GB/卡)

来自 AR+DiT 全流程日志(Model loading took …):

stage TP=2 时每卡 进程总占用/卡
stage 0 = AR 42.7 GiB ~48 GiB
stage 1 = DiT 43.4 GiB ~45 GiB

整个 checkpoint 是 84.98 GiB(vLLM 自己报的数,du -sh 显示 86 GB)。 所以:

  • 4 卡(AR TP2 + DiT TP2):每卡 ~43–49 GiB,余量充足,还可以加大 gpu_memory_utilization。
  • 2 卡(AR TP1 + DiT TP1):每卡要放整份 stage 权重,约 ~85 GiB,141 GB 单卡装得下(已实测)。
  • BF16 基座做不到 2 卡:它的 AR 在 TP2 下就已经把卡吃满,所以基座至少 4 卡(AR TP2 + DiT TP2)。

1. 运行

export CUDA_VISIBLE_DEVICES=0,1,2,3      # 四张空闲卡

cd /tmp                                   # cwd 必须是中立目录
python /path/to/vllm-omni/examples/offline_inference/text_to_image/text_to_image.py \
  --model            /path/to/HunyuanImage-3.0-Instruct-Distil-MXFP8 \
  --deploy-config    ./hunyuan_image_3_moe.yaml \
  --prompt           "A cute cat" \
  --num-inference-steps 8 \
  --guidance-scale   5.0 \
  --output           ./output.png

成功时日志末尾会有 Saved generated image to ./output.png,并且会看到 Using 'MARLIN' MxFp8 MoE backend。

2. 参数说明(重要)

  • **--num-inference-steps 8**:Distil 模型是 8 步蒸馏的,步数不要按 Instruct 版的 50 步来设。
  • **--guidance-scale 对这个模型是"实参"**:产物 config.json 里 cfg_distilled=true, vLLM-Omni 会把 1000 × guidance_scale 作为 guidance embedding 喂进 DiT —— 这个值会实质影响出图。 本目录的验证图用的是 5.0;如果你要和 vLLM-Omni 自带的参照图 (tests/assets/hunyuan_image3/hunyuan_image_distill_ref.png)比 PSNR, 用 2.5(对应 tests/e2e/accuracy/test_hunyuan_image3.py 里 Distil 分支的设置)。
  • --prompt 只走 AR 阶段;AR 会先生成 CoT 与图片比例 token,再连同 KV 缓存交给 DiT。

实测效果

同机、同 prompt(A cute cat)、seed=42、8 步、guidance-scale 5.0、AR 与 DiT 均为 TP=2, 与 BF16 基座 的 AR+DiT 全流程对比:

输出
本产物(MXFP8) mxfp8
BF16 基座(参照) bf16
比较口径 PSNR vs BF16 备注
DiT-only(TP2) 31.17 dB 该口径稳定,这个数是可比的量化误差指标;日志有 Using 'MARLIN' MxFp8 MoE backend,0 条未初始化告警
AR+DiT 全流程(TP2+TP2) 30.7 dB 见上图;该口径抖动大,不能当精度指标,原因见下

与姊妹产物 …-MXFP8-ct 在同一口径下互相 PSNR 28.3 dB —— 两者属于同一水平, 26.2 / 30.7 的先后在这个口径下没有意义(见下)。

复现性说明(重要,请先读)

同一套管线在不同"比较口径"下的可重复性差别很大,实测(同模型、同命令、同 TP 配置重复跑):

比较口径 同模型两次运行的 PSNR 结论
DiT-only 41.5 dB / 39.2 dB 稳定 → 量化误差(~31 dB)是可测且一致的
AR+DiT 全流程 27.8 dB 抖动大 → AR 阶段会重新生成 CoT 与比例 token,量化误差被 run-to-run 抖动淹没

因此:

  • 判断量化精度请用 DiT-only 口径(固定 seed、同 TP、同 prompt)。
  • AR+DiT 的图只看"观感是否正常",不要拿它的 PSNR 当精度指标 —— 上表里 ct 26.2 dB / ar 30.7 dB 之间的差别在这个口径下不具统计意义。
  • 本管线(vLLM-Omni 的 DiT 采样)不是严格确定性的,比较时务必同一次会话内跑 BF16 与量化版; 需要更稳的指标可改用 LPIPS / CLIP 之类的感知度量。

另外:TP 配置也会改变出图

同一产物在 TP1+TP1 与 TP2+TP2 下跑出来的图互相只有 26.0 dB。 比较量化误差时,BF16 与量化版必须用相同的 TP 配置。

已知限制

  1. auto_round(INC)路径需要 vLLM ≥ 0.29.0:0.29.0 才新增 INCMxfp8MoEMethod; 更早的 vLLM 会在 MoE 上报缺 w13_weight_scale / 显存不足。
  2. ViT 保持 BF16:ViT 的 mlp.fc2 输入维度 4304 不能被 32 整除,无法做 MXFP8 分块量化。 所以本产物并不是"全模型 MXFP8"。
  3. AR 与 DiT 共用同一个量化产物,但两者由不同的推理栈加载 (AR 走 vLLM 的普通 LLM 加载器,DiT 走 vLLM-Omni 的 diffusion pipeline)。
  4. 需要较新的 vLLM-Omni:见上文"推理环境"的告警。
  5. 本目录的 images/ 是验证快照,不参与权重加载。

目录内容

config.json                    # 含 quantization_config(auto-round / mx_fp / bits=8)
model-0000N-of-00032.safetensors
model.safetensors.index.json
*.py / tokenizer* / assets/    # 来自基座模型的配置与自定义代码
images/                        # 本 README 用到的验证图
Downloads last month
21
Safetensors
Model size
83B params
Tensor type
BF16
·
F8_E4M3
·
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support