YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
HunyuanImage-3.0-Instruct-Distil · MXFP8 (auto_round)
本目录是 tencent/HunyuanImage-3.0-Instruct-Distil 的 MXFP8 量化产物,
用 AutoRound 的 **--format auto_round**(vLLM INC 路径)导出。
在 vLLM + vLLM-Omni 上,AR(自回归语言模型)与 DiT(扩散 Transformer)两个 stage 都已实测跑通。
上图为
seed=42、8 步、promptA cute cat的 AR+DiT 全流程输出(AR 先产出 CoT + 比例 token,KV 复用给 DiT)。
概览
| 项 | 值 |
|---|---|
| 基座模型 | tencent/HunyuanImage-3.0-Instruct-Distil(Distil 版,cfg_distilled=true / use_meanflow=true) |
| 量化方案 | AutoRound MXFP8(--scheme MXFP8),分块 32 元素、8 bit、对称、共享指数(E8M0 scale) |
| 导出格式 | --format auto_round → quantization_config.quant_method = "auto-round"、data_type = "mx_fp"(vLLM INC 路径) |
| 量化工具 | auto-round 0.15.0(--model_free,不需要校准数据) |
| 磁盘占用 | 86 GB(基座 BF16 为 158 GB,约 0.54×) |
| 权重 dtype | 打包为 BF16 存储的 FP8-E4M3 值 + uint8 MX 分块 scale |
| 已量化 | DiT/AR 的 Linear(含 MoE 专家 gate_and_up_proj / down_proj),作用域 block_name_to_quantize = "model.layers" |
| 未量化(保持 BF16) | vision(ViT)、wte/lm_head、guidance_emb、timestep_emb、timestep_r_emb、final_layer,以及 extra_config 里显式标了 bits=16 的层 —— 含每层 MoE 的 router(…mlp.gate.wg,共 32 层) |
quantization_config 关键字段(见 config.json):
{
"quant_method": "auto-round",
"data_type": "mx_fp",
"bits": 8,
"group_size": 32,
"sym": true,
"packing_format": "auto_round:llm_compressor",
"block_name_to_quantize": "model.layers",
"extra_config": { ...220 条,其中 32 条是 "model.layers.N.mlp.gate.wg": {"bits": 16, ...} ... }
}
与姊妹产物 …-MXFP8-ct 的关系
两者的 on-disk 权重逐字节相同,只有 quantization_config 的元数据写法不同。
(已实测核对:两个目录的 32 个 model-*.safetensors 文件名、大小一致,且 sha256 全部相同。)
元数据差异如下:
本产物(auto_round) |
…-MXFP8-ct(llm_compressor) |
|
|---|---|---|
quant_method |
auto-round(vLLM INC 路径) |
compressed-tensors |
| 作用域/忽略的表达 | block_name_to_quantize + extra_config |
ignore(301 条) |
| 加载器 | vLLM INCConfig → INCMxfp8MoEMethod 等 |
vLLM CompressedTensorsConfig |
| 实测(同环境、同配置) | DiT-only PSNR vs BF16 31.17 dB | DiT-only PSNR vs BF16 31.59 dB |
选哪个? 两者都能跑通、精度同级。llm_compressor(compressed-tensors)是 vLLM 的一等公民格式,
元数据语义更标准;auto_round 这条走的 INC 路径在 vLLM 上需要 ≥ 0.29.0(0.29 才新增 INCMxfp8MoEMethod,
更早版本会在 MoE 上报缺 w13_weight_scale)。若不确定,优先用 …-MXFP8-ct。
量化命令
# 需要:auto-round >= 0.15.0(本产物用 0.15.0),以及一份基座 BF16 权重
auto-round \
--model_name tencent/HunyuanImage-3.0-Instruct-Distil \
--model_free \
--scheme MXFP8 \
--ignore_layers "vision,guidance_emb,timestep_emb,timestep_r_emb,final_layer,wte" \
--format auto_round \
--device cuda:0 \
--output_dir ./HunyuanImage-3.0-Instruct-Distil-MXFP8
- 与
llm_compressor版唯一差别就是--format:权重完全一样,只是元数据换一种写法。 - **
--ignore_layers必须排除vision**:ViT 的mlp.fc2输入维度为 4304,不是 32 的整数倍, MXFP8 分块量化无法表示(否则报requires input_size_per_partition (2152) to be divisible by 32)。 - MoE 的 router 由 AutoRound 自动写进
extra_config("model.layers.N.mlp.gate.wg": {"bits": 16})。
推理环境
| 组件 | 版本 |
|---|---|
| vLLM | 0.29.0(auto_round/INC 路径需要 ≥ 0.29.0,见上) |
| vLLM-Omni | main 最新版(本产物验证于 1c7476ec,版本串 0.29.0rc2.dev161+g1c7476ec1,editable 安装) |
| PyTorch | 2.13.0+cu132 |
| FlashInfer | 0.6.18 |
| GPU | NVIDIA H200(141 GB/卡)× 2 或 × 4 |
⚠️ 需要 vLLM-Omni 包含 HunyuanImage-3.0 MXFP8 加载相关修复的版本(修复正在向上游提 PR)。 更早的 vLLM-Omni 版本下,
extra_config里的高精度标记不会按运行时模块名重写, 会导致每层 MoE router 被误量化、weight_scale未初始化,输出退化为接近纯色(PSNR ≈ 12 dB)。
# 安装(与本产物验证时一致)
pip install vllm==0.29.0
pip install -e /path/to/vllm-omni # main 分支
# 自检:确认加载的是你期望的那份 vllm-omni
python -c "import vllm, vllm_omni, os; print(vllm.__version__); print(os.path.dirname(vllm_omni.__file__))"
环境变量 / 运行目录
# 本机(H200 + 本仓库验证环境)实测需要的两个开关;如果你的机器没有对应问题可省略:
export NCCL_NVLS_ENABLE=0 # 本机 NVLS fabric 不可用,TP>=2 建组会 NCCL error
export VLLM_USE_FLASHINFER_SAMPLER=0 # 本机 CUDA 头文件与 FlashInfer 采样器 JIT 不匹配
# vLLM-Omni 会以子进程 import 模型代码;cwd 必须是中立目录,
# 否则子进程可能 import 到你本地的 vllm-omni 源码副本
cd /tmp
不需要设置 VLLM_ALLREDUCE_USE_FLASHINFER:vLLM ≥ 0.29 默认开启 FlashInfer all-reduce,
vLLM-Omni main 已修好 Diffusion worker 与 vLLM 分布式全局状态的对齐(上游 PR #7676)。
推理示例:AR + DiT 全流程(text2img)
走 vLLM-Omni main 的官方离线入口 examples/offline_inference/text_to_image/text_to_image.py,
配一份两 stage 的 deploy YAML(AR = stage 0,DiT = stage 1,KV 通过共享内存传给 DiT)。
0. 准备(把下面这块存为 hunyuan_image_3_moe.yaml)
# AR (stage 0) + DiT (stage 1),4 张卡:AR 用 2 卡、DiT 用 2 卡
pipeline: hunyuan_image_3_moe
async_chunk: false
trust_remote_code: true
connectors:
shared_memory_connector:
name: SharedMemoryConnector
stages:
- stage_id: 0
is_comprehension: true
final_output: true
final_output_type: text
max_num_seqs: 1
gpu_memory_utilization: 0.9
enforce_eager: true
max_num_batched_tokens: 32768
devices: "0,1"
tensor_parallel_size: 2
hf_overrides:
rope_parameters:
mrope_section: [0, 32, 32]
rope_type: default
omni_kv_config:
need_send_cache: true
output_connectors:
to_stage_1: shared_memory_connector
default_sampling_params:
temperature: 0.0
top_p: 1
top_k: -1
max_tokens: 8192
detokenize: true
skip_special_tokens: false
include_stop_str_in_output: true
- stage_id: 1
max_num_seqs: 1
gpu_memory_utilization: 0.9
enforce_eager: true
devices: "2,3"
distributed_executor_backend: "mp"
omni_kv_config:
need_recv_cache: true
parallel_config:
tensor_parallel_size: 2
enable_expert_parallel: true
input_connectors:
from_stage_0: shared_memory_connector
default_sampling_params:
num_inference_steps: 8
guidance_scale: 0
edges:
- from: 0
to: 1
window_size: -1
max_inflight: 1
devices用本地序号(配合CUDA_VISIBLE_DEVICES使用)。- 只有 2 张卡时,把两个 stage 的
tensor_parallel_size改成1、devices改成"0"/"1"(TP=1 时每个 stage 要把整份权重放单卡,实测 ~85 GiB;141 GB 单卡装得下,见「显存占用」)。- 用
shared_memory_connector(单机)替代官方 YAML 里的 RDMA/Mooncake 传输,免掉外部依赖。- DiT stage 的
guidance_scale: 0只是兜底,实际用命令行传入的值(见下)。
显存占用(实测,H200 141 GB/卡)
来自 AR+DiT 全流程日志(Model loading took …):
| stage | TP=2 时每卡 | 进程总占用/卡 |
|---|---|---|
| stage 0 = AR | 42.7 GiB | ~48 GiB |
| stage 1 = DiT | 43.4 GiB | ~45 GiB |
整个 checkpoint 是 84.98 GiB(vLLM 自己报的数,du -sh 显示 86 GB)。
所以:
- 4 卡(AR TP2 + DiT TP2):每卡 ~43–49 GiB,余量充足,还可以加大
gpu_memory_utilization。 - 2 卡(AR TP1 + DiT TP1):每卡要放整份 stage 权重,约 ~85 GiB,141 GB 单卡装得下(已实测)。
- BF16 基座做不到 2 卡:它的 AR 在 TP2 下就已经把卡吃满,所以基座至少 4 卡(AR TP2 + DiT TP2)。
1. 运行
export CUDA_VISIBLE_DEVICES=0,1,2,3 # 四张空闲卡
cd /tmp # cwd 必须是中立目录
python /path/to/vllm-omni/examples/offline_inference/text_to_image/text_to_image.py \
--model /path/to/HunyuanImage-3.0-Instruct-Distil-MXFP8 \
--deploy-config ./hunyuan_image_3_moe.yaml \
--prompt "A cute cat" \
--num-inference-steps 8 \
--guidance-scale 5.0 \
--output ./output.png
成功时日志末尾会有 Saved generated image to ./output.png,并且会看到
Using 'MARLIN' MxFp8 MoE backend。
2. 参数说明(重要)
- **
--num-inference-steps 8**:Distil 模型是 8 步蒸馏的,步数不要按 Instruct 版的 50 步来设。 - **
--guidance-scale对这个模型是"实参"**:产物config.json里cfg_distilled=true, vLLM-Omni 会把1000 × guidance_scale作为 guidance embedding 喂进 DiT —— 这个值会实质影响出图。 本目录的验证图用的是5.0;如果你要和 vLLM-Omni 自带的参照图 (tests/assets/hunyuan_image3/hunyuan_image_distill_ref.png)比 PSNR, 用 2.5(对应tests/e2e/accuracy/test_hunyuan_image3.py里 Distil 分支的设置)。 --prompt只走 AR 阶段;AR 会先生成 CoT 与图片比例 token,再连同 KV 缓存交给 DiT。
实测效果
同机、同 prompt(A cute cat)、seed=42、8 步、guidance-scale 5.0、AR 与 DiT 均为 TP=2,
与 BF16 基座 的 AR+DiT 全流程对比:
| 比较口径 | PSNR vs BF16 | 备注 |
|---|---|---|
| DiT-only(TP2) | 31.17 dB | 该口径稳定,这个数是可比的量化误差指标;日志有 Using 'MARLIN' MxFp8 MoE backend,0 条未初始化告警 |
| AR+DiT 全流程(TP2+TP2) | 30.7 dB | 见上图;该口径抖动大,不能当精度指标,原因见下 |
与姊妹产物
…-MXFP8-ct在同一口径下互相 PSNR 28.3 dB —— 两者属于同一水平, 26.2 / 30.7 的先后在这个口径下没有意义(见下)。
复现性说明(重要,请先读)
同一套管线在不同"比较口径"下的可重复性差别很大,实测(同模型、同命令、同 TP 配置重复跑):
| 比较口径 | 同模型两次运行的 PSNR | 结论 |
|---|---|---|
| DiT-only | 41.5 dB / 39.2 dB | 稳定 → 量化误差(~31 dB)是可测且一致的 |
| AR+DiT 全流程 | 27.8 dB | 抖动大 → AR 阶段会重新生成 CoT 与比例 token,量化误差被 run-to-run 抖动淹没 |
因此:
- 判断量化精度请用 DiT-only 口径(固定 seed、同 TP、同 prompt)。
- AR+DiT 的图只看"观感是否正常",不要拿它的 PSNR 当精度指标 —— 上表里 ct 26.2 dB / ar 30.7 dB 之间的差别在这个口径下不具统计意义。
- 本管线(vLLM-Omni 的 DiT 采样)不是严格确定性的,比较时务必同一次会话内跑 BF16 与量化版; 需要更稳的指标可改用 LPIPS / CLIP 之类的感知度量。
另外:TP 配置也会改变出图
同一产物在 TP1+TP1 与 TP2+TP2 下跑出来的图互相只有 26.0 dB。 比较量化误差时,BF16 与量化版必须用相同的 TP 配置。
已知限制
auto_round(INC)路径需要 vLLM ≥ 0.29.0:0.29.0 才新增INCMxfp8MoEMethod; 更早的 vLLM 会在 MoE 上报缺w13_weight_scale/ 显存不足。- ViT 保持 BF16:ViT 的
mlp.fc2输入维度 4304 不能被 32 整除,无法做 MXFP8 分块量化。 所以本产物并不是"全模型 MXFP8"。 - AR 与 DiT 共用同一个量化产物,但两者由不同的推理栈加载 (AR 走 vLLM 的普通 LLM 加载器,DiT 走 vLLM-Omni 的 diffusion pipeline)。
- 需要较新的 vLLM-Omni:见上文"推理环境"的告警。
- 本目录的
images/是验证快照,不参与权重加载。
目录内容
config.json # 含 quantization_config(auto-round / mx_fp / bits=8)
model-0000N-of-00032.safetensors
model.safetensors.index.json
*.py / tokenizer* / assets/ # 来自基座模型的配置与自定义代码
images/ # 本 README 用到的验证图
- Downloads last month
- 21

