Confucius4-R2T2 ONNX
ONNX export of netease-youdao/Confucius4-R2T2 (a Qwen3-ASR-1.7B fine-tune) for onnxruntime / onnxruntime-web (WebGPU). Includes fp16, 4-bit and sub-4-bit (2-bit MLP) variants. The demo defaults to encoder_q4f16 + decoder_q2mixf16 (~1.0 GB).
Browser demo: mrfakename/confucius4-r2t2-webgpu
Any modifications made to the original model in this Derivative Work are not endorsed, warranted, or guaranteed by the original right-holder of the original model, and the original right-holder disclaims all liability related to this Derivative Work.
The model weights are distributed under the NetEase Youdao Model Use License (see MODEL_LICENSE). The export / inference code is Apache-2.0.
Files
| file | what |
|---|---|
encoder_fp16.onnx |
audio encoder, fp16 weights, fp32 I/O |
encoder_q8.onnx |
encoder, 8-bit MatMulNBits (block 32), fp32 |
encoder_q4.onnx / encoder_q4f16.onnx |
encoder, 4-bit GPTQ (block 32), fp32 / fp16 |
decoder_fp16.onnx |
Qwen3 text decoder, fp16, GQA with fused q/k-norm (WebGPU / CUDA EPs) |
decoder_q4f16.onnx |
decoder, all linears 4-bit GPTQ, fp16 (WebGPU / CUDA) |
decoder_q2mixf16.onnx |
gate/up 2-bit, q/k/v/o/down 4-bit, fp16 (WebGPU / CUDA) |
decoder_q2mlpf16.onnx |
gate/up/down 2-bit, attention 4-bit, fp16 (WebGPU / CUDA) |
decoder_q4.onnx, decoder_q2mix.onnx, decoder_q2mlp.onnx |
same weights, fp32 graph with unfused q/k-norm (CPU EP) |
decoder_q2mix_wasm.onnx |
decoder_q2mix.onnx with the GatherBlockQuantized embedding rewritten as Gather + nibble unpack (onnxruntime-web WASM has no GBQ kernel); reads decoder_q2mix.onnx_data |
All decoders quantize the tied embedding / lm_head to symmetric 4-bit (GatherBlockQuantized + MatMulNBits share one weight). Large graphs use external data (*.onnx_data, *.onnx_data_1, ...).
Interface
Features: Whisper log-mel, 128 bins, 16 kHz, n_fft 400, hop 160, no 30 s padding (WhisperFeatureExtractor(..., padding="longest", truncation=False)). mel_filters_201x128.f32 holds the filterbank.
Encoder: input_features float32 [n_chunks, 128, 100] โ audio_embeds [n_tokens, 2048]. Split the mel into windows of 800 frames; split each window into 100-frame chunks, zero-padding the last one. Each full chunk produces 13 tokens; a tail chunk of L frames produces ((((L-1)//2+1)-1)//2+1-1)//2+1 tokens. Run windows independently and concatenate them.
Decoder: inputs input_ids int64 [1, S], audio_embeds float32 [1, S, 2048], attention_mask int64 [1, past+S], past_key_values.{0..27}.{key,value} [1, 8, past, 128] (fp16 for *f16 / fp16 graphs). At every position where input_ids == 151676 (<|audio_pad|>), the embedding is replaced by the corresponding row of audio_embeds. On decode steps, pass zeros [1, 1, 2048]. Outputs: logits and present.*.
Prompt: <|im_start|>system\n{context}<|im_end|>\n<|im_start|>user\n<|audio_start|> + <|audio_pad|> ร n_tokens + <|audio_end|><|im_end|>\n<|im_start|>assistant\n. Optionally append language English<asr_text> to force the language. Decode greedily until <|im_end|>; the output looks like language Chinese<asr_text>....
Quality
The ONNX files were checked against PyTorch on 12 LibriSpeech (EN, WER) and 9 WenetSpeech test_net (ZH, CER) clips. vs ref is the error against the human reference; vs pt is the error against the bf16 PyTorch transcript. The scores below come from a fake-quant run in PyTorch using exactly the same integer weights as the ONNX files. ORT end-to-end checks are in eval_results.json.
Note: the 2-bit *f16 graphs run correctly on onnxruntime-web WebGPU (1.30), which is what the demo uses, but onnxruntime-gpu 1.30's CUDA MatMulNBits kernel fails on them (illegal memory access / garbage output). On CUDA, use decoder_q4f16.onnx or decoder_fp16.onnx.
| variant | EN WER vs ref | EN WER vs pt | ZH CER vs ref | ZH CER vs pt |
|---|---|---|---|---|
| pt_bf16 | 5.87 | - | 11.58 | - |
| fp32 | 5.87 | 0.0 | 10.53 | 1.1 |
| dec-q4 | 5.62 | 0.25 | 8.42 | 3.3 |
| dec-q2mix | 5.62 | 1.98 | 9.47 | 6.59 |
| dec-q2mlp | 7.09 | 2.72 | 9.47 | 8.79 |
| dec-q2 | 100.0 | 100.0 | 86.32 | 85.71 |
| enc-q8 | 5.87 | 0.0 | 10.53 | 1.1 |
| enc-q4 | 5.87 | 0.25 | 10.53 | 4.4 |
| dec-q4+enc-q4 | 5.87 | 0.0 | 9.47 | 5.49 |
| dec-q2mix+enc-q4 | 5.87 | 2.22 | 10.53 | 7.69 |
| dec-q2mix+enc-q8 | 5.62 | 1.98 | 9.47 | 6.59 |
Quantization: decoder linears use GPTQ (asymmetric, block 32) with Hessians from ~100 EN/ZH calibration clips (teacher-forced on the model's own transcripts). Encoder q8 uses RTN with an MSE-optimal clip; encoder q4 uses GPTQ.
- Downloads last month
- 20