Confucius4-R2T2 ONNX

ONNX export of netease-youdao/Confucius4-R2T2 (a Qwen3-ASR-1.7B fine-tune) for onnxruntime / onnxruntime-web (WebGPU). Includes fp16, 4-bit and sub-4-bit (2-bit MLP) variants. The demo defaults to encoder_q4f16 + decoder_q2mixf16 (~1.0 GB).

Browser demo: mrfakename/confucius4-r2t2-webgpu

Any modifications made to the original model in this Derivative Work are not endorsed, warranted, or guaranteed by the original right-holder of the original model, and the original right-holder disclaims all liability related to this Derivative Work.

The model weights are distributed under the NetEase Youdao Model Use License (see MODEL_LICENSE). The export / inference code is Apache-2.0.

Files

file what
encoder_fp16.onnx audio encoder, fp16 weights, fp32 I/O
encoder_q8.onnx encoder, 8-bit MatMulNBits (block 32), fp32
encoder_q4.onnx / encoder_q4f16.onnx encoder, 4-bit GPTQ (block 32), fp32 / fp16
decoder_fp16.onnx Qwen3 text decoder, fp16, GQA with fused q/k-norm (WebGPU / CUDA EPs)
decoder_q4f16.onnx decoder, all linears 4-bit GPTQ, fp16 (WebGPU / CUDA)
decoder_q2mixf16.onnx gate/up 2-bit, q/k/v/o/down 4-bit, fp16 (WebGPU / CUDA)
decoder_q2mlpf16.onnx gate/up/down 2-bit, attention 4-bit, fp16 (WebGPU / CUDA)
decoder_q4.onnx, decoder_q2mix.onnx, decoder_q2mlp.onnx same weights, fp32 graph with unfused q/k-norm (CPU EP)
decoder_q2mix_wasm.onnx decoder_q2mix.onnx with the GatherBlockQuantized embedding rewritten as Gather + nibble unpack (onnxruntime-web WASM has no GBQ kernel); reads decoder_q2mix.onnx_data

All decoders quantize the tied embedding / lm_head to symmetric 4-bit (GatherBlockQuantized + MatMulNBits share one weight). Large graphs use external data (*.onnx_data, *.onnx_data_1, ...).

Interface

Features: Whisper log-mel, 128 bins, 16 kHz, n_fft 400, hop 160, no 30 s padding (WhisperFeatureExtractor(..., padding="longest", truncation=False)). mel_filters_201x128.f32 holds the filterbank.

Encoder: input_features float32 [n_chunks, 128, 100] โ†’ audio_embeds [n_tokens, 2048]. Split the mel into windows of 800 frames; split each window into 100-frame chunks, zero-padding the last one. Each full chunk produces 13 tokens; a tail chunk of L frames produces ((((L-1)//2+1)-1)//2+1-1)//2+1 tokens. Run windows independently and concatenate them.

Decoder: inputs input_ids int64 [1, S], audio_embeds float32 [1, S, 2048], attention_mask int64 [1, past+S], past_key_values.{0..27}.{key,value} [1, 8, past, 128] (fp16 for *f16 / fp16 graphs). At every position where input_ids == 151676 (<|audio_pad|>), the embedding is replaced by the corresponding row of audio_embeds. On decode steps, pass zeros [1, 1, 2048]. Outputs: logits and present.*.

Prompt: <|im_start|>system\n{context}<|im_end|>\n<|im_start|>user\n<|audio_start|> + <|audio_pad|> ร— n_tokens + <|audio_end|><|im_end|>\n<|im_start|>assistant\n. Optionally append language English<asr_text> to force the language. Decode greedily until <|im_end|>; the output looks like language Chinese<asr_text>....

Quality

The ONNX files were checked against PyTorch on 12 LibriSpeech (EN, WER) and 9 WenetSpeech test_net (ZH, CER) clips. vs ref is the error against the human reference; vs pt is the error against the bf16 PyTorch transcript. The scores below come from a fake-quant run in PyTorch using exactly the same integer weights as the ONNX files. ORT end-to-end checks are in eval_results.json.

Note: the 2-bit *f16 graphs run correctly on onnxruntime-web WebGPU (1.30), which is what the demo uses, but onnxruntime-gpu 1.30's CUDA MatMulNBits kernel fails on them (illegal memory access / garbage output). On CUDA, use decoder_q4f16.onnx or decoder_fp16.onnx.

variant EN WER vs ref EN WER vs pt ZH CER vs ref ZH CER vs pt
pt_bf16 5.87 - 11.58 -
fp32 5.87 0.0 10.53 1.1
dec-q4 5.62 0.25 8.42 3.3
dec-q2mix 5.62 1.98 9.47 6.59
dec-q2mlp 7.09 2.72 9.47 8.79
dec-q2 100.0 100.0 86.32 85.71
enc-q8 5.87 0.0 10.53 1.1
enc-q4 5.87 0.25 10.53 4.4
dec-q4+enc-q4 5.87 0.0 9.47 5.49
dec-q2mix+enc-q4 5.87 2.22 10.53 7.69
dec-q2mix+enc-q8 5.62 1.98 9.47 6.59

Quantization: decoder linears use GPTQ (asymmetric, block 32) with Hessians from ~100 EN/ZH calibration clips (teacher-forced on the model's own transcripts). Encoder q8 uses RTN with an MSE-optimal clip; encoder q4 uses GPTQ.

Downloads last month
20
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for mrfakename/Confucius4-R2T2-ONNX

Quantized
(11)
this model

Space using mrfakename/Confucius4-R2T2-ONNX 1