Bernini-v2 (int4, MLX)

4-bit variant of mlx-community/Bernini-v2-bf16 — Apple-MLX conversion of ByteDance/Bernini-Diffusers-v2 (revision 399cf6a, Apache-2.0).

Quantization scope: the two Wan2.2-A14B transformer experts only (attention + FFN Linears, 4-bit, group size 64; embeddings/norms/head/modulation stay bf16). The planner plane (mllm/, vit_decoder, planner_glue), umT5 encoder, and VAE pass through at bf16 — weight-only int4 on diffusion transformers costs measurable quality (bf16 is the recommended default; this variant is the memory-constrained opt-in).

See the bf16 repo's card for the full component table, conversion notes, retrained-experts warning (not interchangeable with Bernini-R-*), usage pointers, and provenance.

Downloads last month
10
MLX
Hardware compatibility
Log In to add your hardware

Quantized

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for mlx-community/Bernini-v2-int4

Finetuned
(2)
this model

Collection including mlx-community/Bernini-v2-int4