Instructions to use ISTA-DASLab/Kimi-K3-P48NVFP4-MoESQ with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use ISTA-DASLab/Kimi-K3-P48NVFP4-MoESQ with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="ISTA-DASLab/Kimi-K3-P48NVFP4-MoESQ", trust_remote_code=True)# pip install -U transformers accelerate # Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("ISTA-DASLab/Kimi-K3-P48NVFP4-MoESQ", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use ISTA-DASLab/Kimi-K3-P48NVFP4-MoESQ with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "ISTA-DASLab/Kimi-K3-P48NVFP4-MoESQ" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ISTA-DASLab/Kimi-K3-P48NVFP4-MoESQ", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/ISTA-DASLab/Kimi-K3-P48NVFP4-MoESQ
- SGLang
How to use ISTA-DASLab/Kimi-K3-P48NVFP4-MoESQ with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "ISTA-DASLab/Kimi-K3-P48NVFP4-MoESQ" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ISTA-DASLab/Kimi-K3-P48NVFP4-MoESQ", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "ISTA-DASLab/Kimi-K3-P48NVFP4-MoESQ" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ISTA-DASLab/Kimi-K3-P48NVFP4-MoESQ", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use ISTA-DASLab/Kimi-K3-P48NVFP4-MoESQ with Docker Model Runner:
docker model run hf.co/ISTA-DASLab/Kimi-K3-P48NVFP4-MoESQ
Kimi-K3-P48NVFP4-MoESQ
A W4A4 + paired-4:8 sparse compressed checkpoint of
moonshotai/Kimi-K3, produced with
MoESQ. The routed-expert weights are NVFP4 with paired-4:8 structured sparsity, and they
are stored sparse: only the kept values plus a small mask are on disk, not a dense NVFP4
tensor with zeros in place. The target is NVIDIA Blackwell (SM100 and SM120) sparse tensor cores.
- Base model: moonshotai/Kimi-K3 (MoE, 896 routed experts in a 3,584-wide latent space, top-16, 93 layers; released in MXFP4)
- Compression: NVFP4 W4A4 plus paired-4:8 sparsity on the routed experts (2.72 T parameters)
- Routed experts: 1,347.1 GiB (MXFP4 source) → 871.7 GiB, 0.65× (about 2.75 bits/weight, scales and mask included; 5.8× smaller than the same experts in BF16)
- Checkpoint size: 978.2 GiB, vs 1,453.7 GiB for the MXFP4 source (0.67×): it serves on a single 8× B200 node (132 GiB of weights per GPU), which the MXFP4 release does not fit
- Kernel: paired-4:8 sparse NVFP4 grouped GEMM (CUTLASS, SM100 and SM120) through vLLM's
paired48_nvfp4MoE backend, with Kimi-K3's SiTU activation fused into the GEMM1 epilogue
Links
- Paper: Hardware-Native Joint Sparse-Quantization for Trillion-Scale Mixture-of-Experts (arXiv:2610.02241)
- Code: IST-DASLab/MoESQ. It contains the compression code, the vLLM integration and the kernels.
- Other MoESQ checkpoints: ISTA-DASLab/moesq collection
Usage
This checkpoint does not load in upstream vLLM. The MoESQ repository installs a patched
vLLM v0.30.0 (the paired48_nvfp4 MoE backend and sparse-storage loader) together with the
kernels; Kimi-K3 needs kernels v0.14.0 or later, which add its SiTU activation. It needs an
NVIDIA Blackwell SM100 (B200, GB200) or SM120 (RTX 5090, RTX PRO 6000) GPU and a CUDA toolkit
= 12.8; SM103 (B300) and SM121 (DGX Spark) are not supported. The SiTU path has been tested on SM100 only so far.
git clone --recurse-submodules https://github.com/IST-DASLab/MoESQ.git && cd MoESQ
bash integrations/vllm/install.sh && source .venv-vllm/bin/activate
# 1 node, 8x B200: tensor + expert parallel. Kimi Delta Attention keeps one state slot per
# running sequence, and only ~530 fit beside the weights, so cap --max-num-seqs.
vllm serve ISTA-DASLab/Kimi-K3-P48NVFP4-MoESQ --trust-remote-code --language-model-only \
--tensor-parallel-size 8 --enable-expert-parallel --max-num-seqs 256 --max-model-len 32768 \
--kv-cache-dtype fp8 --attention-backend TOKENSPEED_MLA \
--attention-config '{"use_prefill_query_quantization":true,"mla_prefill_backend":"TOKENSPEED_MLA"}'
# or 2 nodes x 8 B200 (more KV cache): tensor parallel 8 within a node, pipeline parallel 2 across nodes,
# expert parallel inside each tensor-parallel group. Run on node 0:
vllm serve ISTA-DASLab/Kimi-K3-P48NVFP4-MoESQ --trust-remote-code --language-model-only \
--tensor-parallel-size 8 --pipeline-parallel-size 2 --enable-expert-parallel \
--kv-cache-dtype fp8 --attention-backend TOKENSPEED_MLA \
--attention-config '{"use_prefill_query_quantization":true,"mla_prefill_backend":"TOKENSPEED_MLA"}' \
--nnodes 2 --node-rank 0 --master-addr <node-0 address> --master-port 29581
# and on node 1 the same command with --node-rank 1 --headless
vLLM selects the backend automatically. --language-model-only serves text only; the vision
tower is in the checkpoint (BF16) but was not evaluated. See
Serving with vLLM for other
parallel layouts.
Evaluation
The reference is the original Kimi-K3 release (MXFP4 routed experts), served with the same vLLM build and the same arguments on 2 × 8 B200 (TP8 × PP2, expert parallel).
OpenLLM Leaderboard v1: 6-task average
The six tasks are ARC-Challenge (25-shot), HellaSwag (10-shot), MMLU (5-shot), TruthfulQA-MC2 (0-shot), Winogrande (5-shot) and GSM8K (5-shot, greedy). Scores use lm-evaluation-harness with the full test sets, few-shot seed 0 (one run).
| Task | Kimi-K3 (MXFP4) | This model (MoESQ) |
|---|---|---|
| ARC-Challenge | 74.40 | 72.87 |
| HellaSwag | 91.78 | 89.11 |
| MMLU | 91.16 | 88.68 |
| TruthfulQA-MC2 | 60.18 | 60.78 |
| Winogrande | 83.19 | 83.66 |
| GSM8K | 96.59 | 93.63 |
| Average | 82.88 | 81.45 (98.3 % recovery) |
Reasoning and coding
The metric is average pass@1 over N samples per problem (AIME25 N=10, GPQA-Diamond N=5, MATH-500 N=5; LiveCodeBench v6 code generation over 10 seeds). Sampling uses temperature 1.0 and top_p 0.95, with up to 65,536 new tokens, at the default reasoning effort. Both models were run with this same protocol. ± is the lighteval standard error across problems; for LiveCodeBench it is the standard deviation across the 10 seeds.
| Benchmark | Kimi-K3 (MXFP4) | This model (MoESQ) | Recovery |
|---|---|---|---|
| AIME25 | 97.33 ± 1.17 | 98.33 ± 0.69 | 101.0 % |
| GPQA-Diamond | 92.42 ± 1.66 | 89.39 ± 1.85 | 96.7 % |
| MATH-500 | 89.84 ± 1.11 | 90.00 ± 1.14 | 100.2 % |
| LiveCodeBench v6 | 57.89 ± 1.21 | 62.80 ± 1.95 | 108.5 % |
The LiveCodeBench gain is consistent across problems rather than a few outliers (paired over problems: +4.9 points, t = 4.5; better on 50 of 175 problems, worse on 16, equal on 109), and both runs used identical sampling settings with no truncated or empty answers. We report it as measured and do not claim that compression improves coding in general.
Compression details
| Field | Value |
|---|---|
| Weights | NVFP4 (E2M1), one FP8-E4M3 scale per 32 dense K elements, FP32 per-tensor global scale |
| Activations | NVFP4, dynamic per-32 group scales, FP32 per-tensor global scale stored per expert linear |
| Sparsity | Paired 4:8 along K: in every 8 consecutive elements (4 pairs), exactly 2 pairs are kept |
| Compressed layers | Routed experts (gate_proj, up_proj, down_proj, 896 per layer) in layers 1–92 |
| Left in BF16 | Attention (24 MLA layers and 69 Kimi Delta Attention layers), Attention Residual projections, the latent projections around the routed experts (routed_expert_up_proj, routed_expert_down_proj), router (gate), shared experts, layer-0 dense MLP, embeddings, lm_head, norms, vision tower |
| Format | compressed-tensors, nvfp4-pack-quantized, with a paired48_sparse storage marker |
The router's e_score_correction_bias is stored in BF16 (FP32 in the source).
The scale group is 32 because the Blackwell sparse NVFP4 MMA (SM100 and SM120) requires one scale per 32 dense K elements, which is 16 surviving elements after the 4:8 prune.
| Component | Kimi-K3 (MXFP4 source) | This model |
|---|---|---|
| Routed experts (92 layers × 896, 2.72 T params) | 1,347.1 GiB (MXFP4) | 871.7 GiB (paired-4:8 NVFP4, sparse storage) |
| Everything else (BF16) | 106.6 GiB | 106.5 GiB |
| Total | 1,453.7 GiB | 978.2 GiB |
Sparse storage layout (paired48_sparse, pair-bitmask v1)
Packed NVFP4 stores two adjacent K elements per byte, so one pair is one byte. A paired-4:8 weight therefore has exactly 2 nonzero bytes in every 4-byte chunk. Each routed-expert linear stores:
| Tensor | dtype / shape | Contents |
|---|---|---|
weight_sparse_packed |
uint8 [out, K/4] |
the 2 kept bytes of every 4, in K order |
weight_sparse_mask |
uint8 [out, K/16] |
4 bits per 4-byte chunk (exactly 2 set; bit i means byte i is kept); the low nibble is the lower-K chunk |
weight_scale |
float8_e4m3 [out, K/32] |
block scales (linear layout) |
weight_global_scale, input_global_scale |
float32 | per-tensor global scales |
quantization_config carries "paired48_sparse": {"layout": "pair-bitmask", "version": 1}.
If a 4-byte chunk has fewer than 2 nonzero bytes, its lowest-index zero bytes are marked as
kept. The encoding is lossless: converting to and from dense NVFP4 is exact. At load time vLLM
rebuilds each layer's dense packed weight and compresses it into the kernel's own layout.
Load-time memory is therefore the sparse checkpoint plus one dense layer, not the whole dense
model.
Recipe (MoESQ, arm "gw2")
The full MoESQ config for this run is in moe_sq_config.yaml.
- Source: the MXFP4 routed experts are dequantized to BF16 before compression.
- Init: GPTQ, 4-bit, 512 calibration samples, percdamp 0.1, block size 128. Masks come from paired-4:8 pruning.
- Refinement: masks and weight values are learned jointly for 10 epochs on 2,048 mixed
calibration sequences of up to 4,096 tokens, with activations fake-quantized to NVFP4. The
objective is the block-output reconstruction error weighted by the router gate
(
gate_weight_exponent = 2). The settings are Kimi-K2.5's, untuned, except for fewer calibration sequences (Attention Residuals make each activation cache about 10× larger). - Compute: 2 nodes × 8 B200, about 23 hours end to end.
License
This model is a derivative of Kimi-K3 and is released under the Kimi K3 License (Copyright (c) 2026 Moonshot AI), which is included in this repository. Among its conditions: Model-as-a-Service businesses with more than 20 million US dollars in revenue over 12 months need a separate agreement with Moonshot AI for commercial use, and commercial products above 100 million monthly active users or 20 million US dollars in monthly revenue must prominently display "Kimi K3". See the license for the full terms.
Citation
@misc{lee2026hardwarenativejointsparsequantizationtrillionscale,
title={Hardware-Native Joint Sparse-Quantization for Trillion-Scale Mixture-of-Experts},
author={Kwanhee Lee and Namhoon Lee and Dan Alistarh},
year={2026},
eprint={2610.02241},
archivePrefix={arXiv},
primaryClass={cs.AR},
doi={10.48550/arXiv.2610.02241},
url={https://arxiv.org/abs/2610.02241},
}
Please also cite Kimi-K3 as described in the original model card.
Contact
For questions, open a discussion here or contact kwanhee.lee@postech.ac.kr.
- Downloads last month
- 233
Model tree for ISTA-DASLab/Kimi-K3-P48NVFP4-MoESQ
Base model
moonshotai/Kimi-K3