Instructions to use zenlm/zen6-flash with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use zenlm/zen6-flash with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf zenlm/zen6-flash:Q8_0 # Run inference directly in the terminal: llama cli -hf zenlm/zen6-flash:Q8_0
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf zenlm/zen6-flash:Q8_0 # Run inference directly in the terminal: llama cli -hf zenlm/zen6-flash:Q8_0
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf zenlm/zen6-flash:Q8_0 # Run inference directly in the terminal: ./llama-cli -hf zenlm/zen6-flash:Q8_0
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf zenlm/zen6-flash:Q8_0 # Run inference directly in the terminal: ./build/bin/llama-cli -hf zenlm/zen6-flash:Q8_0
Use Docker
docker model run hf.co/zenlm/zen6-flash:Q8_0
- LM Studio
- Jan
- vLLM
How to use zenlm/zen6-flash with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "zenlm/zen6-flash" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "zenlm/zen6-flash", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/zenlm/zen6-flash:Q8_0
- Ollama
How to use zenlm/zen6-flash with Ollama:
ollama run hf.co/zenlm/zen6-flash:Q8_0
- Unsloth Desktop
- Pi
How to use zenlm/zen6-flash with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf zenlm/zen6-flash:Q8_0
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "zenlm/zen6-flash:Q8_0" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use zenlm/zen6-flash with Docker Model Runner:
docker model run hf.co/zenlm/zen6-flash:Q8_0
- Lemonade
How to use zenlm/zen6-flash with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull zenlm/zen6-flash:Q8_0
Run and chat with the model
lemonade run user.zen6-flash-Q8_0
List all available models
lemonade list
- Hermes Agent
How to use zenlm/zen6-flash with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf zenlm/zen6-flash:Q8_0
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default zenlm/zen6-flash:Q8_0
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use zenlm/zen6-flash with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf zenlm/zen6-flash:Q8_0
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "zenlm/zen6-flash:Q8_0" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Zen 6 Flash is the ternary build of Zen 6, from Zen LM, the open model family of Zoo Labs Foundation, a 501(c)(3) non-profit. It is chosen for the same two jobs as Zen 6, where the machine is small:
- Agentic coding that runs on your own machine, in under 8 GB with its vision projector, on an Apple Silicon laptop or one GPU.
- Marketing work, reading images and documents beside the brief.
Zen 6 and Zen 6 Flash are available now; Zen 7 is in research preview:
request access. Zen 6 Flash answers on
api.hanzo.ai as zen6-flash.
What the files are
Every weight is ternary, the embeddings and the LM head included. The language model holds 26,895,998,464 weights (read from the GGUF tensor table), so a file's bits per weight is its size in bits over that count, and its ratio to FP16 is 53.79 GB (the count × 2 bytes) over its size.
| File | Size | Bits per weight | Smaller than FP16 |
|---|---|---|---|
Ternary-Bonsai-2-27B-PTQ1_0.gguf, densely packed |
5.95 GB | 1.77 | 9.0x |
Ternary-Bonsai-2-27B-PQ2_0.gguf, one trit per 2-bit slot, read by GPU and Metal kernels as is |
7.21 GB | 2.14 | 7.5x |
Ternary-Bonsai-2-27B-mmproj-Q8_0.gguf, vision projector |
629 MB | ||
Ternary-Bonsai-2-27B-mmproj-BF16.gguf, vision projector, full precision |
931 MB | ||
Bonsai-2-27B-DFlash2-Q8_0.gguf, speculative drafter |
2.06 GB | ||
SHA256SUMS |
66 KB |
| Context | 262,144 tokens native; about three layers in four are linear attention |
| Inputs | text and images |
| Architecture | GGUF general.architecture: qwen35 |
Accuracy against full precision
| Suite | FP16 | A conventional 2-bit build (IQ2_XXS) | Zen 6 Flash (PQ2_0) | Retained |
|---|---|---|---|---|
| Average of 14 reasoning tests | 86.33 | 72.59 | 84.78 | 98.2% |
| Math (MathVision, GSM8K) | 97.10 | 79.40 | 96.57 | 99.5% |
| Code (HumanEval, LiveCode) | 90.20 | 74.80 | 89.42 | 99.1% |
| Tool calling and schema adherence | 76.80 | 58.10 | 74.92 | 97.6% |
| Footprint | 53.79 GB | 8.8 GB | 7.21 GB | 86.6% smaller |
Measured
- Apple Silicon M4 / M5 Max (Metal): 47.2 tok/s decode alone, 92.4 tok/s with the drafter; 7.84 GB including the vision projector and a 32K cache.
- DGX Spark (CUDA 13.3): 170.04 tok/s decode with the drafter at 54.2% draft acceptance; 32,768 tokens across 4 parallel slots in under 12 GB.
Serve it
llama.cpp, with the vision projector and the drafter:
llama-server \
-m Ternary-Bonsai-2-27B-PQ2_0.gguf \
--mmproj Ternary-Bonsai-2-27B-mmproj-Q8_0.gguf \
-md Bonsai-2-27B-DFlash2-Q8_0.gguf \
--spec-draft-n-max 3 \
-c 32768 \
--port 8080
On a Mac, one prompt:
llama-cli \
-m Ternary-Bonsai-2-27B-PQ2_0.gguf \
--mmproj Ternary-Bonsai-2-27B-mmproj-Q8_0.gguf \
-p "<image>\nDescribe the system architecture shown in this diagram in detail."
Citation
@misc{zenlm2026zen6flash,
title = {Zen 6 Flash},
author = {Zen LM},
year = {2026},
publisher = {Zoo Labs Foundation}
}
License & attribution
Apache-2.0; see LICENSE.
Zen 6 Flash is Ternary Bonsai 2 27B by prism-ml (prism-ml/Ternary-Bonsai-2-27B-gguf,
Apache-2.0), itself built from Qwen3.8-27B by the Qwen team (Apache-2.0), with the
speculative drafter by ProCreations (ProCreations/Ternary-Bonsai-2-27B-DFlash2).
- Downloads last month
- 695
1-bit
2-bit
8-bit
Model tree for zenlm/zen6-flash
Base model
Qwen/Qwen3.8-27B