File size: 6,861 Bytes
abe8eb5
 
 
6ab6114
 
 
 
 
 
 
 
 
 
 
 
 
abe8eb5
6ab6114
 
 
 
abe8eb5
6ab6114
 
 
 
abe8eb5
6ab6114
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
abe8eb5
 
 
 
 
 
 
6ab6114
abe8eb5
 
 
 
 
 
e1acab4
abe8eb5
 
 
 
 
e1acab4
abe8eb5
 
 
 
 
 
c117684
b099f13
 
abe8eb5
 
 
 
25cf91b
 
abe8eb5
 
 
 
 
 
 
7e5e130
abe8eb5
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
c117684
 
 
 
 
 
 
 
 
 
abe8eb5
 
 
 
e1acab4
abe8eb5
 
 
6ab6114
abe8eb5
7e5e130
abe8eb5
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
---
tags:
- 7b
- agentic-coding
- android
- apple-silicon
- attested
- bash
- c
- chain-of-custody
- chinese
- code
- code-completion
- code-generation
- code-infill
- compacted
- compensation-lora
- consumer-gpu
- cpp
- cryptographically-verified
- css
- distillation
- edge-inference
- efficient
- embedded
- english
- forge-alloy
- function-calling
- general
- general-purpose
- go
- head-pruning
- html
- iphone
- java
- javascript
- knowledge-distillation
- kotlin
- llama-cpp
- lm-studio
- local-inference
- lora
- macbook
- mlx
- mobile
- multilingual
- ollama
- on-device
- optimized
- php
- pruned
- python
- qwen
- qwen-coder
- qwen2
- qwen2.5
- qwen2.5-coder
- raspberry-pi
- reproducible
- ruby
- rust
- sql
- swift
- teacher-student
- text-generation
- typescript
- validation-artifact
- versatile
base_model: Qwen/Qwen2.5-Coder-7B
pipeline_tag: text-generation
license: apache-2.0
---

# 12% Pruned, 61.0 HUMANEVAL (base 62.2)

**Qwen2.5-Coder-7B** recovered to within calibration tolerance of the unmodified base via KL-distillation compensation LoRA.

- **HUMANEVAL**: 61.0 (base 62.2, Δ -1.2)
- **HUMANEVAL+PLUS**: 53.0 (base 53.7, Δ -0.7)


<p align="center">
<a href="https://cambriantech.github.io/forge-alloy/verify/#hf.co/continuum-ai/qwen2.5-coder-7b-compacted/resolve/main/v2-7b-coder-compensated.alloy.json@4fe422e9b01fa8f0">
<img src="alloy-qr.png" alt="Verify Chain of Custody" width="160"/>
</a>
</p>

<p align="center">
<a href="https://cambriantech.github.io/forge-alloy/verify/#hf.co/continuum-ai/qwen2.5-coder-7b-compacted/resolve/main/v2-7b-coder-compensated.alloy.json@4fe422e9b01fa8f0"><b>Every claim on this card is verified</b></a><br>
<b>Trust: self-attested</b> · 2 benchmarks · 1 device tested<br>
<a href="https://github.com/CambrianTech/forge-alloy">ForgeAlloy</a> chain of custody · <a href="v2-7b-coder-compensated.alloy.json">Download alloy</a> · Merkle-chained
</p>

---

**Qwen2.5-Coder-7B** with cryptographic provenance via the [ForgeAlloy](https://github.com/CambrianTech/forge-alloy) chain of custody. Scores **61.0 humaneval** against the unmodified base's **62.2**, recovered to within calibration tolerance after head pruning + distillation. Ships with the per-problem evaluation outputs so the score is independently verifiable.


## Benchmarks

| Benchmark | Score | Base | Δ | Verified |
|---|---|---|---|---|
| **humaneval** | **61.0** | 62.2 | -1.2 | ✅ Result hash |
| **humaneval_plus** | **53.0** | 53.7 | -0.7 | ✅ Result hash |


## What Changed (Base → Forged)

| | Base | Forged | Delta |
|---|---|---|---|
| **Pruning** | None | 12% heads (activation-magnitude) | **-12%** params ✅ |
| **compensation-lora** | None | rank=16 | q_proj, k_proj, v_proj, o_proj... |
| **Pipeline** | | prune → lora → lora → eval | 1 cycles |

## Runs On

| Device | Format | Size | Speed |
|--------|--------|------|-------|
| **NVIDIA GeForce RTX 5090** | fp16 | — | Verified |
| MacBook Pro 32GB | fp16 | 8.0GB | Expected |
| MacBook Air 16GB | Q8_0 | ~4.0GB | Expected |
| MacBook Air 8GB | Q4_K_M | ~2.5GB | Expected |
| iPhone / Android | Q4_K_M | ~2.5GB | Expected |

## Quick Start

```python
from transformers import AutoModelForCausalLM, AutoTokenizer

model = AutoModelForCausalLM.from_pretrained("continuum-ai/v2-7b-coder-compensated",
    torch_dtype="auto", device_map="auto")
tokenizer = AutoTokenizer.from_pretrained("continuum-ai/v2-7b-coder-compensated")

inputs = tokenizer("def merge_sort(arr):", return_tensors="pt").to(model.device)
output = model.generate(**inputs, max_new_tokens=200)
print(tokenizer.decode(output[0], skip_special_tokens=True))
```


## Methodology

Produced via head pruning, LoRA fine-tuning, KL-distillation compensation against the unmodified teacher. Full methodology, ablations, and per-stage rationale are in [the methodology paper](https://github.com/CambrianTech/continuum/blob/main/docs/papers/PLASTICITY-COMPACTION.md) and the companion [`MODEL_METHODOLOGY.md`](MODEL_METHODOLOGY.md) in this repository. The pipeline ran as `prune → lora → lora → eval` over 1 cycle on NVIDIA GeForce RTX 5090.

## Limitations

- This model is currently a methodology demonstration rather than a Pareto-optimal artifact at any specific hardware tier. For production code workloads on smaller hardware, the unmodified Qwen2.5-Coder-7B at standard quantization (Q4_K_M / Q5_K_M / Q8_0) may be a better fit pending the larger Qwen3.5+ forges that exercise the pruning dimension where this methodology actually wins.
- Validated on HumanEval / HumanEval+ for English-language Python code completion. Performance on other programming languages, code paradigms (functional, embedded, kernel), or code-adjacent domains (SQL, regex, shell) has not been measured.
- Ships as fp16 only. GGUF quantization tiers (Q5_K_S / Q3_K_M / Q2_K) are not yet published for this artifact; the per-tier comparison from the development log showed base+quant dominates v2+quant at every VRAM tier on the same 7B base, which is why the methodology validation here uses fp16 and the production GGUF publishes are reserved for the Qwen3.5+ forges where the dimension flips.
- Vision modality not yet wired in. The Continuum sensory architecture treats vision as first-class for personas, but this 7B coder artifact is text-only.


## Chain of Custody

Scan the QR or [verify online](https://cambriantech.github.io/forge-alloy/verify/#hf.co/continuum-ai/qwen2.5-coder-7b-compacted/resolve/main/v2-7b-coder-compensated.alloy.json@4fe422e9b01fa8f0). Download the [alloy file](v2-7b-coder-compensated.alloy.json) to verify independently.

| What | Proof |
|------|-------|
| Model weights | `sha256:156247b9f9b25d302651e2540f1dad58d...` |
| Forged on | NVIDIA GeForce RTX 5090, ? |
| Published | [huggingface](https://huggingface.co/continuum-ai/v2-7b-coder-compensated) — 2026-04-08T05:02:57.072577+00:00 |
| Trust level | [`self-attested`](https://github.com/CambrianTech/forge-alloy/blob/main/docs/ATTESTATION.md) |
| Spec | [ForgeAlloy](https://github.com/CambrianTech/forge-alloy) — Rust/Python/TypeScript |

## Make Your Own

Forged with [Continuum](https://github.com/CambrianTech/continuum) — a distributed AI world that runs on your hardware.

<p align="center">
<a href="https://github.com/CambrianTech/continuum"><img src="https://raw.githubusercontent.com/CambrianTech/continuum/main/docs/images/factory.png" alt="Continuum Model Factory" width="400"/></a>
</p>

The Factory configurator lets you design and forge custom models visually — context extension, pruning, LoRA, quantization, vision/audio modalities. Pick your target devices, the system figures out what fits.

[GitHub](https://github.com/CambrianTech/continuum) · [All Models](https://huggingface.co/continuum-ai) · [Forge-Alloy](https://github.com/CambrianTech/forge-alloy)

## License

apache-2.0