Vega-1.6-128M-Base

Note: AI was used in the creation of this project. But that's why you're here, isn't it?


This is the base model for Vega-1.6-128M, an experimental SLM trained on various datasets. (Ha! I've learned from last time!)

First corpus was 3B toks of 33.3/33.3/33.3 fineweb, fineweb-edu, and cosmopedia. Second corpus was 5B tokens taken from a 35/30/20/15 split of NPset-2-Python-Edu, open-web-math, cosmopedia, and fineweb-edu.

Benchmarks

1019 elo on Bananamind Base Bench 1.1, putting it in 17th place as of writing, ahead of GPT-X-125M and behind Supra2-100M.

  • Hellaswag: 29.62%
  • ARC-E: 39.64%
  • ARC-C: 23.89%
  • PIQA: 59.08%
  • Arithmark-3: 35.60%

Model specs

  • Layers: 16
  • Hidden size: 768
  • Attention heads: 12
  • KV heads: 6
  • Context length: Trained on 2,048 tokens
  • Intermediate size: 2048
  • Vocabulary: 32,000
  • Parameters: 128,410,368
Downloads last month
139
Safetensors
Model size
0.1B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for CNWPlayer/Vega-1.6-128M-Base

Finetunes
1 model

Datasets used to train CNWPlayer/Vega-1.6-128M-Base