TinyStories GPT (27.6M)
A small GPT-style transformer, built and trained from scratch, that generates short children's stories. Pretrained on TinyStories and instruction-tuned on TinyStories-Instruct to follow prompts specifying required words, features, and a summary.
This model was built as part of a hands-on ML Systems Engineering learning project — architecture, tokenizer, training loop, and fine-tuning were all implemented from scratch rather than using existing training frameworks.
Model Details
- Architecture: GPT-style decoder-only transformer (custom implementation)
- Parameters: ~27.6M
- Vocabulary size: 4,096 (custom-trained BPE tokenizer)
- Context length: 512 tokens
- Embedding dimension: 512
- Attention heads: 8
- Decoder layers: 8
- Tokenizer: Custom BPE, trained on the TinyStories corpus (not a standard HF tokenizer — see Usage)
Training
- Pretraining on TinyStories — next-token prediction over short children's stories.
- Supervised fine-tuning on TinyStories-Instruct — trained to generate a story given structured prompts specifying required words, a narrative feature (e.g. dialogue, twist), and a summary. Loss was masked on the prompt tokens so only completion tokens contributed to training.
Example
Prompt:
Features: Twist
Words: cat, jump, happy
Summary: A cat jumps on a table and makes a friend.
Story:
Output:
One day, a cat and a dog were playing in the park. The cat was happy and
excited. She was running and jumping all day long. The dog was happy too.
They both loved to play together.
While playing, the cat saw a big table. The table had a high roof. The cat
thought it was a good place to rest. She jumped on the table and pretended
to be a cat. The dog barked and the cat laughed.
Then, the cat jumped off the table and went to sleep. The cat and the dog
were very happy. They played together until it was time to go home. They
knew they would have fun soon.
Usage
This model uses a custom architecture and tokenizer, so it can't be loaded directly with AutoModel/AutoTokenizer. You'll need the model class definition and the tokenizers library.
from huggingface_hub import hf_hub_download
from tokenizers import Tokenizer
from safetensors.torch import load_file
import torch
# Download files
model_path = hf_hub_download(repo_id="harshit147/tinystories-gpt-27m", filename="model_sft.safetensors")
tokenizer_path = hf_hub_download(repo_id="harshit147/tinystories-gpt-27m", filename="tinystories_bpe.json")
# Load tokenizer
tok = Tokenizer.from_file(tokenizer_path)
# Load model (requires the GPT class definition from the source repo)
model = GPT(vocab_size=4096, d_model=512, context_length=512, num_heads=8, num_layers=8)
state_dict = load_file(model_path)
model.load_state_dict(state_dict)
model.eval()
# Generate
prompt = "Features: Twist\nWords: cat, jump, happy\nSummary: A cat jumps on a table and makes a friend.\nStory:"
tokens = tok.encode(prompt).ids
context = torch.tensor(tokens).unsqueeze(0)
# ... run your generate() function
The model architecture and generation code are available at [https://github.com/HarshitP147/tinystories-gpt-27m].
Prompt Format
The model expects prompts in this structure:
Features: <narrative feature, e.g. Dialogue, Twist, MoralValue>
Words: <comma-separated words to include>
Summary: <one-line summary of the story>
Story:
Limitations
- Trained exclusively on TinyStories — vocabulary and world knowledge are limited to the simple, child-directed language and concepts in that dataset (small vocabulary, short sentences, simple objects/emotions/events).
- Not suitable for factual questions, reasoning tasks, or anything outside the TinyStories domain.
- Context window is capped at 512 tokens.
Project Context
Built as a from-scratch learning project covering transformer architecture, custom tokenizer training, pretraining, and instruction fine-tuning. Follow-up work includes inference optimization (KV-caching, quantization, batching) — see [GitHub repo link] for progress.
- Downloads last month
- 194