TinyStories GPT (51M)

A small GPT-2-style language model (51.2M parameters) trained from scratch on TinyStories. Give it the start of a story and it continues it in the simple, child-friendly style of the dataset.

It was trained in about 2.2 hours on a single free-tier NVIDIA T4 GPU. It is a learning project and a demo, not a general-purpose assistant.

Prompt:  Once upon a time
Output:  Once upon a time, there was a little boy named Timmy. He loved to play with his toys all day long. One day, ...

Quick start

This is a custom PyTorch implementation, so it does not load through transformers. Everything needed to run it is in two files: model.pt (weights) and model_runner.py (model code + generation CLI, no other imports from this repo).

# 1. download both files
pip install huggingface_hub
hf download 57Ajay/tinystories-gpt-51m --local-dir tinystories-gpt
cd tinystories-gpt

# 2. generate (uv installs torch + tiktoken automatically)
uv run model_runner.py --prompt "Once upon a time" --max_new_tokens 200

More examples:

uv run model_runner.py                                   # interactive mode, Ctrl-D to quit
uv run model_runner.py --prompt "One day, a little robot" -t 0.7 --top_p 0.9 -n 3
uv run model_runner.py --prompt "" --seed 42             # unconditional story
Option Default Meaning
--prompt none Text to continue. Omit for interactive mode
--max_new_tokens, -m 200 Maximum number of tokens to generate
--temperature, -t 0.8 Randomness. 0 = greedy (deterministic)
--top_k 50 Sample only from the k most likely tokens (0 disables)
--top_p 1.0 Nucleus sampling threshold (1.0 disables)
--num_samples, -n 1 Number of independent continuations
--seed none Fix the random seed for reproducible output
--device auto cuda if available, otherwise cpu

The model is small enough to run comfortably on a CPU. Generation stops early when the model emits its end-of-story token (<|endoftext|>).

Model details

Architecture Decoder-only Transformer (GPT-2 style), pre-LayerNorm
Parameters 51,237,888 (about 25.2M in the transformer blocks, about 25.8M in the shared embedding)
Layers / heads / hidden size 8 / 8 / 512
Context length 512 tokens
Tokenizer GPT-2 BPE (tiktoken, gpt2 encoding)
Vocabulary 50,304 (GPT-2's 50,257, padded to a multiple of 64; padding ids are never produced)
Position encoding Learned absolute embeddings
Activation GELU (tanh approximation)
Embeddings Input and output embeddings are tied
Dropout None
Weight file model.pt: a PyTorch dict with model (state dict), config (architecture) and val_loss

model.pt can be read with torch.load("model.pt", weights_only=True), so loading it does not execute arbitrary code.

Training

Data

The train split of roneneldan/TinyStories (about 2.1M short stories written by GPT-3.5 / GPT-4 using a vocabulary a young child would understand; see Eldan & Li, 2023), tokenized into about 474M tokens. Each story was prefixed with <|endoftext|> and the stories were concatenated into one stream. Training saw roughly 328M tokens, so about 0.69 epochs (each token was seen at most once).

Procedure

Steps 5,000
Tokens per step 65,536 (micro-batch 16 x 512 tokens, 8 gradient-accumulation steps)
Total tokens about 328M
Optimizer AdamW, betas (0.9, 0.95), eps 1e-8, weight decay 0.1 (applied to matrices only, not to biases or LayerNorm)
Learning rate 1e-3 peak, 100 steps linear warmup, then cosine decay to 1e-4
Gradient clipping Global norm 1.0
Initialization Normal(0, 0.02); residual output projections scaled by (2 x layers)^-0.5
Precision fp16 mixed precision with loss scaling (a T4 has no bf16)
Compilation torch.compile
Hardware 1x NVIDIA T4 (Google Colab), about 41-47k tokens/s, about 2.2 hours total

Evaluation

Metric Value
Validation loss (cross-entropy, nats per token) 1.352
Validation perplexity about 3.86

Validation loss was measured on the first about 164K tokens (20 batches of 16 x 512) of the TinyStories validation split, using the same batches at every evaluation. It was still improving slowly at the end of training (1.355 at step 4,750, 1.352 at step 5,000), so a longer run would likely do better. The numbers come from a small slice of the validation set and are meant as a sanity check, not a benchmark. Loss is measured per GPT-2 token, so it is not comparable to results that use a different tokenizer.

Sample outputs

Generated with the training-time sampler (prompt Once upon a time, temperature 0.8, top-k 50). Openings only:

Once upon a time, there was a little girl named Lily. She had a small notebook where she drew beautiful ...

Once upon a time, there was a little boy named Timmy. He loved to play with his toys all day long. One day, ...

Intended use

  • Learning, teaching and experimenting with how small language models are trained
  • A small, fast baseline for research on tiny models or on TinyStories itself
  • A starting point for fine-tuning or architecture experiments

Limitations and out-of-scope use

  • Narrow domain. The model only knows the style and content of TinyStories: very simple English, young-child vocabulary, and a handful of recurring themes. Prompts far from simple children's stories usually produce off-topic or incoherent text.
  • Repetitive. Expect recurring names (Lily, Timmy, ...) and plot patterns, and loss of coherence over longer outputs.
  • Not an assistant. It cannot follow instructions, answer questions, write code, or be relied on for facts.
  • English only, with a 512-token context window.
  • No safety tuning or filtering. The training data is synthetic and child-friendly, but outputs are not guaranteed to be appropriate and may reflect biases in the data. Do not use it for anything safety-critical or user-facing without your own review.

Files

File Purpose
model.pt Trained weights and architecture config
model_runner.py Single-file inference script (model definition + CLI)
README.md This model card

License

The model weights and code are released under the MIT license. The training data (TinyStories) has its own license, CDLA-Sharing-1.0; please review it if you plan to use or redistribute the model commercially.

Acknowledgements and citation

@article{eldan2023tinystories,
  title   = {TinyStories: How Small Can Language Models Be and Still Speak Coherent English?},
  author  = {Eldan, Ronen and Li, Yuanzhi},
  journal = {arXiv preprint arXiv:2305.07759},
  year    = {2023}
}
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Dataset used to train 57Ajay/tinystories-gpt-51m

Paper for 57Ajay/tinystories-gpt-51m