Sprout story model
A decoder-only model trained from scratch on TinyStoriesV2-GPT4 and instruction tuned on TinyStoriesInstruct. It has 29,368,832 parameters, 8 layers, width 512, 8 attention heads, a 512-token context, rope positions, RMSNorm, SwiGLU, tied embeddings, and a 4096-token byte-level BPE vocabulary.
This bundle contains float32 inference weights at instruction step 18,457. It excludes optimizer state, training data, credentials, and training dependencies. The tied embedding/output tensor is stored once and restored with safetensors.torch.load_model. config.json records weight and tokenizer SHA-256 checksums, source checkpoint SHA-256, source code commit, and the original training fingerprint.
Try the live browser demo. Float32 ONNX inference runs on your device; prompts stay local.
Run locally
Download all files from this model repository into an empty directory, then run:
python -m pip install -r requirements.txt
python app.py
Use the included loading module directly:
from inference import load, stream
model, tok, config = load(".")
for text in stream(model, tok, {"Summary": "A little fox learns to be patient."}):
print(text, end="\r", flush=True)
load also accepts a Hub model ID and an optional revision commit. The included code is the loading contract; this is not a Transformers AutoModel repository. Loading does not require any training corpus files. The demo uses the exact training template, keeps the full instruction in context, and supports streamed output, stopping, temperature, and top-k.
Training and evaluation
The source checkpoint's instruction validation loss is 0.956147. The training run records configuration and losses. See the project README for the fixed behavioral evaluation, comparison with the pretrained model, and base-validation regression. Low validation loss does not establish reliable constraint following.
Pretraining uses the story-level seeded split in the project, with 537,686,596 training tokens and 5,402,089 validation tokens. Instruction preparation excludes malformed and duplicate stories, held-out story overlap, and conversations beyond the context limit. TinyStoriesInstruct is pinned to ee050ed1f8720795be342921335e821856a2b42e. These are synthetic English stories generated by GPT models, not a general conversational corpus.
Limitations and intended use
An educational story-generation experiment. It often misses requested words and sentences, repeats itself, and loses narrative details. Instruction tuning regresses base-story validation loss. Long instructions leave less room for the story. It has no general chat or factual-answering capability, no content filter, and no suitability evaluation for children. Outputs can contain biases or disturbing story content from its synthetic source corpus.
License and provenance
Code and exported weights use the included MIT license. The upstream TinyStories dataset card lists CDLA-Sharing-1.0; TinyStoriesInstruct is the separate instruction corpus. No source stories are redistributed in this bundle. The dataset creators' paper describes TinyStories.
Published behavioral results
On 20 fixed held-out instruction prompts using matching CUDA bf16 settings, the tuned model included all requested words in 6/20 cases, included the exact requested sentence in 4/20, and emitted the assistant end token in 19/20. Base-story validation loss rose 11.81% after instruction tuning. This is a completed educational experiment with limited instruction following. Complete unedited outputs and the project results include settings and limitations.
- Downloads last month
- 50