💻 Instruct tuning simllama 1.1
simonko
simonko912
AI & ML interests
guy with low end hardware learning about llm's and learning about finetuning
my hardware:
intel xeon E5v4
32gb ddr4 ram
amd rx 5700 8gb vram (Hf allows to set from 6600 lol)
512gb ssd
queueing models for mradermacher
contact:
dc: simonko_11015
github: Simonko-912
Recent Activity
repliedto Banaxi-Tech's post about 6 hours ago
Introducing BananaMind 2 Nano
BananaMind 2 Nano is the smallest member of the BananaMind 2.0 family — a 10M-parameter language model that shows how much you can squeeze out of a tiny footprint. It uses the family's digit-isolated tokenizer, so it keeps solid arithmetic despite its size, and it's small enough to run just about anywhere.
Trained on 30B tokens in about a day on a single RTX 5070 Ti (16GB), 4096-token context.
Benchmarks:
Average 35.77
ARC Easy 36.20
PIQA 55.98
ARC Challenge 23.38
HellaSwag 27.50
That 35.77 average edges out Pythia-31M (~34.79) at roughly a third the parameters.
Released under Apache 2.0 on Hugging Face: https://huggingface.co/BananaMind/BananaMind-2-Nano — weights, tokenizer, and config included.