TinyBrainBot-25M-v4-Base

A 25M-parameter language model trained from scratch on 11B tokens, part of the TinyBrainBot v4 series. 25.2M parameters in the transformer itself, plus a 5.4M tied embedding for its 14k vocabulary (30.6M total). It is a base model: it continues text.

It beats Pythia-70M, which has more than twice the parameters and saw 300B tokens, on 10 of 13 benchmarks (ties a further one), and averages 38.5 against Pythia-70M's 35.9.

Benchmark (0-shot) TinyBrainBot-25M-v4 Pythia-31M Pythia-70M
ARC-Easy 42.7 33.7 36.0
ARC-Easy (raw acc) 47.3 37.2 38.2
ARC-Challenge 24.9 21.0 21.8
ARC-Challenge (raw acc) 21.0 16.6 17.6
SciQ 65.9 51.5 56.4
SciQ (raw acc) 73.7 60.0 64.0
OpenBookQA 31.2 26.8 25.4
OpenBookQA (raw acc) 17.2 13.0 12.6
MMLU 23.7 22.9 22.9
HellaSwag 37.2 33.6 33.8
HellaSwag (raw acc) 32.5 31.2 30.8
PIQA 60.2 56.6 59.2
PIQA (raw acc) 60.0 57.5 59.8
WinoGrande 52.6 49.3 52.9
Social IQa 37.2 35.1 35.4
CommonsenseQA 20.3 19.7 19.8
LAMBADA (OpenAI) 23.4 17.8 23.4
BoolQ 58.8 51.0 59.6
MathQA 21.9 19.5 20.4
MathQA (raw acc) 23.2 19.2 20.9
Average 38.5 33.7 35.9

EleutherAI lm-eval-harness v0.4.13, 0-shot, up to 2,000 examples per task, acc_norm where the task reports it and acc otherwise. All three models were run by us on the same harness and settings.

Quick start

from transformers import AutoModelForCausalLM, AutoTokenizer

repo = "nkthebass/tinybrainbot-25m-v4-base"
tok = AutoTokenizer.from_pretrained(repo)
model = AutoModelForCausalLM.from_pretrained(repo)

ids = tok("Once upon a time, a little fox", return_tensors="pt").input_ids
out = model.generate(ids, max_new_tokens=80)       # defaults from generation_config.json
print(tok.decode(out[0], skip_special_tokens=True))

GGUF (llama.cpp, LM Studio, Ollama): tinybrainbot-25m-v4-base-F16.gguf (59 MB) and tinybrainbot-25m-v4-base-Q8_0.gguf (31 MB) are in this repo. They tokenize and generate identically to the transformers version.

llama-cli -m tinybrainbot-25m-v4-base-Q8_0.gguf -p "Once upon a time, a little fox" -n 80 --temp 0.7 --repeat-penalty 1.15

Default sampling: temperature 0.7, top_p 0.95, top_k 40, repetition_penalty 1.15.

Samples

Unedited, at the default settings.

Once upon a time, a little fox named Mia decided to start her own small business selling handmade crafts at the local community center. She wanted to make sure that everyone had enough supplies and opportunities to buy their favorite items. So, she gathered some friends and helped them find a creative idea for making their own goods.

Dear diary, today was a fun and exciting day with lots of memories! I found out that every single day, we learned about the fascinating world of animal behavior and how it affects us all. So let's continue exploring this amazing subject together!

Photosynthesis is the process by which light energy is converted into chemical energy (e.g., glycolysis). Photosynthetic bacteria convert sunlight, carbon dioxide, and water to glucose (a sugar-sweetened food), oxygen, and nitrogen, respectively.

Model

Architecture Llama-compatible decoder, RoPE, SwiGLU, RMSNorm, grouped-query attention
Shape hidden 384, 16 layers, 6 heads of 64, 2 KV heads, FFN 1024
Parameters 25.18M non-embedding + 5.38M tied embedding = 30.55M
Context 1,024 tokens
Vocabulary 14,000

Deep and narrow on purpose: at this size, adding layers helps more than adding width.

Tokenizer

A new 14k byte-level BPE trained for this model on a sample of its own training mix:

  • Newlines and indentation survive exactly, so code, lists and verse keep their shape.
  • Digits are split one per token. "1234" and "1243" share their pieces, which is what small models need to learn arithmetic.
  • No unknown tokens. Every byte is in the base alphabet.
  • Chat and reasoning markers (<|user|>, <|assistant|>, <|end|>, <think>, </think>) are single tokens, reserved for fine-tuning. <|endoftext|> (id 0) separates documents.

Training

11.0B tokens on 2x Tesla V100 16GB, about 22 hours, fp16, AdamW (0.9, 0.95), weight decay 0.1, peak LR 2e-3, 131k tokens per step, warmup-stable-decay schedule.

Stage 1 - 9.7B tokens, general mix. Every source under 3 epochs.

Group Share Sources
Web 37% FineWeb-Edu 28, DCLM 9
Reference 13.5% Wikipedia article leads 12, full Wikipedia 1.5
Narrative 11% Project Gutenberg 8, BookCorpusOpen 3
Code 9% The Stack (smol) 7.5, Python/JS 1.5
Math 9% MegaMath-Web-Pro 6, FineMath-4+ 3
Synthetic textbooks 7.5% Cosmopedia (Stanford, OpenStax, auto-math)
Q&A and facts 7% distilled Q&A, short facts, cause-and-effect explanations
Educational PDFs 5.5% FinePDFs-Edu
Reasoning 0.5% distilled short reasoning

Stage 2 - 500M tokens. The same mix with 7.8% added: everyday physical-commonsense writing (2.5%), procedural "goal and method" writing (1.8%), worked order-of-operations problems (2%) and short facts (+1.5%). All of it is generated text written for this purpose; none of it is benchmark data. Stage 1 ended with a short LR decay, which stage 2 re-warmed.

Stage 3 - 803M token anneal. LR decays to zero over a fact-dense mix: Cosmopedia, FineMath, FinePDFs-Edu, Wikipedia leads, distilled Q&A and reasoning, plus the stage 2 additions.

The packed corpus was block-shuffled and its document density checked shard by shard before training, so every stage is an even mix rather than one source after another.

Reproduce the benchmarks

lm_eval --model hf --model_args pretrained=nkthebass/tinybrainbot-25m-v4-base,dtype=float32 \
  --tasks arc_easy,arc_challenge,sciq,openbookqa,mmlu,hellaswag,piqa,winogrande,social_iqa,commonsense_qa,lambada_openai,boolq,mathqa \
  --num_fewshot 0 --limit 2000 --batch_size 8 --seed 1234

TinyBrainBot v4

Downloads last month
286
Safetensors
Model size
30.6M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Datasets used to train nkthebass/tinybrainbot-25m-v4-base

Collection including nkthebass/tinybrainbot-25m-v4-base

Evaluation results