Pythia 1.4B -- 16,750 Facts via Neural Reclamation

This model has 16,750 facts written directly into its weights using jBlaze's Neural Reclamation pipeline. This is a capacity stress test -- how far can a tiny 1.4B model go before it breaks?

The facts are sourced from CounterFact-Tracing (Wikidata relation types -- locations, citizenships, languages, genres, and more). The full training corpus is included in this repo as facts_maxout.json (27,081 entries total -- this model was trained on the first 16,750).

The Answer: It Did Not Break

Metric Baseline (0 facts) After 16,750 Facts
General Capability 60.0% 55.0%
Perplexity 13.3 18.3
Direct Recall -- 49.8%
Paraphrase Recall -- 49.3%

After absorbing 16,750 facts, the model lost only 5 percentage points of general capability and perplexity increased by just 5 points. The model still produces coherent, natural English. It still reasons. It still works.

Nearly half of all 16,750 facts are directly retrievable -- on a model with only 1.4 billion parameters. That is approximately one retrievable fact per 170,000 parameters.

For Comparison

Gradient-based training destroyed this same architecture at 50 facts (LoRA v1) or 125 facts (LoRA v2 with maximum conservative settings). Neural Reclamation loaded 335x more facts than gradient training's breaking point, and the model is still functional.

Method Facts General Cap Perplexity Model Status
Gradient (LoRA v1) 50 10.0% 459.6 Dead
Gradient (LoRA v2) ~125 35.0% 96.5 Dead
Neural Reclamation (500 novel facts) 500 Stable Stable Alive, 98% recall
Neural Reclamation (this model) 16,750 55.0% 18.3 Alive

Cohort Recall

Facts are loaded in batches. Earlier cohorts get partially overwritten as new facts compete for the same weight capacity. This is expected behavior -- the model has a finite number of parameters.

Cohort Recall
1-1,000 54%
1,001-1,500 50%
1,501-2,000 56%
2,001-2,500 60%
2,501-3,000 46%
3,001-3,500 40%
3,501-4,000 54%
4,001-4,500 46%
4,501-5,000 52%
5,001-5,500 40%
5,501-6,000 42%
6,001-6,500 64%
6,501-7,000 32%
7,001-7,500 46%
7,501-8,000 40%
8,001-8,500 52%
8,501-9,000 46%
9,001-9,500 52%
9,501-10,000 50%
10,001-10,500 50%
10,501-11,000 50%
11,001-11,500 54%
11,501-12,000 52%
12,001-12,500 54%
12,501-13,000 54%
13,001-13,500 42%
13,501-14,000 46%
14,001-14,500 42%
14,501-15,000 54%
15,001-15,500 56%
15,501-16,000 60%
16,001-16,500 56%
16,501-16,750 82%

Recall is distributed across all cohorts -- the model did not just memorize the last batch and forget everything else. Knowledge is genuinely distributed through the weights. The final cohort (16,501-16,750) shows elevated recall because it was the most recently written.

See Also

Technology

Built with jBlaze -- weight-level surgery for large language models.

Usage

from transformers import AutoModelForCausalLM, AutoTokenizer

model = AutoModelForCausalLM.from_pretrained("ApolloRaines/Pythia-1.4B-DNP-16750-Facts")
tokenizer = AutoTokenizer.from_pretrained("ApolloRaines/Pythia-1.4B-DNP-16750-Facts")

A Note on Our Released Models

Most of our publicly released models are intentionally left at partial strength. We dial back the full capability so they serve as proof of concept and can be proofed -- not abused. The point is to show what's possible, not to hand it out at full power. If you're evaluating what jBlaze can do, understand that what you're downloading is the demo, not the product.

Downloads last month
1,196
Safetensors
Model size
1B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for ApolloRaines/Pythia-1.4B-DNP-16750-Facts

Finetuned
(76)
this model