🌿 Nori-2: High-Potency Reasoning Model (Gemma 4 E4B)

Nori-2 is an advanced compact reasoning model built on top of google/gemma-4-E4B-it.

Unlike conventional reasoning models that undergo massive reinforcement learning (RL) runs or extensive instruction tuning directly on the target architecture, Nori-2 was produced without traditional training.

Instead, Nori-2 was created using **Isometric Manifold Transport (IMT)**—a framework that directly transplants the post-trained reasoning manifolds of advanced Qwen models into the geometric coordinate space of Gemma 4.


🔬 Architectural Upgrades

Nori-2 incorporates three mathematical upgrades:

  1. Riemannian Null-Space Completion: Eliminates the 512-dimension orthogonal blind spot between Qwen (2,048) and Gemma (2,560) using covariance regression ($\mathcal{K}$), allowing reasoning updates to flow across 100% of Gemma's 2,560-dimensional residual stream.
  2. Spectral Wiener Filtering (Rank 96): Replaces harsh SVD cutoffs with quadratic soft-shrinkage, dampening high-frequency coordinate phase noise while preserving dense math circuits.
  3. Target-Norm Calibration (2.5%): Calibrates every layer’s update to exactly 2.5% of its native baseline Frobenius norm, preventing attention-head saturation.

📊 Benchmark Results

Evaluated across 13,524 total generation trajectories on the full test sets using vLLM with FP8 KV-caching:
temperature = 0.6, top_p = 0.95, max_tokens = 8192, N = 4 Rollouts.

Benchmark Total Problems Single-Pass (Pass@1) Majority Vote (Consensus) Search Ceiling (Pass@4)
AIME 90 40.56% 43.33% (39 / 90) 53.33% (48 / 90)
MATH-500 500 54.85% 57.00% (285 / 500) 59.20% (296 / 500)
ARC-Challenge 1,172 94.03% 94.71% (1,110 / 1,172) 97.27% (1,140 / 1,172)
SVAMP 300 84.50% 86.33% (259 / 300) 89.00% (267 / 300)
GSM8K 1,319 79.15% 79.30% (1,046 / 1,319) 84.00% (1,108 / 1,319)

⚔️ Benchmark Comparisons

1. Comparison Across the Nori Evolution & Base Gemma 4

Benchmark Base Gemma 4 E4B Nori-1 Nori-2 Net Gain over Base
MATH-500 32.50% 54.00% 57.00% +24.50%
AIME (Consensus) ~25.00% 40.00% 43.33% +18.33%
AIME (Pass@4) ~30.00% — 53.33% +23.33%
ARC-Challenge 84.20% 94.45% 94.71% +10.51%
SVAMP 76.50% 82.67% 86.33% +9.83%
GSM8K 76.30% 79.45% 79.30% +3.00%

2. Comparison with Modern Reasoning Models (Majority Vote Consensus)

All models compared using their reported Majority Vote / Consensus metrics:

Model Parameters AIME (Consensus) MATH-500 ARC-Challenge SVAMP GSM8K
Nori-2 ~4B 43.33% 57.00% 94.71% 86.33% 79.30%
Nanbeige4.2-3B 3B 48.20% 67.60% 88.50% 85.00% 92.70%
Nanbeige4.1-3B 3B 38.60% 51.20% 86.40% 81.50% 85.90%
VibeThinker-3B 3B 94.30%* 83.40% Not Reported Not Reported 91.80%
Qwen-2.5-3B-Instruct 3B 3.30% 38.50% 78.50% 78.00% 79.20%
Gemma 4 E4B (Base) ~4B ~25.00% 32.50% 84.20% 76.50% 76.30%
Llama-3.2-3B-Instruct 3B 1.10% 26.20% 74.10% 72.50% 68.80%

*VibeThinker's 94.30% AIME score was achieved on AIME 2026 using its official sampling configuration.


🚀 Quickstart & Inference

Using vLLM (High Throughput)

from vllm import LLM, SamplingParams

llm = LLM(
    model="BIBLIOKLEPT/Nori-2",
    max_model_len=16384,
    dtype="bfloat16",
    kv_cache_dtype="fp8"
)

sampling_params = SamplingParams(
    temperature=0.6,
    top_p=0.95,
    max_tokens=8192,
    skip_special_tokens=False
)

outputs = llm.generate(
    ["A farmer has 17 sheep, and all but 9 die. How many sheep are left?"], 
    sampling_params
)

print(outputs[0].outputs[0].text)

Using Transformers

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "BIBLIOKLEPT/Nori-2"

tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
    model_id,
    torch_dtype=torch.bfloat16,
    device_map="auto"
)

prompt = "A farmer has 17 sheep, and all but 9 die. How many sheep are left?"
conv = [{"role": "user", "content": prompt}]

formatted = tokenizer.apply_chat_template(
    conv, 
    tokenize=False, 
    add_generation_prompt=True, 
    enable_thinking=True
)
inputs = tokenizer(formatted, return_tensors="pt").to("cuda")

outputs = model.generate(
    **inputs,
    max_new_tokens=2048,
    temperature=0.6,
    top_p=0.95,
    do_sample=True
)

response = tokenizer.decode(outputs[0][inputs.input_ids.shape[1]:], skip_special_tokens=False)
print(response)

📜 Citation & Attribution

@misc{nori2_2026,
  author = {BIBLIOKLEPT},
  title = {Nori-2: High-Potency Reasoning Model via Isometric Manifold Transport and Null-Space Completion},
  year = {2026},
  publisher = {Hugging Face},
  journal = {Hugging Face Hub},
  howpublished = {https://huggingface.co/BIBLIOKLEPT/Nori-2}
}
Downloads last month
53
Safetensors
Model size
8B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for BIBLIOKLEPT/Nori-2

Finetuned
(410)
this model
Quantizations
1 model

Evaluation results