Instructions to use BIBLIOKLEPT/Nori-2 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use BIBLIOKLEPT/Nori-2 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="BIBLIOKLEPT/Nori-2") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("BIBLIOKLEPT/Nori-2") model = AutoModelForMultimodalLM.from_pretrained("BIBLIOKLEPT/Nori-2", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use BIBLIOKLEPT/Nori-2 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "BIBLIOKLEPT/Nori-2" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "BIBLIOKLEPT/Nori-2", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/BIBLIOKLEPT/Nori-2
- SGLang
How to use BIBLIOKLEPT/Nori-2 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "BIBLIOKLEPT/Nori-2" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "BIBLIOKLEPT/Nori-2", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "BIBLIOKLEPT/Nori-2" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "BIBLIOKLEPT/Nori-2", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use BIBLIOKLEPT/Nori-2 with Docker Model Runner:
docker model run hf.co/BIBLIOKLEPT/Nori-2
🌿 Nori-2: High-Potency Reasoning Model (Gemma 4 E4B)
Nori-2 is an advanced compact reasoning model built on top of google/gemma-4-E4B-it.
Unlike conventional reasoning models that undergo massive reinforcement learning (RL) runs or extensive instruction tuning directly on the target architecture, Nori-2 was produced without traditional training.
Instead, Nori-2 was created using **Isometric Manifold Transport (IMT)**—a framework that directly transplants the post-trained reasoning manifolds of advanced Qwen models into the geometric coordinate space of Gemma 4.
🔬 Architectural Upgrades
Nori-2 incorporates three mathematical upgrades:
- Riemannian Null-Space Completion: Eliminates the 512-dimension orthogonal blind spot between Qwen (2,048) and Gemma (2,560) using covariance regression ($\mathcal{K}$), allowing reasoning updates to flow across 100% of Gemma's 2,560-dimensional residual stream.
- Spectral Wiener Filtering (Rank 96): Replaces harsh SVD cutoffs with quadratic soft-shrinkage, dampening high-frequency coordinate phase noise while preserving dense math circuits.
- Target-Norm Calibration (2.5%): Calibrates every layer’s update to exactly 2.5% of its native baseline Frobenius norm, preventing attention-head saturation.
📊 Benchmark Results
Evaluated across 13,524 total generation trajectories on the full test sets using vLLM with FP8 KV-caching:temperature = 0.6, top_p = 0.95, max_tokens = 8192, N = 4 Rollouts.
| Benchmark | Total Problems | Single-Pass (Pass@1) | Majority Vote (Consensus) | Search Ceiling (Pass@4) |
|---|---|---|---|---|
| AIME | 90 | 40.56% | 43.33% (39 / 90) | 53.33% (48 / 90) |
| MATH-500 | 500 | 54.85% | 57.00% (285 / 500) | 59.20% (296 / 500) |
| ARC-Challenge | 1,172 | 94.03% | 94.71% (1,110 / 1,172) | 97.27% (1,140 / 1,172) |
| SVAMP | 300 | 84.50% | 86.33% (259 / 300) | 89.00% (267 / 300) |
| GSM8K | 1,319 | 79.15% | 79.30% (1,046 / 1,319) | 84.00% (1,108 / 1,319) |
⚔️ Benchmark Comparisons
1. Comparison Across the Nori Evolution & Base Gemma 4
| Benchmark | Base Gemma 4 E4B | Nori-1 | Nori-2 | Net Gain over Base |
|---|---|---|---|---|
| MATH-500 | 32.50% | 54.00% | 57.00% | +24.50% |
| AIME (Consensus) | ~25.00% | 40.00% | 43.33% | +18.33% |
| AIME (Pass@4) | ~30.00% | — | 53.33% | +23.33% |
| ARC-Challenge | 84.20% | 94.45% | 94.71% | +10.51% |
| SVAMP | 76.50% | 82.67% | 86.33% | +9.83% |
| GSM8K | 76.30% | 79.45% | 79.30% | +3.00% |
2. Comparison with Modern Reasoning Models (Majority Vote Consensus)
All models compared using their reported Majority Vote / Consensus metrics:
| Model | Parameters | AIME (Consensus) | MATH-500 | ARC-Challenge | SVAMP | GSM8K |
|---|---|---|---|---|---|---|
| Nori-2 | ~4B | 43.33% | 57.00% | 94.71% | 86.33% | 79.30% |
| Nanbeige4.2-3B | 3B | 48.20% | 67.60% | 88.50% | 85.00% | 92.70% |
| Nanbeige4.1-3B | 3B | 38.60% | 51.20% | 86.40% | 81.50% | 85.90% |
| VibeThinker-3B | 3B | 94.30%* | 83.40% | Not Reported | Not Reported | 91.80% |
| Qwen-2.5-3B-Instruct | 3B | 3.30% | 38.50% | 78.50% | 78.00% | 79.20% |
| Gemma 4 E4B (Base) | ~4B | ~25.00% | 32.50% | 84.20% | 76.50% | 76.30% |
| Llama-3.2-3B-Instruct | 3B | 1.10% | 26.20% | 74.10% | 72.50% | 68.80% |
*VibeThinker's 94.30% AIME score was achieved on AIME 2026 using its official sampling configuration.
🚀 Quickstart & Inference
Using vLLM (High Throughput)
from vllm import LLM, SamplingParams
llm = LLM(
model="BIBLIOKLEPT/Nori-2",
max_model_len=16384,
dtype="bfloat16",
kv_cache_dtype="fp8"
)
sampling_params = SamplingParams(
temperature=0.6,
top_p=0.95,
max_tokens=8192,
skip_special_tokens=False
)
outputs = llm.generate(
["A farmer has 17 sheep, and all but 9 die. How many sheep are left?"],
sampling_params
)
print(outputs[0].outputs[0].text)
Using Transformers
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "BIBLIOKLEPT/Nori-2"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
model_id,
torch_dtype=torch.bfloat16,
device_map="auto"
)
prompt = "A farmer has 17 sheep, and all but 9 die. How many sheep are left?"
conv = [{"role": "user", "content": prompt}]
formatted = tokenizer.apply_chat_template(
conv,
tokenize=False,
add_generation_prompt=True,
enable_thinking=True
)
inputs = tokenizer(formatted, return_tensors="pt").to("cuda")
outputs = model.generate(
**inputs,
max_new_tokens=2048,
temperature=0.6,
top_p=0.95,
do_sample=True
)
response = tokenizer.decode(outputs[0][inputs.input_ids.shape[1]:], skip_special_tokens=False)
print(response)
📜 Citation & Attribution
@misc{nori2_2026,
author = {BIBLIOKLEPT},
title = {Nori-2: High-Potency Reasoning Model via Isometric Manifold Transport and Null-Space Completion},
year = {2026},
publisher = {Hugging Face},
journal = {Hugging Face Hub},
howpublished = {https://huggingface.co/BIBLIOKLEPT/Nori-2}
}
- Downloads last month
- 53
Model tree for BIBLIOKLEPT/Nori-2
Evaluation results
- Accuracy (Majority Consensus) on AIMEAIME 2026 Leaderboard43.330
- Accuracy (Majority Consensus) on MATH-500test set MATH-500 Benchmark57.000
- Accuracy (Majority Consensus) on ARC Challengetest set Open LLM Leaderboard94.710
- Accuracy (Majority Consensus) on SVAMPtest set SVAMP Benchmark86.330
- Accuracy (Majority Consensus) on GSM8Ktest set Open LLM Leaderboard79.300