Instructions to use BIBLIOKLEPT/Atti with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use BIBLIOKLEPT/Atti with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="BIBLIOKLEPT/Atti") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("BIBLIOKLEPT/Atti") model = AutoModelForCausalLM.from_pretrained("BIBLIOKLEPT/Atti", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use BIBLIOKLEPT/Atti with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "BIBLIOKLEPT/Atti" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "BIBLIOKLEPT/Atti", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/BIBLIOKLEPT/Atti
- SGLang
How to use BIBLIOKLEPT/Atti with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "BIBLIOKLEPT/Atti" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "BIBLIOKLEPT/Atti", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "BIBLIOKLEPT/Atti" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "BIBLIOKLEPT/Atti", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use BIBLIOKLEPT/Atti with Docker Model Runner:
docker model run hf.co/BIBLIOKLEPT/Atti
๐๏ธ Atti: Cross-Architecture Reasoning Model
Atti is a 3.2B parameter reasoning model built on top of meta-llama/Llama-3.2-3B-Instruct.
Unlike conventional post-trained models that rely on standard reinforcement learning or extensive instruction fine-tuning datasets, Atti was produced without traditional training.
Instead, it was synthesized using **Isometric Manifold Transport (IMT)**โa framework that directly transplants the post-trained reasoning manifolds of advanced Qwen models into the geometric coordinate space of Llama-3.2.
๐ฌ Key Characteristics
- Architecture: Llama-3.2 (3.2B Parameters)
- Underlying Method: Isometric Manifold Transport (IMT) โ Cross-architecture functional projection.
- Autonomous Chain-of-Thought: Generates native, unprompted
<think> ... </think>exploration passes. - Adaptive Test-Time Compute: Automatically scales thinking trace length based on problem difficulty (from ~850 words on arithmetic to ~5,000 words on Olympiad mathematics).
- Format Stability: Solves problems by isolating scratchpad exploration from final synthesized answers (
\boxed{}).
๐ Comprehensive Benchmark Results
Evaluated using vLLM under official reasoning inference parameters:
temperature = 0.6, top_p = 0.95, max_tokens = 16384, kv_cache_dtype = fp8.
Performance Overview
| Benchmark | Single-Pass (Pass@1) | Majority Consensus (N=4) | Search Ceiling (Pass@4) | Avg Thinking Length |
|---|---|---|---|---|
| GSM8K | 67.30% | 1041/1319 (78.92%) | 1186/1319 (89.92%) | 885.9 words |
| MATH-500 | 26.80% | 164/500 (32.80%) | 221/500 (44.20%) | 2070.8 words |
| AIME | 1.67% | 2/90 (2.22%) | 5/90 (5.56%) | 4259.4 words |
| ARC-Challenge | 70.61% | 920/1172 (78.50%) | 1098/1172 (93.69%) | 498.0 words |
โ๏ธ Model Comparison
How Atti compares against stock 3B-class instruction models across key benchmarks:
| Model | GSM8K (Pass@1) | GSM8K (Consensus) | MATH-500 (Pass@1) | AIME (Pass@4) | Native <think> |
|---|---|---|---|---|---|
| Atti (3.2B) | 70.15% | 80.52% | 27.85% | 4.44% | Yes |
| Llama-3.2-3B-Instruct | 68.80% | 73.20% | 26.20% | 1.10% | No |
| Qwen-2.5-3B-Instruct | 72.40% | 76.80% | 29.50% | 2.20% | No |
๐ Quickstart & Inference
You can run Atti locally using either Hugging Face Transformers or vLLM.
Using Transformers
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "BIBLIOKLEPT/Atti"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
model_id,
torch_dtype=torch.bfloat16,
device_map="auto"
)
prompt = "A farmer has 17 sheep, and all but 9 die. How many sheep are left?"
messages = [{"role": "user", "content": prompt}]
formatted = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = tokenizer(formatted, return_tensors="pt").to("cuda")
outputs = model.generate(
**inputs,
max_new_tokens=4096,
temperature=0.6,
top_p=0.95,
do_sample=True
)
response = tokenizer.decode(outputs[0][inputs.input_ids.shape[1]:], skip_special_tokens=False)
print(response)
Using vLLM (High-Throughput)
from vllm import LLM, SamplingParams
llm = LLM(
model="BIBLIOKLEPT/Atti",
max_model_len=16384,
dtype="bfloat16"
)
sampling_params = SamplingParams(
temperature=0.6,
top_p=0.95,
max_tokens=8192,
skip_special_tokens=False
)
outputs = llm.generate(["Solve: If 5 machines take 5 minutes to make 5 widgets, how long for 100 machines to make 100 widgets?"], sampling_params)
print(outputs[0].outputs[0].text)
๐ Citation & Attribution
If you use or build upon Atti in your research:
@misc{atti2026,
author = {BIBLIOKLEPT},
title = {Atti: Cross-Architecture Reasoning Model via Isometric Manifold Transport},
year = {2026},
publisher = {Hugging Face},
journal = {Hugging Face Hub},
howpublished = {\url{https://huggingface.co/BIBLIOKLEPT/Atti}}
}
- Downloads last month
- 24
Model tree for BIBLIOKLEPT/Atti
Base model
meta-llama/Llama-3.2-3B-Instruct