⚛️ The Open Quantum Challenge is live. Both seasons open today.
Quantum computers are expensive, queued, and reachable by only a few. So we removed the barrier. With a single GPU and a public harness, anyone can contribute to core quantum-computing problems with no quantum hardware at all. That is the point of this challenge: to move quantum research from a handful of labs to everyone.
• Season 1, Quantum simulation: reproduce a target quantum system on a GPU. • Season 2, QEC decoder: correct errors on coded quantum states.
Answers are held privately and every submission is auto-scored against a frozen ground truth, so the ranking is reproducible and hardware-independent.
🏆 Prize: 2,000 USD total (1,000 USD per season, 🥇600 / 🥈300 / 🥉100)
Getting started takes one step: copy the participation guide from the Space and paste it into Codex or Claude Code, then submit against the harness.
🧠 We just released Darwin-27B-ZTC, a judgment engine that reaches a verdict without generating anything.
Most LLMs answer by generating, decoding one token at a time. Darwin-27B-ZTC takes a different route.
⚙️ How it works 🔹 It makes its call in a single forward pass. 🔹 Zero generated tokens, and no decoding loop. 🔹 That keeps latency and cost far below what a generative model needs.
🎯 What it judges 🔹 It handles several question types: free-form correctness (noul), multiple choice (choice), and scoring (score). 🔹 For each one it hands back a calibrated confidence, not just an answer.
📊 How well calibrated (measured) 🔹 KL 0.204, Brier 0.097, so the confidence it reports lines up with what actually happens. 🔹 0.743 accuracy (zero-shot, general split), across 2,000 judgments with zero errors. 🔹 By type: noul 0.847, choice 0.723, score 0.675. 🔹 None of the benchmark's train split went into it. It is pure zero-shot.
🚀 Where it fits 🔹 Grading at scale, model routing, safety gating, anywhere you want a fast decision without paying for generation.
🏆 It currently sits at #1 on the official typed-decisions leaderboard on Hugging Face (0.743 accuracy, zero-shot).
💻 Data-center AI, now on a laptop: POCKET-Darwin-180B
We're releasing a 4-bit GGUF build of Darwin-180B-RSI, #1 on seven official Hugging Face leaderboards (self-reported), that runs without a GPU.
📦 360 GB → 111 GB (4-bit GGUF, 4 files) 🖥️ No GPU: one server CPU (16 threads) at 18.4–21.0 tokens/s 💻 RTX 5060 laptop (8 GB VRAM) + 32 GB RAM: 4.17 tokens/s 🧊 128 GB mini PC: whole model in memory, no GPU needed 🎯 MMLU-Pro, 2,000 questions, paired: original 87.65% = 4-bit 87.65%
How? · Only ~3B of 180B parameters are active per token (10 of 512 experts) · llama.cpp streams just the needed experts from SSD, so 32 GB RAM is enough · Graft quantization: we took the proven Unsloth UD-Q4_K_XL base build and swapped in only the 300 tensors our RSI training changed (300/300 verified)
Under the hood is Model-level Recursive Self-Improvement. The model solves verifiable problems, keeps only its own solutions that check out as correct, and trains on them. No human-written solutions or reasoning traces.
Built for teams that can't send data to an external cloud (defense, finance, public sector) to run a top-tier model fully offline.