DistilQwen
Proof-weighted distillation, Qwen3-30B to 1.7B/0.6B. Three teachers: Instruct, Thinking, Coder. The core method series. DOI 10.57967/hf/8165
Text Generation • 2B • Updated • 4.76k • 2Note 30B teacher, 1.7B student. Proof-weighted KD at 2.25× on reasoning.
reaperdoesntknow/Qwen3-1.7B-Distilled-30B-A3B-SFT-GGUF
Text Generation • 2B • Updated • 2.12kNote Most downloaded GGUF in the collection. CPU-friendly.
reaperdoesntknow/Qwen3-1.7B-Distilled-30B-A3B-SFT
2B • Updated • 146Note Instruct distillation + SFT. Full precision. BF16 H100.
reaperdoesntknow/Qwen3-0.6B-Distilled-30B-A3B
Text Generation • 0.8B • Updated • 4.56kNote 50× compression: 30B → 0.6B. Smallest in the distil family.
reaperdoesntknow/Qwen3-0.6B-Distilled-30B-A3B-Thinking-SFT
Text Generation • 0.8B • Updated • 4.68k • 2Note Thinking teacher + SFT at 0.6B. Extended deliberation traces.
reaperdoesntknow/Qwen3-0.6B-Distilled-30B-A3B-Thinking-SFT-GGUF
Text Generation • 0.8B • Updated • 2.16kNote Edge deployment of extended-thinking at 0.6B. Apache 2.0.
reaperdoesntknow/Qwen3-1.7B-Coder-Distilled-SFT
Text Generation • 2B • Updated • 4.7k • 1Note Coder teacher produces uniquely structured distributions.
reaperdoesntknow/Qwen3-1.7B-Coder-Distilled-SFT-GGUF
Text Generation • 2B • Updated • 2.55k • 2Note Structured reasoning for edge deployment. Apache 2.0.
reaperdoesntknow/DistilQwen3-1.7B-uncensored
Text Generation • 2B • Updated • 3.77kNote Uncensored base distillation. No alignment filtering.
reaperdoesntknow/TopologicalQwen
Text Generation • 2B • Updated • 5.54k • 1Note TKD flagship. BV decomposition → jump detection → curriculum.
reaperdoesntknow/DiStil-Qwen3-1.7B-uncensored
2B • Updated • 401 • 1Note Bridge between base distil and topology-aware models.
reaperdoesntknow/Disctil-Qwen3-1.7B
Text Generation • 2B • Updated • 3.46k • 1Note DISC-refined. Discrepancy-aware training produces cleaner signal.
reaperdoesntknow/DistilQwen3-1.7B-uncensored-GGUF
2B • Updated • 2.91k • 3Note Community validated — third-party quantizations exist.
reaperdoesntknow/Qwen3-1.7B-Thinking-Distil
Text Generation • 2B • Updated • 4.89k • 2Note The most popular model. Thinking teacher = richest signal.
reaperdoesntknow/LFM2.5-1.2B-Distilled-SFT
Text Generation • 1B • Updated • 3.76kNote Proves TKD works across architecture families, not just within Qwen.
reaperdoesntknow/Discrepancy_Calculus
UpdatedNote Continuous Thought Dynamics — mathematical backbone of DualMind.