Kestrel Decider 230M

A 230M-parameter typed decision model by Kestrelyn: give it a state and a question, get back a calibrated yes/no, choice among N options, or score on a rubric — never free text. It is built with the strands-decider recipe, retrained on LiquidAI/LFM2.5-230M-Base instead of Qwen3.5-2B-Base: about 1/9th the parameters of strands-decider-2B-hobson-v19, for routing, triage, tool selection and guardrails where latency, memory or CPU-only serving matter more than the last points of accuracy.

Independent release. Not made, reviewed or endorsed by the strands-decider authors or by Liquid AI. Licensed under the LFM Open License v1.0 (see License) — not Apache-2.0 like v19.

Versions

  • v1.1 (this) — LoRA rank 128 (from 16) plus the 8 depthwise short convolutions trained in full; every BoardgameQA train row (+12,000); strands-decider v19 as teacher on the answer-adequacy, generated-question and BoardgameQA rows where v19 is right.
  • v1.0 — First release: strands-decider's v19 recipe and corpus on LFM2.5-230M-Base, rank-16 LoRA, 3 epochs.

Earlier versions are git tags. To use one, download it and load the folder: hf download Kestrelyn/kestrel-decider-230m --revision v1.0 --local-dir kestrel-v1.0, then StrandsDeciderModel.load("kestrel-v1.0").

Quick start

pip install strands-decider
strands-decider ask Kestrelyn/kestrel-decider-230m \
  --state "Help! My payouts have been failing for 3 days!" \
  --choice "Which team?=billing,sales,retail"
from strands_decider.infer import SystemOneEngine, EngineConfig
from strands_decider.modeling import StrandsDeciderModel
from strands_decider.schema import ChoiceQuestion, NoulQuestion

model = StrandsDeciderModel.load("Kestrelyn/kestrel-decider-230m")
engine = SystemOneEngine(model, EngineConfig(device="cuda"))  # or "cpu"
r = engine.ask("Help! My payouts have been failing for 3 days!", {
    "team": ChoiceQuestion(instructions="Which team?",
                           criteria={"billing": None, "sales": None, "retail": None}),
    "urgent": NoulQuestion(instructions="The customer needs a reply today."),
})
print(r.answers["team"].choice, r.answers["team"].probabilities)

strands-decider serve Kestrelyn/kestrel-decider-230m exposes the same OpenAI-style /v1/systemone endpoint as v19. The 2 MB head and ~124 MB LoRA adapter load on top of the base model, which is fetched from LiquidAI/LFM2.5-230M-Base (0.5 GB) on first use. Requires transformers>=5.15.

Results

v1.1 (this model) v1.0 v19 (Qwen3.5-2B, reference)
JevBench v1 public (231 tasks)³ 0.645 0.636 0.723
JevBench Brier ↓ 0.474 0.487 0.342
JevBench ECE ↓ 0.064 0.090 0.052
Unseen classification tasks, overall¹ 0.558 0.567 0.647
  yes/no · choice · score 0.559 · 0.606 · 0.442 0.572 · 0.618 · 0.439 0.596 · 0.725 · 0.499
Unseen classification ECE ↓ 0.058 0.052 0.054
MuSiQue (multi-hop QA) 0.828 0.871 0.879
ContractNLI 0.827 0.828 0.862
BoardgameQA (rule reasoning) 0.818 0.596 0.810
HotpotQA (never trained on) 0.543 0.552 0.726
Generated document questions, v16 set · v18 set² 0.751 · 0.696 0.691 · 0.697 0.846 · 0.757
Answer adequacy, HelpSteer2 held-out (balanced) 0.598 0.564 0.739

¹ Four classification tasks absent from training (emotion, MASSIVE intents, sarcasm, hate-speech severity), 6,000 rows. Rebuilt from the Hub in October 2026: the file's sha256 differs from the one strands-decider recorded (an upstream dataset revision), so compare this row to v19 approximately; JevBench and the multi-step sets are exact. ² Pooled over each set's skills, weighted by count. ³ Accuracy on the JevBench v1 public set (original 72 + easy 48 + hard 111 tasks), run with the official harness at commit 1bcc55e against strands-decider serve at a 3,072-token window — the same measurement as v19's published 0.723. It is not the Capability Score of the current JevBench board (v1.5.5: 1,624 decisions incl. 720 sealed), which this model has not been submitted to.

Latency (JevBench, one RTX 3060 12 GB, batch of one request): p50 19 ms, p95 126 ms.

Where it is good: facts, extraction, intent, routing and tool selection over a provided state (JevBench families at 0.83–1.00), and question answering over a document it is given (MuSiQue 0.828 vs v19's 0.879; ContractNLI 0.827); rule reasoning, BoardgameQA 0.818, level with v19's 0.810. Its confidence is calibrated on short classification (ECE 0.058 on unseen tasks, v19 0.054), so thresholding on confidence works there; on JevBench's harder mix it is less well calibrated (ECE 0.064 vs v19's 0.052).

Where it is weak — use v19 or larger instead: judging whether a response is adequate (HelpSteer2 0.598, chance 0.5), chaining facts across sources it was not trained on (HotpotQA; JevBench multi_hop), and ambiguous or hard-judgement items.

Known regression in v1.1: MuSiQue fell from 0.871 to 0.828, all of it on the answerable questions (0.843 → 0.763; the unanswerable ones held): v1.1 more readily answers that a multi-hop document question cannot be answered. If that is your main use, use the v1.0 tag (see Versions).

JevBench by family

family accuracy tasks
routing_hard 1.000 5
tool_selection 1.000 12
extraction 0.917 24
routing 0.917 12
fact 0.833 12
intent 0.833 24
tradeoff 0.667 6
trap 0.625 8
ambiguous 0.571 7
long_policy 0.526 19
adversarial 0.500 6
ordinal 0.500 12
probability 0.500 10
judge_hard 0.471 17
temporal_numeric 0.467 15
adequacy 0.417 12
policy 0.417 12
multi_hop 0.389 18

Training data

The strands-decider v19 corpus, rebuilt from public sources with the recipe's own scripts (training/recipe.sh build fetch multistep generated adequacy); the training files matched the recipe's recorded sha256 byte for byte. 135,339 rows: 131,263 training, 4,076 validation.

Short classification tasks (100,449 rows, at most 24 options each), rendered as typed questions:

question type datasets
choice fancyzhx/ag_news, legacy-datasets/banking77, clinc/clinc_oos, fancyzhx/dbpedia_14, papluca/language-identification, community-datasets/yahoo_answers_topics
yes/no ucirvine/sms_spam, google/civil_comments, nyu-mll/glue (MNLI, WNLI), google/boolq, google-research-datasets/paws, tals/vitaminc, qiaojin/PubMedQA, tasksource/ruletaker (depths 0–2)
score Yelp/yelp_review_full, SetFit/sst5, sealuzh/app_reviews, osyvokon/pavlick-formality-scores

Multi-step documents (12,909 rows): ContractNLI (CC BY 4.0, Koreeda & Manning 2021), MuSiQue (CC BY 4.0, Trivedi et al. 2022; answerable and unanswerable versions of each question), BoardgameQA (tasksource/Boardgame-QA, CC BY 4.0, Kazemi et al. 2023). Plus 12,000 further BoardgameQA train rows that the corpus build caps out (every remaining train question that fits the window), so all ~15,000 are used.

Generated document questions (3,815 rows) and answer adequacy (6,166 rows): questions written by Qwen3.6-27B and kept when Qwen3.5-397B-A17B agreed (committed in the strands-decider repo), and nvidia/HelpSteer2 human ratings (CC BY 4.0) thresholded to adequate / inadequate.

Teacher distributions (training targets, not rows): 90,350 rows in all. Frozen Qwen/Qwen3.5-4B on the 59,525 yes/no and choice short-task rows where it agrees with the gold label; strands-decider v14 on the ContractNLI and MuSiQue rows; and strands-decider v19 (Apache-2.0) on 20,916 BoardgameQA, answer-adequacy and generated-question rows, kept only where v19's answer is the gold label.

Held out, never trained on: dair-ai/emotion, mteb/amazon_massive_intent, raquiba/Sarcasm_News_Headline, ucberkeley-dlab/measuring-hate-speech (calibration and the unseen-task evaluation), tasksource/ruletaker depths 3, 5 and NatLang, and hotpotqa/hotpot_qa (CC BY-SA 4.0, evaluation only).

Per-source licences and attribution: strands-decider's data/sources.md and THIRD_PARTY_NOTICES.md. Check each source's terms before redistributing data built from it.

How it was trained

  • Recipe: strands-decider's v19 recipe and training corpus plus every BoardgameQA train row — the v5 classification corpus, 24,909 multi-step rows (ContractNLI, MuSiQue, BoardgameQA), generated document questions and answer-adequacy rows; option order shuffled per example. Config: train_config.json (reproduces the run with strands-decider train --config train_config.json).
  • Torso: LFM2.5-230M-Base (8 short-conv + 6 attention layers), frozen, with rank-128 LoRA (alpha 256) on q_proj, k_proj, v_proj, out_proj, in_proj, w1, w2, w3; the 8 depthwise short convolutions (conv.conv, which LoRA cannot target) are trained in full and saved with the adapter (31.6M trainable parameters).
  • Readout: pointer head (dim 256) scoring each option from its own hidden state.
  • Teacher: KL toward the frozen Qwen3.5-4B's distributions on the yes/no and choice short-task rows where it agrees with the gold label, and toward strands-decider v14's distributions on the ContractNLI and MuSiQue rows (strands-decider's v20 teacher file); toward strands-decider v19's on the BoardgameQA, adequacy and generated-question rows where v19 is right (rows v19 gets wrong train on the gold label alone); plus a KL anchor (weight 0.3) to the untouched base's own option-number readout.
  • Schedule: 3 epochs, effective batch 32, LoRA lr 0.0001, head lr 0.001, cosine; max_length 3072.
  • Hardware: 2× RTX 3060 12 GB (DDP), Ubuntu 22.04, torch 2.11 / transformers 5.18 / peft 0.21.
  • Calibration: temperatures fitted per question type on held-out data (yes/no 1.13, choice 0.91, score 1.19), stored in strands_decider_config.json.

Training is single-seed; differences of a couple of points between runs are within noise.

Files

file
lora/ PEFT LoRA adapter for the LFM2.5 decoder (incl. the trained short convolutions)
head.safetensors pointer readout head
strands_decider_config.json model config incl. calibration
train_config.json, history.json training config and loss/validation curve
evals/, results.json JevBench summary, internal eval output, every number above
LICENSE, NOTICE LFM Open License v1.0; attribution and modification notice
MANIFEST.sha256 sha256 of every file

License

This model is a Derivative Work of LFM2.5-230M-Base and is distributed under the LFM Open License v1.0. In short (the LICENSE governs): use, modification and redistribution are permitted, but commercial use is licensed only to entities with annual revenue under US$10 million; redistributions must include the license and mark modifications. See NOTICE for attribution, including the strands-decider recipe (Apache-2.0) and its training-data sources.

Credits

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Kestrelyn/kestrel-decider-230m

Adapter
(4)
this model

Datasets used to train Kestrelyn/kestrel-decider-230m

Evaluation results