Instructions to use Kestrelyn/kestrel-decider-230m with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use Kestrelyn/kestrel-decider-230m with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
Kestrel Decider 230M
A 230M-parameter typed decision model by Kestrelyn: give it a state and a question, get back a calibrated yes/no, choice among N options, or score on a rubric — never free text. It is built with the strands-decider recipe, retrained on LiquidAI/LFM2.5-230M-Base instead of Qwen3.5-2B-Base: about 1/9th the parameters of strands-decider-2B-hobson-v19, for routing, triage, tool selection and guardrails where latency, memory or CPU-only serving matter more than the last points of accuracy.
Independent release. Not made, reviewed or endorsed by the strands-decider authors or by Liquid AI. Licensed under the LFM Open License v1.0 (see License) — not Apache-2.0 like v19.
Versions
- v1.1 (this) — LoRA rank 128 (from 16) plus the 8 depthwise short convolutions trained in full; every BoardgameQA train row (+12,000); strands-decider v19 as teacher on the answer-adequacy, generated-question and BoardgameQA rows where v19 is right.
- v1.0 — First release: strands-decider's v19 recipe and corpus on LFM2.5-230M-Base, rank-16 LoRA, 3 epochs.
Earlier versions are git tags. To use one, download it and load the folder: hf download Kestrelyn/kestrel-decider-230m --revision v1.0 --local-dir kestrel-v1.0, then StrandsDeciderModel.load("kestrel-v1.0").
Quick start
pip install strands-decider
strands-decider ask Kestrelyn/kestrel-decider-230m \
--state "Help! My payouts have been failing for 3 days!" \
--choice "Which team?=billing,sales,retail"
from strands_decider.infer import SystemOneEngine, EngineConfig
from strands_decider.modeling import StrandsDeciderModel
from strands_decider.schema import ChoiceQuestion, NoulQuestion
model = StrandsDeciderModel.load("Kestrelyn/kestrel-decider-230m")
engine = SystemOneEngine(model, EngineConfig(device="cuda")) # or "cpu"
r = engine.ask("Help! My payouts have been failing for 3 days!", {
"team": ChoiceQuestion(instructions="Which team?",
criteria={"billing": None, "sales": None, "retail": None}),
"urgent": NoulQuestion(instructions="The customer needs a reply today."),
})
print(r.answers["team"].choice, r.answers["team"].probabilities)
strands-decider serve Kestrelyn/kestrel-decider-230m exposes the same OpenAI-style /v1/systemone
endpoint as v19. The 2 MB head and ~124 MB LoRA adapter load on top of the base model,
which is fetched from 0.5 GB) on first use.
Requires LiquidAI/LFM2.5-230M-Base (transformers>=5.15.
Results
| v1.1 (this model) | v1.0 | v19 (Qwen3.5-2B, reference) | |
|---|---|---|---|
| JevBench v1 public (231 tasks)³ | 0.645 | 0.636 | 0.723 |
| JevBench Brier ↓ | 0.474 | 0.487 | 0.342 |
| JevBench ECE ↓ | 0.064 | 0.090 | 0.052 |
| Unseen classification tasks, overall¹ | 0.558 | 0.567 | 0.647 |
| yes/no · choice · score | 0.559 · 0.606 · 0.442 | 0.572 · 0.618 · 0.439 | 0.596 · 0.725 · 0.499 |
| Unseen classification ECE ↓ | 0.058 | 0.052 | 0.054 |
| MuSiQue (multi-hop QA) | 0.828 | 0.871 | 0.879 |
| ContractNLI | 0.827 | 0.828 | 0.862 |
| BoardgameQA (rule reasoning) | 0.818 | 0.596 | 0.810 |
| HotpotQA (never trained on) | 0.543 | 0.552 | 0.726 |
| Generated document questions, v16 set · v18 set² | 0.751 · 0.696 | 0.691 · 0.697 | 0.846 · 0.757 |
| Answer adequacy, HelpSteer2 held-out (balanced) | 0.598 | 0.564 | 0.739 |
¹ Four classification tasks absent from training (emotion, MASSIVE intents, sarcasm,
hate-speech severity), 6,000 rows. Rebuilt from the Hub in October 2026: the file's
sha256 differs from the one strands-decider recorded (an upstream dataset revision), so
compare this row to v19 approximately; JevBench and the multi-step sets are exact.
² Pooled over each set's skills, weighted by count.
³ Accuracy on the JevBench v1 public set (original 72 + easy 48 + hard 111 tasks), run
with the official harness at commit 1bcc55e against strands-decider serve at a
3,072-token window — the same measurement as v19's published 0.723. It is not the
Capability Score of the current JevBench board (v1.5.5: 1,624 decisions incl. 720 sealed),
which this model has not been submitted to.
Latency (JevBench, one RTX 3060 12 GB, batch of one request): p50 19 ms, p95 126 ms.
Where it is good: facts, extraction, intent, routing and tool selection over a provided state (JevBench families at 0.83–1.00), and question answering over a document it is given (MuSiQue 0.828 vs v19's 0.879; ContractNLI 0.827); rule reasoning, BoardgameQA 0.818, level with v19's 0.810. Its confidence is calibrated on short classification (ECE 0.058 on unseen tasks, v19 0.054), so thresholding on confidence works there; on JevBench's harder mix it is less well calibrated (ECE 0.064 vs v19's 0.052).
Where it is weak — use v19 or larger instead: judging whether a response is
adequate (HelpSteer2 0.598, chance 0.5), chaining
facts across sources it was not trained on (HotpotQA; JevBench multi_hop), and
ambiguous or hard-judgement items.
Known regression in v1.1: MuSiQue fell from 0.871 to 0.828, all of it on the answerable questions (0.843 → 0.763; the unanswerable ones held): v1.1 more readily answers that a multi-hop document question cannot be answered. If that is your main use, use the v1.0 tag (see Versions).
JevBench by family
| family | accuracy | tasks |
|---|---|---|
| routing_hard | 1.000 | 5 |
| tool_selection | 1.000 | 12 |
| extraction | 0.917 | 24 |
| routing | 0.917 | 12 |
| fact | 0.833 | 12 |
| intent | 0.833 | 24 |
| tradeoff | 0.667 | 6 |
| trap | 0.625 | 8 |
| ambiguous | 0.571 | 7 |
| long_policy | 0.526 | 19 |
| adversarial | 0.500 | 6 |
| ordinal | 0.500 | 12 |
| probability | 0.500 | 10 |
| judge_hard | 0.471 | 17 |
| temporal_numeric | 0.467 | 15 |
| adequacy | 0.417 | 12 |
| policy | 0.417 | 12 |
| multi_hop | 0.389 | 18 |
Training data
The strands-decider v19 corpus, rebuilt from public sources with the recipe's own
scripts (training/recipe.sh build fetch multistep generated adequacy); the training
files matched the recipe's recorded sha256 byte for byte. 135,339 rows: 131,263 training, 4,076 validation.
Short classification tasks (100,449 rows, at most 24 options each), rendered as typed questions:
| question type | datasets |
|---|---|
| choice | fancyzhx/ag_news, legacy-datasets/banking77, clinc/clinc_oos, fancyzhx/dbpedia_14, papluca/language-identification, community-datasets/yahoo_answers_topics |
| yes/no | ucirvine/sms_spam, google/civil_comments, nyu-mll/glue (MNLI, WNLI), google/boolq, google-research-datasets/paws, tals/vitaminc, qiaojin/PubMedQA, tasksource/ruletaker (depths 0–2) |
| score | Yelp/yelp_review_full, SetFit/sst5, sealuzh/app_reviews, osyvokon/pavlick-formality-scores |
Multi-step documents (12,909 rows): ContractNLI (CC BY 4.0, Koreeda & Manning 2021),
MuSiQue (CC BY 4.0, Trivedi et al. 2022; answerable and unanswerable versions of each
question), BoardgameQA (tasksource/Boardgame-QA, CC BY 4.0, Kazemi et al. 2023).
Plus 12,000 further BoardgameQA train rows that the corpus build caps out (every remaining train question that fits the window), so all ~15,000 are used.
Generated document questions (3,815 rows) and answer adequacy (6,166 rows):
questions written by Qwen3.6-27B and kept when Qwen3.5-397B-A17B agreed (committed in
the strands-decider repo), and nvidia/HelpSteer2 human ratings (CC BY 4.0) thresholded
to adequate / inadequate.
Teacher distributions (training targets, not rows): 90,350 rows in all. Frozen Qwen/Qwen3.5-4B on the 59,525 yes/no and choice short-task rows where it agrees with the gold label; strands-decider v14 on the ContractNLI and MuSiQue rows; and strands-decider v19 (Apache-2.0) on 20,916 BoardgameQA, answer-adequacy and generated-question rows, kept only where v19's answer is the gold label.
Held out, never trained on: dair-ai/emotion, mteb/amazon_massive_intent,
raquiba/Sarcasm_News_Headline, ucberkeley-dlab/measuring-hate-speech (calibration and
the unseen-task evaluation), tasksource/ruletaker depths 3, 5 and NatLang, and
hotpotqa/hotpot_qa (CC BY-SA 4.0, evaluation only).
Per-source licences and attribution: strands-decider's
data/sources.md
and THIRD_PARTY_NOTICES.md. Check each source's terms before redistributing data built
from it.
How it was trained
- Recipe: strands-decider's v19 recipe and training corpus plus every BoardgameQA train row — the v5
classification corpus, 24,909 multi-step rows (ContractNLI, MuSiQue, BoardgameQA),
generated document questions and answer-adequacy rows; option order shuffled per
example. Config:
train_config.json(reproduces the run withstrands-decider train --config train_config.json). - Torso: LFM2.5-230M-Base (8 short-conv + 6 attention layers), frozen, with rank-128
LoRA (alpha 256) on
q_proj, k_proj, v_proj, out_proj, in_proj, w1, w2, w3; the 8 depthwise short convolutions (conv.conv, which LoRA cannot target) are trained in full and saved with the adapter (31.6M trainable parameters). - Readout: pointer head (dim 256) scoring each option from its own hidden state.
- Teacher: KL toward the frozen Qwen3.5-4B's distributions on the yes/no and choice short-task rows where it agrees with the gold label, and toward strands-decider v14's distributions on the ContractNLI and MuSiQue rows (strands-decider's v20 teacher file); toward strands-decider v19's on the BoardgameQA, adequacy and generated-question rows where v19 is right (rows v19 gets wrong train on the gold label alone); plus a KL anchor (weight 0.3) to the untouched base's own option-number readout.
- Schedule: 3 epochs, effective batch 32, LoRA lr 0.0001, head lr 0.001,
cosine;
max_length3072. - Hardware: 2× RTX 3060 12 GB (DDP), Ubuntu 22.04, torch 2.11 / transformers 5.18 / peft 0.21.
- Calibration: temperatures fitted per question type on held-out data
(yes/no 1.13, choice 0.91, score 1.19), stored in
strands_decider_config.json.
Training is single-seed; differences of a couple of points between runs are within noise.
Files
| file | |
|---|---|
lora/ |
PEFT LoRA adapter for the LFM2.5 decoder (incl. the trained short convolutions) |
head.safetensors |
pointer readout head |
strands_decider_config.json |
model config incl. calibration |
train_config.json, history.json |
training config and loss/validation curve |
evals/, results.json |
JevBench summary, internal eval output, every number above |
LICENSE, NOTICE |
LFM Open License v1.0; attribution and modification notice |
MANIFEST.sha256 |
sha256 of every file |
License
This model is a Derivative Work of LFM2.5-230M-Base and is distributed under the
LFM Open License v1.0. In short (the LICENSE governs): use, modification
and redistribution are permitted, but commercial use is licensed only to entities with
annual revenue under US$10 million; redistributions must include the license and mark
modifications. See NOTICE for attribution, including the strands-decider recipe
(Apache-2.0) and its training-data sources.
Credits
- strands-decider — the recipe, data pipeline, training and serving code (Apache-2.0).
- Liquid AI — LFM2.5-230M-Base.
- JevBench — external benchmark.
- Downloads last month
- -
Model tree for Kestrelyn/kestrel-decider-230m
Base model
LiquidAI/LFM2.5-230M-BaseDatasets used to train Kestrelyn/kestrel-decider-230m
google-research-datasets/paws
fancyzhx/ag_news
Evaluation results
- accuracy (149/231) on JevBench public, served at 3072self-reported0.645
- accuracy (n=6000) on eval: held-out short tasksself-reported0.558
- accuracy (n=900) on eval: boardgameself-reported0.818
- accuracy (n=1026) on eval: contractnliself-reported0.827
- accuracy (n=959) on eval: hotpotqa (held out)self-reported0.543
- accuracy (n=1199) on eval: musiqueself-reported0.828
- accuracy (n=350) on eval: generated_v16_evalself-reported0.751
- accuracy (n=247) on eval: generated_v18_evalself-reported0.696