Decision 1.0 β€” Kai 0.6B

Decision-1.0-Kai-0.6B

Kai, from kairos β€” the right moment to choose.

Open Decision Foundation Models

Choose an action, judge a condition, or score against your own rubric. Kai reads the context and candidate descriptions together, then returns structured decisions and probability distributions.

Decision collection Β· Download

Measured decisions

53.52 overall β€” above both general Laya models, and +7.04 points over the previous Kai release.

Benchmark group Kai Β· 0.6B Laya English Laya Multilingual
Decisions 57.96 56.54 47.25
Composition 40.83 35.33 38.92
Reading 54.69 51.41 50.78
Inference 69.79 63.75 57.29
Transfer 48.37 53.06 47.13
Weighted overall 53.52 51.03 47.19

On the benchmark's 160-question BoolQ slice, Kai scores 74.38%, versus 69.38% for each Laya reference.

Same 54 tasks and 3,766 scored questions, with fixed source/family weights. Both Laya references use released general weights, without target-dataset fine-tuning. This is an observed regression suite; results vary by task. Full matrix, paired intervals, and methods.

Three ways to decide

Choice Noul Score
Choose among your actions or categories. Check whether a condition is true. Rate against ordered criteria.
Candidate IDs and probabilities. Probability of yes. Level distribution and expected score.

Use the System One format: state / model / questions β†’ answers. Ask many questions about one context, or apply shared questions across a batch of contexts. Candidates are supplied at runtime.

SDK and curl examples Β· Training-source attribution

128 mixed questions in 163 ms β€” 58% lower latency. Automatic typed scheduling accelerates the measured SystemOne runtime with the same weights. Paired local AMD measurements on a fixed workload. Latency and scaling.

Download for local inference

hf download vllm-sr/Decision-1.0-Kai-0.6B --local-dir Decision-1.0-Kai-0.6B

This repository contains model files, provenance and the Transformers loading code. Serving requires a compatible vLLM Semantic Router Decision runtime, distributed separately. Check its hardware support before serving. The runtime must support vllm-sr-decision format version 1 and the file map in config.json. For local inference without a server, see Use with πŸ€— Transformers below.

Use with πŸ€— Transformers

The repository includes its inference code, so stock Transformers can download and run the complete model locally with trust_remote_code=True. system_one takes and returns the same System One request and response bodies as the Decision runtime; nothing is generated.

pip install "transformers>=4.57" torch safetensors huggingface_hub
from transformers import AutoModel

model = AutoModel.from_pretrained("vllm-sr/Decision-1.0-Kai-0.6B", trust_remote_code=True)
response = model.system_one(
    state="The parcel arrived damaged. Please send a replacement today.",
    questions={
        "route": {
            "type": "choice",
            "instructions": "Which team should handle this request?",
            "criteria": {"delivery": "Damaged or missing parcels", "billing": "Payments and invoices"},
        },
        "urgent": {
            "type": "noul",
            "instructions": "Does the customer request action today?",
        },
    },
)
print(response["answers"]["route"]["choice"], response["answers"]["urgent"]["noul"])

pipeline("decision", model="vllm-sr/Decision-1.0-Kai-0.6B", trust_remote_code=True) accepts the same request body. A malformed question is answered with an invalid_question error. The model loads on the first GPU when one is visible, otherwise on the CPU (pass device="cpu" or device="cuda:0" to choose); weights and arithmetic are FP32. A complete question, its candidates and the state are limited to 1,024 tokens; if a question is longer, every question of the request is answered with a max_length_exceeded error and nothing is truncated.

Use

Replace the placeholder with a SystemOne-compatible endpoint configured to serve Decision-1.0-Kai-0.6B, and set DECISION_API_KEY to that endpoint's key.

pip install typesafe-sdk
import os
from typesafe_sdk import TypeSafeClient, Choice, Noul

client = TypeSafeClient(
    api_key=os.environ["DECISION_API_KEY"],
    base_url="https://your-decision-endpoint.example",
    model="Decision-1.0-Kai-0.6B",
)
questions = {
    "route": Choice(instructions="Which team should handle this request?",
                    criteria={"delivery": "Damaged or missing parcels", "billing": "Payments and invoices"}),
    "urgent": Noul(instructions="Does the customer request action today?"),
}
response = client.system_one(state="The parcel arrived damaged. Please send a replacement today.", questions=questions)
print(response.choices["route"].choice, response.nouls["urgent"].noul)

The same request with curl:

curl -X POST https://your-decision-endpoint.example/v1/systemone \
  -H "Authorization: Bearer $DECISION_API_KEY" \
  -H "Content-Type: application/json" \
  --data '{
    "model": "Decision-1.0-Kai-0.6B",
    "state": "The parcel arrived damaged. Please send a replacement today.",
    "questions": {
      "route": {"type": "choice", "instructions": "Which team should handle this request?", "criteria": {"delivery": "Damaged or missing parcels", "billing": "Payments and invoices"}},
      "urgent": {"type": "noul", "instructions": "Does the customer request action today?"}
    }
  }'

Official Python SDK Β· HTTP API

Architecture

Kai architecture

Three 22-layer bidirectional paths share multilingual input embeddings. Each decision type has its own interaction layers and candidate readout. Candidates within a question are scored together; questions are processed in batches.

Choice and Score were updated while preserving the released Noul path exactly. The model files retain the three-path architecture and complete-input contract.

Architecture and readout diagrams Β· Training and release notes

The complete 1,024-token budget includes context, questions, candidates, and special tokens. Longer inputs are rejected. This release's headline evaluation includes English and Chinese; broader multilingual results from previous weights are historical evidence.

Transfer remains behind Laya English; Reading is 1.25 points below the previous Kai release. Probabilities are not guarantees, and candidate order can affect predictions. Full results and limitations Β· AMD runtime measurements

Built on Vela Encoder. Attribution Β· License scope

Downloads last month
100
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for vllm-sr/Decision-1.0-Kai-0.6B

Finetuned
(13)
this model
Finetunes
2 models
Quantizations
1 model

Spaces using vllm-sr/Decision-1.0-Kai-0.6B 4

Collection including vllm-sr/Decision-1.0-Kai-0.6B