Xunzhuo commited on
Commit
d074d4b
·
1 Parent(s): fdc57f3

feat(decision): add strict public API and streamed Tetris races

Browse files

Signed-off-by: Xunzhuo <Xunzhuo@users.noreply.huggingface.co>

.dockerignore CHANGED
@@ -7,6 +7,7 @@
7
  !engine.py
8
  !relay.py
9
  !model_registry.py
 
10
  !examples.json
11
  !static/
12
  !static/**
 
7
  !engine.py
8
  !relay.py
9
  !model_registry.py
10
+ !tetris_arena.py
11
  !examples.json
12
  !static/
13
  !static/**
.gitignore ADDED
@@ -0,0 +1,4 @@
 
 
 
 
 
1
+ __pycache__/
2
+ *.py[cod]
3
+ .ruff_cache/
4
+ .pytest_cache/
DEPLOYMENT.md CHANGED
@@ -1,8 +1,8 @@
1
  # Gateway and resident workers
2
 
3
- The Docker Space is a CPU-only gateway: no weights or tokenization. One gateway process maintains up to four independent bounded queues, selected by an explicit model ID and pinned manifest. Keep one replica because queue state is in memory.
4
 
5
- Kai retains `decision-nano-preview`; the other wire IDs are `decision-lex`, `decision-sol` and `decision-nox`. `/v1/models` exposes actual sizes, complete-input limits and per-model examples. `/api/status?model=…` reports the selected queue. Omitting `model` selects Kai.
6
 
7
  Configure `DECISION_BACKEND=pull_queue`, the private `DECISION_WORKER_TOKEN` secret, and `DECISION_MODEL_REGISTRY_V2`. Each registry item includes `id`, `label`, `version`, `manifest_sha256`, `complete_input_tokens` and optional presentation/example metadata. The old `DECISION_MODEL_REGISTRY` remains a fallback for legacy two-model deployments. The newer variable allows existing code to ignore it safely during a coordinated upgrade.
8
 
@@ -14,8 +14,22 @@ For upgrades, verify idle queues and preserve exact previous sources and registr
14
 
15
  ## Shared-question context batches
16
 
17
- The gateway accepts both the legacy `state` request and the new mutually exclusive `states` request. Batch support is negotiated by authenticated worker heartbeat (`context_batch_v1`); `/api/status` exposes `context_batch`. A legacy worker continues serving single-context requests and is never given a batch it cannot handle.
18
 
19
  Deploy a compatible gateway first. Upgrade one idle worker at a time, retaining its exact released manifest and runtime. Stop the old process explicitly and let its heartbeat expire before the replacement claims that model. Keep decoder workers under their existing shared GPU lock. Enable the new batch interface after each registered model reports the capability. A model mismatch, expired lease or missing result is an error, never a trigger for fallback inference or automatic replay.
20
 
21
- User request bodies remain bounded at 256 KiB. Worker result messages have an 8 MiB bound for repeated labels and legends; other internal messages retain their prior bound. Credentials remain exclusively in the existing private worker secret mount and Space secret. No model weights are shipped in this Space.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
  # Gateway and resident workers
2
 
3
+ The Docker Space is a CPU-only gateway: no weights or tokenization. One gateway process maintains up to six independent bounded queues, selected by an explicit model ID and pinned manifest. Keep one replica because queue state is in memory.
4
 
5
+ Internal queues retain short wire IDs such as `decision-nano-preview` for Kai. `/v1/models` exposes canonical Hugging Face repository IDs, actual sizes, complete-input limits, and per-model examples; `wire_id` is supplied for the Studio editor. `/api/status?model=…` reports the selected queue. Public `/v1/systemone` and `/v1/decision/batches` calls must select a canonical model explicitly.
6
 
7
  Configure `DECISION_BACKEND=pull_queue`, the private `DECISION_WORKER_TOKEN` secret, and `DECISION_MODEL_REGISTRY_V2`. Each registry item includes `id`, `label`, `version`, `manifest_sha256`, `complete_input_tokens` and optional presentation/example metadata. The old `DECISION_MODEL_REGISTRY` remains a fallback for legacy two-model deployments. The newer variable allows existing code to ignore it safely during a coordinated upgrade.
8
 
 
14
 
15
  ## Shared-question context batches
16
 
17
+ The public `/v1/systemone` endpoint accepts one `state`; `/v1/decision/batches` accepts ordered `states` with shared questions. Batch support is negotiated by authenticated worker heartbeat (`context_batch_v1`); `/api/status` exposes `context_batch`. A legacy worker continues serving single-context requests and is never given a batch it cannot handle.
18
 
19
  Deploy a compatible gateway first. Upgrade one idle worker at a time, retaining its exact released manifest and runtime. Stop the old process explicitly and let its heartbeat expire before the replacement claims that model. Keep decoder workers under their existing shared GPU lock. Enable the new batch interface after each registered model reports the capability. A model mismatch, expired lease or missing result is an error, never a trigger for fallback inference or automatic replay.
20
 
21
+ Public single-state request bodies remain bounded at 256 KiB; public batch bodies are bounded at 2 MiB, with at most 1,024 state/question decisions. Worker result messages have an 8 MiB bound for repeated labels and legends; other internal messages retain their prior bound. Credentials remain exclusively in the existing private worker secret mount and Space secret. No model weights are shipped in this Space.
22
+
23
+ ## Tetris Race
24
+
25
+ `/tetris/` is a passive SSE client for server-authoritative races. Starting a race creates two independent asynchronous runners immediately; connecting, reconnecting, or slowing the event stream does not pause either model. A disconnect starts a bounded reconnection lease, after which an abandoned race is cancelled. Only the seeded tetromino sequence is shared. The server retains separate board snapshots, request states, model actions, and traces at `/api/tetris/races/{id}/trace`. An illegal model choice ends that side and is never replaced by a heuristic choice or another model.
26
+
27
+ All endpoints and credentials are server-only Space variables or secrets. Configure the shared Decision gateway with `TETRIS_LOCAL_API_URL` pointing to its `/v1/systemone` route, or configure resident `drun` endpoints independently with `TETRIS_LOCAL_LUX_API_URL`, `TETRIS_LOCAL_NOX_API_URL`, `TETRIS_LOCAL_SOL_API_URL`, `TETRIS_LOCAL_EOS_API_URL`, `TETRIS_LOCAL_KAI_API_URL`, and `TETRIS_LOCAL_LEX_API_URL`. Optional local bearer secrets use the corresponding `_API_KEY` name or the shared `TETRIS_LOCAL_API_KEY`. A Decision response must return the exact Hugging Face canonical model ID selected by the player, such as `llm-semantic-router/Decision-1.0-Lux-9B`; a different or missing identity ends that side instead of being displayed under the wrong model name. This strict response identity is a deployment prerequisite. Legacy wire IDs such as `decision-lux` are not compatibility aliases here and deliberately fail closed during integration with an outdated Gateway.
28
+
29
+ Configure Jev Cloud with `TETRIS_JEV_API_URL`, the `TETRIS_JEV_API_KEY` secret, and optional `TETRIS_JEV_MODEL`. Configure the alternate SystemOne provider with `TETRIS_SYSTEMONE_API_URL`, the `TETRIS_SYSTEMONE_API_KEY` secret, and optional `TETRIS_SYSTEMONE_MODEL`. Existing `JEV_API_*`, `JEV_MIRROR_API_*`, and `LOCAL_API_URL` variables remain accepted while the deployment migrates. The `jev-latest` alias accepts only itself or a versioned `jev-x.y.z` response; an explicitly pinned Jev model must match exactly. `/api/tetris/config` exposes only readiness and display metadata; it never returns endpoint URLs or credentials.
30
+
31
+ Step races accept 1 through `TETRIS_MAX_STEPS` pieces. The default server limit is 120, and the configured limit can be raised to 2,000; `/api/tetris/config` supplies the active limit to the browser. `first_failure` has no user-visible step target and stops after the first game-over or model failure. Its operational guard defaults to the configured step limit and can be lowered with `TETRIS_FIRST_FAILURE_GUARD`. Each race reserves at most two provider calls per allowed step. `TETRIS_CLIENT_CALLS_PER_WINDOW` defaults to at least one full configured race (4,000 calls at the 2,000-step limit), and a lower explicit value is rejected. Retained events and traces are bounded by that race limit; completed races expire or are evicted from the in-memory cache, while quota records shed their trace references after completion.
32
+
33
+ `TETRIS_RACE_TIMEOUT_SECONDS` sets a 1–3,600 second whole-race deadline (default 900); raise it deliberately for slow long runs. `TETRIS_STREAM_GRACE_SECONDS` sets the 1–300 second reconnection lease after the last event stream disconnects (default 30); reconnecting resumes retained SSE events, and the browser can recover the latest board and terminal result from the race snapshot. An abandoned race is cancelled. `TETRIS_COMPLETED_TTL_SECONDS` controls completed-race retention from 30–3,600 seconds (default 300). Shutdown cancels and awaits every active race before closing the provider transport. A failed, timed-out, or otherwise incomplete race does not award a score winner. The speed percentage is available only when both sides complete the same step target; it compares the sums of server-measured request-to-response times for both providers. Local `inference_ms`, when returned, is retained per turn in the trace for diagnosis and never enters that comparison. Whole-run durations remain in the race summary as diagnostics.
34
+
35
+ `TETRIS_REQUEST_TIMEOUT_SECONDS` bounds each server-to-provider call to 0.1–300 seconds (default 45). The adapter sends optional bearer credentials only from server configuration, rejects redirects, ignores ambient proxy settings, bounds response bytes, and never returns provider URLs or credentials to the browser. Validate each configured endpoint with its selected canonical model and expected response identity before enabling that competitor; publishing a Space commit rebuilds the live Space.
Dockerfile CHANGED
@@ -6,7 +6,7 @@ RUN useradd --create-home --uid 1000 studio
6
  WORKDIR /app
7
  COPY requirements.txt /app/requirements.txt
8
  RUN pip install -r /app/requirements.txt
9
- COPY --chown=studio:studio app.py systemone_api.py contract.py engine.py relay.py model_registry.py examples.json LICENSE /app/
10
  COPY --chown=studio:studio static /app/static
11
  USER studio
12
  EXPOSE 7860
 
6
  WORKDIR /app
7
  COPY requirements.txt /app/requirements.txt
8
  RUN pip install -r /app/requirements.txt
9
+ COPY --chown=studio:studio app.py systemone_api.py contract.py engine.py relay.py model_registry.py tetris_arena.py examples.json LICENSE /app/
10
  COPY --chown=studio:studio static /app/static
11
  USER studio
12
  EXPOSE 7860
README.md CHANGED
@@ -17,6 +17,12 @@ license: mit
17
 
18
  [Decision collection](https://huggingface.co/collections/llm-semantic-router/decision-10-6ab12177bd0002394d8409f9)
19
 
 
 
 
 
 
 
20
  | Model | Focus | Model size | Complete input per question |
21
  |---|---|---:|---:|
22
  | [Kai](https://huggingface.co/llm-semantic-router/Decision-1.0-Kai-0.6B) | Multilingual encoder | 0.6B | 1K |
 
17
 
18
  [Decision collection](https://huggingface.co/collections/llm-semantic-router/decision-10-6ab12177bd0002394d8409f9)
19
 
20
+ The Space also includes a server-authoritative [Tetris Race](https://decision.vllm-sr.ai/tetris/) playground. It can compare any two configured Decision models, Jev Cloud, or a SystemOne upstream without exposing provider credentials to the browser. Both runners receive the same seeded piece sequence but keep independent boards, requests, actions, timing, and traces.
21
+
22
+ The race length follows the server's configured limit (120 steps by default, up to 2,000 when enabled). A provider failure or incomplete run has no score winner. For equal completed step races, the speed percentage compares server-measured request-to-response time on both sides; a provider's local inference time is only diagnostic.
23
+
24
+ Tetris local endpoints must run the strict Decision API and echo the selected Hugging Face repository ID exactly (for example, `llm-semantic-router/Decision-1.0-Lux-9B`) in the response `model` field. Legacy wire aliases such as `decision-lux` are intentionally rejected; this makes an outdated or misrouted Gateway fail closed instead of attributing one model's play to another.
25
+
26
  | Model | Focus | Model size | Complete input per question |
27
  |---|---|---:|---:|
28
  | [Kai](https://huggingface.co/llm-semantic-router/Decision-1.0-Kai-0.6B) | Multilingual encoder | 0.6B | 1K |
SYSTEMONE_API.md CHANGED
@@ -4,6 +4,7 @@ Call Lux, Nox, Sol, Eos, Kai, and Lex from your computer using the same `state /
4
 
5
  **Base URL:** `https://decision.vllm-sr.ai`
6
  **Inference:** `POST /v1/systemone`
 
7
  **Models:** `GET /v1/models`
8
 
9
  This is a public demo endpoint backed by the released models on AMD GPUs. It requires no API key. Inference runs on the connected server; your computer sends the request. It is a shared, bounded service rather than a dedicated production deployment. Responses are real model outputs; there is no chat-completion endpoint.
@@ -13,7 +14,7 @@ This is a public demo endpoint backed by the released models on AMD GPUs. It req
13
  ```bash
14
  curl --fail-with-body https://decision.vllm-sr.ai/v1/systemone \
15
  -H 'Content-Type: application/json' \
16
- -d '{"model":"Lux","state":"My subscription was charged twice. Please refund the duplicate charge.","questions":{"billing":{"type":"noul","instructions":"Does the customer report a billing problem?"},"team":{"type":"choice","instructions":"Which team should handle this message?","criteria":{"Billing":"Charges, invoices, and refunds","Accounts":"Login, passwords, and account access"}}}}'
17
  ```
18
 
19
  Ask multiple named questions in one call. The same state is used for every question and the worker processes the admitted decisions in GPU batches. This does not imply shared-state encoder caching.
@@ -32,7 +33,7 @@ from typesafe_sdk import Choice, Noul, TypeSafeClient, RetryPolicy
32
  with TypeSafeClient(
33
  base_url="https://decision.vllm-sr.ai",
34
  api_key="decision-public", # SDK-required placeholder, not a credential
35
- model="Lux",
36
  timeout=60,
37
  retry=RetryPolicy(max_retries=0),
38
  ) as client:
@@ -52,7 +53,7 @@ with TypeSafeClient(
52
  print(result.choices["team"].choice)
53
  ```
54
 
55
- Change `model` to `Nox`, `Sol`, `Eos`, `Kai`, or `Lex`. Short names are case-insensitive; full Hugging Face repository names and existing wire IDs also work. The returned `model` identifies the actual worker. SDK calls explicitly select a model; raw HTTP requests that omit it use Lux. A standalone single-model Kai deployment still defaults to Kai.
56
 
57
  | Model | Parameters | Complete input per question |
58
  |---|---:|---:|
@@ -68,10 +69,10 @@ Change `model` to `Nox`, `Sol`, `Eos`, `Kai`, or `Lex`. Short names are case-ins
68
  ## Responses and limits
69
 
70
  - Noul returns P(true); Choice returns its winning label and probabilities; Score returns the expected ordinal level and its distribution.
71
- - Choice/Score `confidence` is **max(probabilities)**, the native peak probability. It is not a calibrated probability of correctness, and no equivalence to TypeSafe's proprietary confidence calculation is claimed. Raw distributions and predictions are unchanged. The response describes this definition in `profile`.
72
  - The hosted endpoint accepts one or more named questions without a fixed question-count cap. GPU execution uses microbatches of up to eight decisions. Choice accepts 2–255 options, Score accepts 2–10 ordered levels, and JSON requests are limited to 256 KiB. Expanded state/question input is limited to 16 MiB so large Cartesian workloads stay bounded. Provide explicit nonempty instructions. State plus question plus all candidates must fit the model limit; overflowing requests are rejected without truncation.
73
  - Synchronous requests have a 45-second server wait. Each model admits up to eight queued/running requests. A busy or timed-out service may return 429/503/504. Completed requests do not consume admission capacity. Do not treat transport errors as predictions. SDK example retries are disabled to avoid duplicate inference.
74
- - Studio additionally supports `states: [{id, state}, ...]`, with shared questions, without fixed context or decision-count caps; the same byte and complete-input budgets apply. This is a documented Decision extension; it is not the standard SDK response shape. With the standard SDK, call `system_one` separately for each state.
75
 
76
  See [SYSTEM_ONE_MAPPING.md](SYSTEM_ONE_MAPPING.md) for the complete adapter contract. The standard single-state call, all three answer types, and model discovery are validated with the official SDK.
77
 
 
4
 
5
  **Base URL:** `https://decision.vllm-sr.ai`
6
  **Inference:** `POST /v1/systemone`
7
+ **Shared-question batches:** `POST /v1/decision/batches`
8
  **Models:** `GET /v1/models`
9
 
10
  This is a public demo endpoint backed by the released models on AMD GPUs. It requires no API key. Inference runs on the connected server; your computer sends the request. It is a shared, bounded service rather than a dedicated production deployment. Responses are real model outputs; there is no chat-completion endpoint.
 
14
  ```bash
15
  curl --fail-with-body https://decision.vllm-sr.ai/v1/systemone \
16
  -H 'Content-Type: application/json' \
17
+ -d '{"model":"llm-semantic-router/Decision-1.0-Lux-9B","state":"My subscription was charged twice. Please refund the duplicate charge.","questions":{"billing":{"type":"noul","instructions":"Does the customer report a billing problem?"},"team":{"type":"choice","instructions":"Which team should handle this message?","criteria":{"Billing":"Charges, invoices, and refunds","Accounts":"Login, passwords, and account access"}}}}'
18
  ```
19
 
20
  Ask multiple named questions in one call. The same state is used for every question and the worker processes the admitted decisions in GPU batches. This does not imply shared-state encoder caching.
 
33
  with TypeSafeClient(
34
  base_url="https://decision.vllm-sr.ai",
35
  api_key="decision-public", # SDK-required placeholder, not a credential
36
+ model="llm-semantic-router/Decision-1.0-Lux-9B",
37
  timeout=60,
38
  retry=RetryPolicy(max_retries=0),
39
  ) as client:
 
53
  print(result.choices["team"].choice)
54
  ```
55
 
56
+ Select any exact Hugging Face repository ID returned as `id` by `/v1/models`. The model is required on public inference requests; short names, wire aliases, and case variants are rejected. The response repeats the selected canonical ID. The Studio editor uses separate same-origin endpoints for its internal queue and diagnostics.
57
 
58
  | Model | Parameters | Complete input per question |
59
  |---|---:|---:|
 
69
  ## Responses and limits
70
 
71
  - Noul returns P(true); Choice returns its winning label and probabilities; Score returns the expected ordinal level and its distribution.
72
+ - Choice/Score `confidence` is **max(probabilities)**, the native peak probability. It is not a calibrated probability of correctness, and no equivalence to TypeSafe's proprietary confidence calculation is claimed. Raw distributions and predictions are unchanged. Public responses contain only `model`, `answers`, and `usage`.
73
  - The hosted endpoint accepts one or more named questions without a fixed question-count cap. GPU execution uses microbatches of up to eight decisions. Choice accepts 2–255 options, Score accepts 2–10 ordered levels, and JSON requests are limited to 256 KiB. Expanded state/question input is limited to 16 MiB so large Cartesian workloads stay bounded. Provide explicit nonempty instructions. State plus question plus all candidates must fit the model limit; overflowing requests are rejected without truncation.
74
  - Synchronous requests have a 45-second server wait. Each model admits up to eight queued/running requests. A busy or timed-out service may return 429/503/504. Completed requests do not consume admission capacity. Do not treat transport errors as predictions. SDK example retries are disabled to avoid duplicate inference.
75
+ - `POST /v1/decision/batches` applies one shared `questions` map to ordered `states: [{id, state}, ...]`. It requires the same explicit canonical `model`, unique state IDs, at most 1,024 states, questions, and total decisions, a 2 MiB JSON body, and the same 16 MiB expanded-input and per-question token budgets. Its atomic response contains `model`, ordered `results: [{id, answers, usage}, ...]`, and aggregate `usage`. This Decision extension is separate from the SDK's single-state method. There is no `/v1/systemone/batch` alias.
76
 
77
  See [SYSTEM_ONE_MAPPING.md](SYSTEM_ONE_MAPPING.md) for the complete adapter contract. The standard single-state call, all three answer types, and model discovery are validated with the official SDK.
78
 
SYSTEM_ONE_MAPPING.md CHANGED
@@ -7,18 +7,18 @@ Studio exposes a **SystemOne-format Decision service** with official Python SDK
7
  | Input | Exact local mapping |
8
  |---|---|
9
  | `state` text/object/array | Native `state_text`; structured content serializes as UTF-8 JSON with sorted object keys and compact separators. No truncation. |
10
- | `states: [{id, state}, ...]` | Decision Studio extension for shared-question context batches; mutually exclusive with `state`. IDs are unique output keys and never added to model text. |
11
  | Question map key | Native question ID and returned answer key. Never added to model text. |
12
  | `instructions` text/object/array | Native question text, with the same deterministic JSON rule. |
13
  | Choice `criteria` map | Each option name is semantic: null description → name; otherwise name + `: ` + full description. Object insertion order is retained. External native candidate IDs equal the names, but the names reach the model only through this explicit semantic text. |
14
  | Score `criteria` array | Ordered native levels with IDs and values 0…K−1. Returned score is the probability-weighted level index. |
15
  | Noul `criteria.false/true` | Native false/true descriptions. If absent, the original native default no/yes descriptions are used. |
16
 
17
- The adapter has no fixed question/context-count caps. It accepts 2–255 Choice options, 2–10 Score levels, a 256 KiB JSON request, and up to 16 MiB of expanded context/question input. Records execute in GPU microbatches of up to eight. The complete-input limit is 1,024 tokens for Kai/Lex and 16,384 for Eos/Sol/Nox/Lux **per question**, including state and all candidate descriptions. The selected model’s published loader and numerical engine are retained. Unsupported fields and models are rejected. This is an intentional bounded subset of the upstream interface.
18
 
19
- Single-context responses preserve named `answers`, Choice `choice/probabilities`, Score `score/legend/probabilities`, and Noul `noul`. `usage.output_tokens=0` because these models do not generate text. Studio adds actual timing/profile/source fields.
20
 
21
- The synchronous `/v1/systemone` response supplies the SDK-required `confidence` as **max(probabilities)**, matching the native peak-probability convention. This is explicitly recorded in `profile.confidence_definition`; it is not calibrated correctness or a claim of equivalence to TypeSafe’s unspecified confidence formula. `top_probability` remains available and the original distribution is unchanged. Noul returns its original probability of yes. Internal worker and asynchronous Studio envelopes retain their existing fields.
22
 
23
  This mapping is independently testable with small metadata fixtures. It does not change the model's packing, native graph, weights, scoring arithmetic, or export format.
24
 
@@ -28,7 +28,7 @@ Sol and Nox use their published `question_row` and `predict_rows` path, the same
28
 
29
  ```json
30
  {
31
- "model": "decision-nano-preview",
32
  "states": [
33
  {"id": "T-101", "state": "My subscription was charged twice. Please refund the duplicate payment."},
34
  {"id": "T-102", "state": "I forgot my password and need help signing in."}
@@ -46,6 +46,6 @@ Sol and Nox use their published `question_row` and `predict_rows` path, the same
46
  }
47
  ```
48
 
49
- Batch responses replace top-level `answers` with ordered `results: [{id, answers, usage}, ...]`. Context order and shared-question order are retained. Top-level `usage` sums all pairs; `timing` also reports `contexts`, `questions` and `decisions`. A row can expand to show the original probabilities. Score is an expected ordinal index; Noul is P(Yes), not an independent confidence estimate.
50
 
51
  The worker admits every complete context/question input before its first forward. Any overflow rejects the whole request: no truncation or partial result. Admitted pairs are flattened into one predictor call with physical batches of up to eight. This is actual GPU batching, not browser-side request fanout or shared-state caching. The native numerical engine, weights, candidate order and published temperatures remain unchanged.
 
7
  | Input | Exact local mapping |
8
  |---|---|
9
  | `state` text/object/array | Native `state_text`; structured content serializes as UTF-8 JSON with sorted object keys and compact separators. No truncation. |
10
+ | `states: [{id, state}, ...]` | `POST /v1/decision/batches` extension for shared-question context batches. IDs are unique output keys and never added to model text. |
11
  | Question map key | Native question ID and returned answer key. Never added to model text. |
12
  | `instructions` text/object/array | Native question text, with the same deterministic JSON rule. |
13
  | Choice `criteria` map | Each option name is semantic: null description → name; otherwise name + `: ` + full description. Object insertion order is retained. External native candidate IDs equal the names, but the names reach the model only through this explicit semantic text. |
14
  | Score `criteria` array | Ordered native levels with IDs and values 0…K−1. Returned score is the probability-weighted level index. |
15
  | Noul `criteria.false/true` | Native false/true descriptions. If absent, the original native default no/yes descriptions are used. |
16
 
17
+ The public single-state endpoint requires exactly `model`, `state`, and `questions`, with a 256 KiB JSON body. The separate batch endpoint requires exactly `model`, `states`, and `questions`, with a 2 MiB JSON body and at most 1,024 states, questions, and total decisions. Both accept 2–255 Choice options, 2–10 Score levels, and up to 16 MiB of expanded context/question input. Records execute in GPU microbatches of up to eight. The complete-input limit is 1,024 tokens for Kai/Lex and 16,384 for Eos/Sol/Nox/Lux **per question**, including state and all candidate descriptions. The selected model’s published loader and numerical engine are retained. Unsupported fields and models are rejected.
18
 
19
+ Single-context responses preserve named `answers`, Choice `choice/probabilities`, Score `score/legend/probabilities`, and Noul `noul`. `usage.output_tokens=0` because these models do not generate text. The public inference envelopes contain only `model`, answers or results, and usage. Studio's internal endpoints additionally retain timing/profile/source fields.
20
 
21
+ The public inference responses supply the SDK-required `confidence` as **max(probabilities)**, matching the native peak-probability convention. It is not calibrated correctness or a claim of equivalence to TypeSafe’s unspecified confidence formula. Noul returns its original probability of yes. Internal worker and Studio envelopes retain diagnostic fields, including `top_probability` and timing.
22
 
23
  This mapping is independently testable with small metadata fixtures. It does not change the model's packing, native graph, weights, scoring arithmetic, or export format.
24
 
 
28
 
29
  ```json
30
  {
31
+ "model": "llm-semantic-router/Decision-1.0-Kai-0.6B",
32
  "states": [
33
  {"id": "T-101", "state": "My subscription was charged twice. Please refund the duplicate payment."},
34
  {"id": "T-102", "state": "I forgot my password and need help signing in."}
 
46
  }
47
  ```
48
 
49
+ Submit this payload to `POST /v1/decision/batches`. Batch responses replace top-level `answers` with ordered `results: [{id, answers, usage}, ...]`. Context order and shared-question order are retained. Top-level `usage` sums all pairs. Studio's internal response also reports `contexts`, `questions`, and `decisions` in `timing`. A row can expand to show the original probabilities. Score is an expected ordinal index; Noul is P(Yes), not an independent confidence estimate.
50
 
51
  The worker admits every complete context/question input before its first forward. Any overflow rejects the whole request: no truncation or partial result. Admitted pairs are flattened into one predictor call with physical batches of up to eight. This is actual GPU batching, not browser-side request fanout or shared-state caching. The native numerical engine, weights, candidate order and published temperatures remain unchanged.
app.py CHANGED
@@ -1,23 +1,76 @@
1
  """Same-origin Studio: local native mode or an outbound-worker CPU gateway."""
2
  import asyncio
 
3
  import json
4
  import logging
5
  import os
6
  import time
 
7
  from pathlib import Path
 
8
  from fastapi import FastAPI, HTTPException, Request
9
- from fastapi.responses import JSONResponse
10
  from fastapi.staticfiles import StaticFiles
11
  from starlette.concurrency import run_in_threadpool
 
12
  from contract import MODEL, to_records
13
- from engine import DEFAULT_MANIFEST, Engine, Unavailable, Busy
14
  from relay import Relay, RelayError
15
- from systemone_api import arrange_models, public_models, resolve_model, sdk_response
 
 
 
 
 
 
 
 
 
 
 
 
 
 
16
 
17
  ROOT = Path(__file__).resolve().parent
18
  logger = logging.getLogger("decision.studio")
19
 
20
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
21
  def unique_object(pairs):
22
  result = {}
23
  for key, value in pairs:
@@ -51,7 +104,16 @@ async def read_input(request):
51
  return payload, records
52
 
53
 
54
- def create_app(mode=None, *, native_engine=None, relay=None, registry=None, relays=None):
 
 
 
 
 
 
 
 
 
55
  mode = mode or os.getenv("DECISION_BACKEND", "native")
56
  if mode not in {"native", "pull_queue"}:
57
  raise ValueError("DECISION_BACKEND must be native or pull_queue")
@@ -60,7 +122,32 @@ def create_app(mode=None, *, native_engine=None, relay=None, registry=None, rela
60
  registry = [{'id': MODEL, 'label': 'Kai', 'version': '1.0', 'manifest_sha256': relay.manifest}]
61
  registry = arrange_models(model_registry(registry))
62
  default_model = next(iter(registry))
63
- api = FastAPI(title="Decision Studio", version="0.6.0", docs_url="/api/docs", redoc_url=None)
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
64
  api.state.engine = native_engine or Engine()
65
  if mode == "native" and (len(registry) != 1 or api.state.engine.model != MODEL
66
  or api.state.engine.manifest != registry[MODEL]['manifest_sha256']):
@@ -79,14 +166,27 @@ def create_app(mode=None, *, native_engine=None, relay=None, registry=None, rela
79
  api.state.relay = relays.get(MODEL) # backward-compatible default for local tooling
80
  api.state.registry = registry
81
 
 
 
82
  def select(model):
83
  if not isinstance(model, str) or model not in registry:
84
  raise HTTPException(422, "This model is not available in this Studio")
85
  return relays.get(model)
86
 
87
- async def input_for(request):
88
- payload = await read_json(request)
89
- model = resolve_model(payload.get('model', default_model) if isinstance(payload, dict) else default_model, registry)
 
 
 
 
 
 
 
 
 
 
 
90
  selected = select(model)
91
  if isinstance(payload, dict):
92
  payload = dict(payload, model=model)
@@ -94,7 +194,7 @@ def create_app(mode=None, *, native_engine=None, relay=None, registry=None, rela
94
  records = to_records(payload, model=model)
95
  except (ValueError, TypeError, UnicodeError, RecursionError) as exc:
96
  raise HTTPException(422, str(exc)) from None
97
- return payload, records, selected
98
 
99
  @api.exception_handler(RelayError)
100
  async def relay_error(request, exc):
@@ -110,14 +210,100 @@ def create_app(mode=None, *, native_engine=None, relay=None, registry=None, rela
110
  def examples():
111
  return json.loads((ROOT / "examples.json").read_text())
112
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
113
  @api.get("/v1/models")
114
  def models():
115
- return {"default": default_model, "models": public_models(registry), "limits": {"request_bytes": 256 * 1024, "expanded_input_bytes": 16 * 1024 * 1024, "questions": None, "contexts": None, "active_requests_per_model": 8, "gpu_microbatch": 8, "synchronous_wait_seconds": 45}}
116
 
117
  @api.post("/v1/systemone")
 
 
 
 
 
 
 
118
  @api.post("/api/evaluate")
119
- async def evaluate(request: Request):
120
- payload, records, selected = await input_for(request)
 
 
 
 
 
121
  if selected:
122
  job = await selected.submit(payload, keep_until_read=True)
123
  deadline = time.monotonic() + 45
@@ -126,7 +312,10 @@ def create_app(mode=None, *, native_engine=None, relay=None, registry=None, rela
126
  remaining = deadline - time.monotonic()
127
  current = await selected.wait(job["id"], max(0.0, min(0.25, remaining)))
128
  if current["status"] == "succeeded":
129
- return sdk_response(current["result"])
 
 
 
130
  if current["status"] not in Relay.ACTIVE:
131
  raise HTTPException(503, current.get("detail", "The request did not finish"))
132
  if await request.is_disconnected():
@@ -135,7 +324,11 @@ def create_app(mode=None, *, native_engine=None, relay=None, registry=None, rela
135
  finally:
136
  await selected.release(job["id"])
137
  try:
138
- return sdk_response(await run_in_threadpool(api.state.engine.evaluate, payload, records))
 
 
 
 
139
  except Unavailable as exc:
140
  raise HTTPException(503, str(exc)) from None
141
  except Busy as exc:
@@ -152,7 +345,7 @@ def create_app(mode=None, *, native_engine=None, relay=None, registry=None, rela
152
  if relays:
153
  @api.post("/api/jobs", status_code=202)
154
  async def submit(request: Request):
155
- payload, _, selected = await input_for(request)
156
  return await selected.submit(payload)
157
 
158
  @api.get("/api/jobs/{job_id}")
 
1
  """Same-origin Studio: local native mode or an outbound-worker CPU gateway."""
2
  import asyncio
3
+ import ipaddress
4
  import json
5
  import logging
6
  import os
7
  import time
8
+ from contextlib import asynccontextmanager
9
  from pathlib import Path
10
+
11
  from fastapi import FastAPI, HTTPException, Request
12
+ from fastapi.responses import JSONResponse, StreamingResponse
13
  from fastapi.staticfiles import StaticFiles
14
  from starlette.concurrency import run_in_threadpool
15
+
16
  from contract import MODEL, to_records
17
+ from engine import Busy, Engine, Unavailable
18
  from relay import Relay, RelayError
19
+ from systemone_api import (
20
+ arrange_models,
21
+ public_models,
22
+ resolve_canonical_model,
23
+ resolve_model,
24
+ public_response,
25
+ sdk_response,
26
+ )
27
+ from tetris_arena import (
28
+ ArenaQuotaError,
29
+ ArenaValidationError,
30
+ HTTPDecisionAdapter,
31
+ RaceManager,
32
+ encode_sse,
33
+ )
34
 
35
  ROOT = Path(__file__).resolve().parent
36
  logger = logging.getLogger("decision.studio")
37
 
38
 
39
+ def trusted_proxy_networks(value: str):
40
+ if not value.strip():
41
+ return ()
42
+ try:
43
+ return tuple(
44
+ ipaddress.ip_network(part.strip(), strict=False)
45
+ for part in value.split(",")
46
+ )
47
+ except ValueError as exc:
48
+ raise ValueError("TETRIS_TRUSTED_PROXY_CIDRS must contain IP CIDRs") from exc
49
+
50
+
51
+ def client_identity(request: Request, trusted_proxies=()) -> str:
52
+ """Use X-Forwarded-For only when every traversed hop is trusted."""
53
+ peer = request.client.host if request.client else ""
54
+ try:
55
+ current = ipaddress.ip_address(peer)
56
+ except ValueError:
57
+ return "unknown"
58
+ header = request.headers.get("x-forwarded-for", "")
59
+ if not header or len(header) > 1024 or not trusted_proxies:
60
+ return str(current)
61
+ try:
62
+ hops = [ipaddress.ip_address(part.strip()) for part in header.split(",")]
63
+ except ValueError:
64
+ return str(current)
65
+ if not 1 <= len(hops) <= 16:
66
+ return str(current)
67
+ for hop in reversed(hops):
68
+ if not any(current in network for network in trusted_proxies):
69
+ break
70
+ current = hop
71
+ return str(current)
72
+
73
+
74
  def unique_object(pairs):
75
  result = {}
76
  for key, value in pairs:
 
104
  return payload, records
105
 
106
 
107
+ def create_app(
108
+ mode=None,
109
+ *,
110
+ native_engine=None,
111
+ relay=None,
112
+ registry=None,
113
+ relays=None,
114
+ tetris_adapter=None,
115
+ tetris_manager=None,
116
+ ):
117
  mode = mode or os.getenv("DECISION_BACKEND", "native")
118
  if mode not in {"native", "pull_queue"}:
119
  raise ValueError("DECISION_BACKEND must be native or pull_queue")
 
122
  registry = [{'id': MODEL, 'label': 'Kai', 'version': '1.0', 'manifest_sha256': relay.manifest}]
123
  registry = arrange_models(model_registry(registry))
124
  default_model = next(iter(registry))
125
+ trusted_proxies = trusted_proxy_networks(
126
+ os.getenv("TETRIS_TRUSTED_PROXY_CIDRS", "")
127
+ )
128
+ owns_tetris_manager = tetris_manager is None
129
+ owns_tetris_adapter = owns_tetris_manager and tetris_adapter is None
130
+ if tetris_manager is None:
131
+ tetris_adapter = tetris_adapter or HTTPDecisionAdapter.from_environment()
132
+ tetris_manager = RaceManager.from_environment(tetris_adapter)
133
+
134
+ @asynccontextmanager
135
+ async def lifespan(_api):
136
+ yield
137
+ if owns_tetris_manager:
138
+ try:
139
+ await tetris_manager.aclose()
140
+ finally:
141
+ if owns_tetris_adapter:
142
+ await tetris_adapter.close()
143
+
144
+ api = FastAPI(
145
+ title="Decision Studio",
146
+ version="0.6.0",
147
+ docs_url="/api/docs",
148
+ redoc_url=None,
149
+ lifespan=lifespan,
150
+ )
151
  api.state.engine = native_engine or Engine()
152
  if mode == "native" and (len(registry) != 1 or api.state.engine.model != MODEL
153
  or api.state.engine.manifest != registry[MODEL]['manifest_sha256']):
 
166
  api.state.relay = relays.get(MODEL) # backward-compatible default for local tooling
167
  api.state.registry = registry
168
 
169
+ api.state.tetris = tetris_manager
170
+
171
  def select(model):
172
  if not isinstance(model, str) or model not in registry:
173
  raise HTTPException(422, "This model is not available in this Studio")
174
  return relays.get(model)
175
 
176
+ async def input_for(request, *, public_contract=None):
177
+ payload = await read_json(
178
+ request, limit=2 * 1024 * 1024 if public_contract == 'batch' else 256 * 1024
179
+ )
180
+ if public_contract is not None:
181
+ expected = {'model', 'questions', 'states' if public_contract == 'batch' else 'state'}
182
+ if not isinstance(payload, dict) or set(payload) != expected:
183
+ raise HTTPException(422, f"Provide exactly {', '.join(sorted(expected))}.")
184
+ requested_model = payload.get('model') if isinstance(payload, dict) else None
185
+ model = (
186
+ resolve_canonical_model(requested_model, registry)
187
+ if public_contract is not None
188
+ else resolve_model(requested_model or default_model, registry)
189
+ )
190
  selected = select(model)
191
  if isinstance(payload, dict):
192
  payload = dict(payload, model=model)
 
194
  records = to_records(payload, model=model)
195
  except (ValueError, TypeError, UnicodeError, RecursionError) as exc:
196
  raise HTTPException(422, str(exc)) from None
197
+ return payload, records, selected, registry[model]['repo_id']
198
 
199
  @api.exception_handler(RelayError)
200
  async def relay_error(request, exc):
 
210
  def examples():
211
  return json.loads((ROOT / "examples.json").read_text())
212
 
213
+ @api.get("/api/tetris/config")
214
+ def tetris_config():
215
+ # Credentials and upstream URLs intentionally never enter this response.
216
+ return api.state.tetris.public_config()
217
+
218
+ @api.post("/api/tetris/races", status_code=201)
219
+ async def create_tetris_race(request: Request):
220
+ payload = await read_json(request)
221
+ try:
222
+ race = await api.state.tetris.create(
223
+ payload,
224
+ client_id=client_identity(request, trusted_proxies),
225
+ )
226
+ except ArenaValidationError as exc:
227
+ headers = (
228
+ {"Retry-After": str(exc.retry_after)}
229
+ if isinstance(exc, ArenaQuotaError)
230
+ else None
231
+ )
232
+ raise HTTPException(exc.status_code, str(exc), headers=headers) from None
233
+ return {
234
+ "id": race.id,
235
+ "status": race.status,
236
+ "events_url": f"/api/tetris/races/{race.id}/events",
237
+ "result_url": f"/api/tetris/races/{race.id}",
238
+ }
239
+
240
+ def get_tetris_race(race_id):
241
+ try:
242
+ return api.state.tetris.get(race_id)
243
+ except KeyError:
244
+ raise HTTPException(404, "Race not found") from None
245
+
246
+ @api.get("/api/tetris/races/{race_id}")
247
+ async def tetris_race(race_id: str):
248
+ return get_tetris_race(race_id).snapshot()
249
+
250
+ @api.get("/api/tetris/races/{race_id}/trace")
251
+ async def tetris_race_trace(race_id: str):
252
+ return get_tetris_race(race_id).snapshot(include_traces=True)
253
+
254
+ @api.get("/api/tetris/races/{race_id}/events")
255
+ async def tetris_race_events(race_id: str, request: Request, after: int = 0):
256
+ race = get_tetris_race(race_id)
257
+ header_cursor = request.headers.get("last-event-id")
258
+ if header_cursor:
259
+ try:
260
+ after = max(after, int(header_cursor))
261
+ except ValueError:
262
+ raise HTTPException(400, "Last-Event-ID must be an integer") from None
263
+ if after < 0:
264
+ raise HTTPException(400, "after must not be negative")
265
+
266
+ async def stream():
267
+ async for event in race.iter_events(after):
268
+ yield encode_sse(event)
269
+
270
+ return StreamingResponse(
271
+ stream(),
272
+ media_type="text/event-stream",
273
+ headers={
274
+ "Cache-Control": "no-store",
275
+ "Connection": "keep-alive",
276
+ "X-Accel-Buffering": "no",
277
+ },
278
+ )
279
+
280
+ @api.delete("/api/tetris/races/{race_id}", status_code=204)
281
+ async def cancel_tetris_race(race_id: str):
282
+ try:
283
+ await api.state.tetris.cancel(race_id)
284
+ except KeyError:
285
+ raise HTTPException(404, "Race not found") from None
286
+
287
  @api.get("/v1/models")
288
  def models():
289
+ return {"default": registry[default_model]['repo_id'], "models": public_models(registry), "limits": {"request_bytes": 256 * 1024, "expanded_input_bytes": 16 * 1024 * 1024, "questions": None, "contexts": None, "active_requests_per_model": 8, "gpu_microbatch": 8, "synchronous_wait_seconds": 45}}
290
 
291
  @api.post("/v1/systemone")
292
+ async def evaluate_systemone(request: Request):
293
+ return await evaluate(request, public_contract='single')
294
+
295
+ @api.post("/v1/decision/batches")
296
+ async def evaluate_decision_batches(request: Request):
297
+ return await evaluate(request, public_contract='batch')
298
+
299
  @api.post("/api/evaluate")
300
+ async def evaluate_studio(request: Request):
301
+ return await evaluate(request, public_contract=None)
302
+
303
+ async def evaluate(request: Request, *, public_contract: str | None):
304
+ payload, records, selected, public_model = await input_for(
305
+ request, public_contract=public_contract
306
+ )
307
  if selected:
308
  job = await selected.submit(payload, keep_until_read=True)
309
  deadline = time.monotonic() + 45
 
312
  remaining = deadline - time.monotonic()
313
  current = await selected.wait(job["id"], max(0.0, min(0.25, remaining)))
314
  if current["status"] == "succeeded":
315
+ if public_contract is not None:
316
+ return public_response(current['result'], public_model=public_model,
317
+ batch=public_contract == 'batch')
318
+ return sdk_response(current['result'])
319
  if current["status"] not in Relay.ACTIVE:
320
  raise HTTPException(503, current.get("detail", "The request did not finish"))
321
  if await request.is_disconnected():
 
324
  finally:
325
  await selected.release(job["id"])
326
  try:
327
+ result = await run_in_threadpool(api.state.engine.evaluate, payload, records)
328
+ if public_contract is not None:
329
+ return public_response(result, public_model=public_model,
330
+ batch=public_contract == 'batch')
331
+ return sdk_response(result)
332
  except Unavailable as exc:
333
  raise HTTPException(503, str(exc)) from None
334
  except Busy as exc:
 
345
  if relays:
346
  @api.post("/api/jobs", status_code=202)
347
  async def submit(request: Request):
348
+ payload, _, selected, _ = await input_for(request)
349
  return await selected.submit(payload)
350
 
351
  @api.get("/api/jobs/{job_id}")
contract.py CHANGED
@@ -5,6 +5,8 @@ import math
5
  MODEL = "decision-nano-preview"
6
  MAX_EXPANDED_INPUT_BYTES = 16 * 1024 * 1024
7
  MAX_REQUEST_BYTES = 256 * 1024
 
 
8
 
9
 
10
  def content(value, label):
@@ -107,8 +109,9 @@ def contexts(body, *, model=MODEL):
107
  raise ValueError("Provide exactly one of state or states")
108
  if body.get("model", MODEL) != model:
109
  raise ValueError("Request model does not match the selected model")
110
- if len(json.dumps(body, ensure_ascii=False, separators=(",", ":"), allow_nan=False).encode("utf-8")) > MAX_REQUEST_BYTES:
111
- raise ValueError("Request exceeds 256 KiB; no field is truncated")
 
112
  questions = body.get("questions")
113
  if not isinstance(questions, dict) or not questions:
114
  raise ValueError("Provide at least one named question")
@@ -119,6 +122,8 @@ def contexts(body, *, model=MODEL):
119
  states = body["states"]
120
  if not isinstance(states, list) or not states:
121
  raise ValueError("Provide at least one context")
 
 
122
  result, seen = [], set()
123
  for item in states:
124
  if not isinstance(item, dict) or set(item) != {"id", "state"}:
 
5
  MODEL = "decision-nano-preview"
6
  MAX_EXPANDED_INPUT_BYTES = 16 * 1024 * 1024
7
  MAX_REQUEST_BYTES = 256 * 1024
8
+ MAX_BATCH_REQUEST_BYTES = 2 * 1024 * 1024
9
+ MAX_BATCH_DECISIONS = 1024
10
 
11
 
12
  def content(value, label):
 
109
  raise ValueError("Provide exactly one of state or states")
110
  if body.get("model", MODEL) != model:
111
  raise ValueError("Request model does not match the selected model")
112
+ limit = MAX_BATCH_REQUEST_BYTES if "states" in body else MAX_REQUEST_BYTES
113
+ if len(json.dumps(body, ensure_ascii=False, separators=(",", ":"), allow_nan=False).encode("utf-8")) > limit:
114
+ raise ValueError("Request exceeds the endpoint byte limit; no field is truncated")
115
  questions = body.get("questions")
116
  if not isinstance(questions, dict) or not questions:
117
  raise ValueError("Provide at least one named question")
 
122
  states = body["states"]
123
  if not isinstance(states, list) or not states:
124
  raise ValueError("Provide at least one context")
125
+ if len(states) > MAX_BATCH_DECISIONS or len(questions) > MAX_BATCH_DECISIONS or len(states) * len(questions) > MAX_BATCH_DECISIONS:
126
+ raise ValueError("A batch may contain at most 1024 states, questions, and decisions")
127
  result, seen = [], set()
128
  for item in states:
129
  if not isinstance(item, dict) or set(item) != {"id", "state"}:
model_registry.py CHANGED
@@ -53,7 +53,13 @@ def model_registry(value=None):
53
  or any(not isinstance(x, str) or not 1 <= len(x) <= 128 for x in ids)
54
  or len(set(ids)) != len(ids) or item['default_example_id'] not in ids):
55
  raise ValueError('Invalid example allowlist/default')
56
- result[item['id']] = dict(item, complete_input_tokens=cap)
 
 
 
 
 
 
57
  if MODEL not in result:
58
  raise ValueError('Retain the existing Kai wire alias as the default')
59
  return result
 
53
  or any(not isinstance(x, str) or not 1 <= len(x) <= 128 for x in ids)
54
  or len(set(ids)) != len(ids) or item['default_example_id'] not in ids):
55
  raise ValueError('Invalid example allowlist/default')
56
+ # The queue uses the short ID internally, but public SystemOne traffic
57
+ # is identified exclusively by the immutable Hub repository ID.
58
+ result[item['id']] = dict(
59
+ item,
60
+ repo_id=profile['repo_id'],
61
+ complete_input_tokens=cap,
62
+ )
63
  if MODEL not in result:
64
  raise ValueError('Retain the existing Kai wire alias as the default')
65
  return result
requirements.txt CHANGED
@@ -1,2 +1,3 @@
1
  fastapi==0.115.12
 
2
  uvicorn==0.34.2
 
1
  fastapi==0.115.12
2
+ httpx==0.28.1
3
  uvicorn==0.34.2
static/app.js CHANGED
@@ -109,7 +109,9 @@ async function refreshMenuStatuses() {
109
  }
110
  async function loadModels() {
111
  const result=validateRegistry(await requestJSON('/v1/models',{cache:'no-store'}));
112
- models=result.models;selectedModel=result.default;
 
 
113
  modelMenu.identity(modelInfo(),{enabled:false});
114
  }
115
  function saveWorkspace() {
@@ -249,7 +251,7 @@ async function requestJSON(url, options = {}) {
249
  return body;
250
  }
251
  async function nativePrediction(payload, current) {
252
- return requestJSON('/v1/systemone', {method:'POST', headers:{'Content-Type':'application/json'}, body:JSON.stringify(payload), signal:current.signal});
253
  }
254
  function pollPause(signal) {
255
  return new Promise((resolve, reject) => {
 
109
  }
110
  async function loadModels() {
111
  const result=validateRegistry(await requestJSON('/v1/models',{cache:'no-store'}));
112
+ const selected=result.models.find(model=>model.id===result.default);
113
+ models=result.models.map(model=>({...model,id:model.wire_id,repo_id:model.id}));
114
+ selectedModel=selected.wire_id;
115
  modelMenu.identity(modelInfo(),{enabled:false});
116
  }
117
  function saveWorkspace() {
 
251
  return body;
252
  }
253
  async function nativePrediction(payload, current) {
254
+ return requestJSON('/api/evaluate', {method:'POST', headers:{'Content-Type':'application/json'}, body:JSON.stringify(payload), signal:current.signal});
255
  }
256
  function pollPause(signal) {
257
  return new Promise((resolve, reject) => {
static/model-menu.js CHANGED
@@ -3,10 +3,10 @@ const esc = value => String(value).replace(/[&<>"']/g,x=>({'&':'&amp;','<':'&lt;
3
  export function validateRegistry(result) {
4
  if(!Array.isArray(result?.models)||!result.models.length||!result.models.some(m=>m.id===result.default))throw Error('Invalid model registry.');
5
  const text=(v,max)=>typeof v==='string'&&v.trim()&&v.length<=max;
6
- const ids=new Set();
7
  for(const m of result.models){
8
- if(!text(m.id,128)||ids.has(m.id)||!text(m.label,64)||!text(m.version,32)||!/^[0-9a-f]{64}$/.test(m.manifest_sha256)||!Number.isInteger(m.complete_input_tokens)||m.complete_input_tokens<1)throw Error('Invalid model identity.');
9
- ids.add(m.id);
10
  if(m.description!==undefined&&!text(m.description,240)||m.parameter_label!==undefined&&!text(m.parameter_label,32))throw Error('Invalid model description.');
11
  if(m.example_ids!==undefined&&(!Array.isArray(m.example_ids)||!m.example_ids.length||m.example_ids.some(x=>!text(x,128))||new Set(m.example_ids).size!==m.example_ids.length))throw Error('Invalid model examples.');
12
  if(m.default_example_id!==undefined&&!text(m.default_example_id,128))throw Error('Invalid default example.');
 
3
  export function validateRegistry(result) {
4
  if(!Array.isArray(result?.models)||!result.models.length||!result.models.some(m=>m.id===result.default))throw Error('Invalid model registry.');
5
  const text=(v,max)=>typeof v==='string'&&v.trim()&&v.length<=max;
6
+ const ids=new Set(),wireIds=new Set();
7
  for(const m of result.models){
8
+ if(!text(m.id,128)||!text(m.wire_id,128)||ids.has(m.id)||wireIds.has(m.wire_id)||!text(m.label,64)||!text(m.version,32)||!/^[0-9a-f]{64}$/.test(m.manifest_sha256)||!Number.isInteger(m.complete_input_tokens)||m.complete_input_tokens<1)throw Error('Invalid model identity.');
9
+ ids.add(m.id);wireIds.add(m.wire_id);
10
  if(m.description!==undefined&&!text(m.description,240)||m.parameter_label!==undefined&&!text(m.parameter_label,32))throw Error('Invalid model description.');
11
  if(m.example_ids!==undefined&&(!Array.isArray(m.example_ids)||!m.example_ids.length||m.example_ids.some(x=>!text(x,128))||new Set(m.example_ids).size!==m.example_ids.length))throw Error('Invalid model examples.');
12
  if(m.default_example_id!==undefined&&!text(m.default_example_id,128))throw Error('Invalid default example.');
static/tetris/app.js ADDED
@@ -0,0 +1,460 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ const elements = {
2
+ leftModel: document.querySelector('#left-model'),
3
+ rightModel: document.querySelector('#right-model'),
4
+ raceMode: document.querySelector('#race-mode'),
5
+ customWrap: document.querySelector('#custom-wrap'),
6
+ customSteps: document.querySelector('#custom-steps'),
7
+ seed: document.querySelector('#seed'),
8
+ start: document.querySelector('#start'),
9
+ status: document.querySelector('#status'),
10
+ summary: document.querySelector('#pk-summary'),
11
+ summaryLabel: document.querySelector('#summary-label'),
12
+ summaryLeft: document.querySelector('#summary-left'),
13
+ summaryRight: document.querySelector('#summary-right'),
14
+ summarySpeed: document.querySelector('#summary-speed'),
15
+ summaryScore: document.querySelector('#summary-score'),
16
+ closeSummary: document.querySelector('#close-summary'),
17
+ }
18
+
19
+ const views = Object.fromEntries(
20
+ ['left', 'right'].map((side) => {
21
+ const root = document.querySelector(`[data-side="${side}"]`)
22
+ const board = root.querySelector('[data-board]')
23
+ const cells = Array.from({ length: 200 }, () => {
24
+ const cell = document.createElement('span')
25
+ cell.className = 'cell'
26
+ cell.dataset.piece = '.'
27
+ board.append(cell)
28
+ return cell
29
+ })
30
+ return [side, {
31
+ root,
32
+ name: root.querySelector('[data-name]'),
33
+ board,
34
+ cells,
35
+ latency: root.querySelector('[data-latency]'),
36
+ score: root.querySelector('[data-score]'),
37
+ lines: root.querySelector('[data-lines]'),
38
+ pieces: root.querySelector('[data-pieces]'),
39
+ finish: root.querySelector('[data-finish]'),
40
+ resultScore: root.querySelector('[data-result-score]'),
41
+ resultLines: root.querySelector('[data-result-lines]'),
42
+ resultPieces: root.querySelector('[data-result-pieces]'),
43
+ }]
44
+ }),
45
+ )
46
+
47
+ let config = null
48
+ let raceId = null
49
+ let events = null
50
+ let rendering = false
51
+ let lifecycle = 'idle'
52
+ let operation = 0
53
+ let createController = null
54
+ const latest = { left: null, right: null }
55
+ const seenSteps = { left: 0, right: 0 }
56
+
57
+ function setStatus(message) {
58
+ elements.status.textContent = message
59
+ }
60
+
61
+ function addOption(select, competitor) {
62
+ const option = document.createElement('option')
63
+ option.value = competitor.id
64
+ option.textContent = competitor.option_label
65
+ select.append(option)
66
+ }
67
+
68
+ function selectedCompetitor(side) {
69
+ const value = side === 'left' ? elements.leftModel.value : elements.rightModel.value
70
+ return config?.competitors.find((competitor) => competitor.id === value) ?? null
71
+ }
72
+
73
+ function updateNames() {
74
+ for (const side of ['left', 'right']) {
75
+ views[side].name.textContent = selectedCompetitor(side)?.display_name ?? 'Unavailable'
76
+ }
77
+ }
78
+
79
+ function renderControls() {
80
+ const locked = lifecycle !== 'idle'
81
+ for (const control of [
82
+ elements.leftModel,
83
+ elements.rightModel,
84
+ elements.raceMode,
85
+ elements.customSteps,
86
+ elements.seed,
87
+ ]) control.disabled = locked
88
+ const hasCompetitor = Boolean(config?.competitors.some((item) => item.ready))
89
+ elements.start.disabled = !hasCompetitor || lifecycle === 'starting' || lifecycle === 'stopping'
90
+ elements.start.textContent = {
91
+ idle: 'Start race',
92
+ starting: 'Starting…',
93
+ running: 'Stop',
94
+ stopping: 'Stopping…',
95
+ }[lifecycle]
96
+ }
97
+
98
+ function setLifecycle(next) {
99
+ lifecycle = next
100
+ renderControls()
101
+ }
102
+
103
+ function clearBoard(side) {
104
+ for (const cell of views[side].cells) cell.dataset.piece = '.'
105
+ }
106
+
107
+ function resetView() {
108
+ latest.left = null
109
+ latest.right = null
110
+ seenSteps.left = 0
111
+ seenSteps.right = 0
112
+ for (const side of ['left', 'right']) {
113
+ clearBoard(side)
114
+ views[side].latency.textContent = '— ms'
115
+ views[side].score.textContent = '0'
116
+ views[side].lines.textContent = '0'
117
+ views[side].pieces.textContent = '0'
118
+ views[side].finish.hidden = true
119
+ }
120
+ elements.summary.hidden = true
121
+ updateNames()
122
+ }
123
+
124
+ function paint(side, step) {
125
+ const view = views[side]
126
+ const game = step.game
127
+ game.board.forEach((row, y) => {
128
+ Array.from(row).forEach((piece, x) => {
129
+ view.cells[(y * 10) + x].dataset.piece = piece
130
+ })
131
+ })
132
+ view.score.textContent = String(game.score)
133
+ view.lines.textContent = String(game.lines)
134
+ view.pieces.textContent = String(game.pieces)
135
+ if (Number.isFinite(step.request_ms)) {
136
+ view.latency.textContent = `${Math.round(step.request_ms)} ms`
137
+ }
138
+ if (step.cleared_rows?.length) {
139
+ view.board.classList.remove('is-clearing')
140
+ requestAnimationFrame(() => {
141
+ view.board.classList.add('is-clearing')
142
+ setTimeout(() => view.board.classList.remove('is-clearing'), 140)
143
+ })
144
+ }
145
+ }
146
+
147
+ function scheduleRender() {
148
+ if (rendering) return
149
+ rendering = true
150
+ requestAnimationFrame(() => {
151
+ rendering = false
152
+ for (const side of ['left', 'right']) {
153
+ if (latest[side]) paint(side, latest[side])
154
+ latest[side] = null
155
+ }
156
+ })
157
+ }
158
+
159
+ function showSideResult(side, result) {
160
+ const view = views[side]
161
+ view.finish.querySelector('[data-result-status]').textContent = {
162
+ finished: 'Finished',
163
+ game_over: 'Game over',
164
+ guard_reached: 'Limit reached',
165
+ error: 'Failed',
166
+ stopped: 'Stopped',
167
+ }[result.status] ?? 'Incomplete'
168
+ view.resultScore.textContent = String(result.score)
169
+ view.resultLines.textContent = String(result.lines)
170
+ view.resultPieces.textContent = String(result.pieces)
171
+ view.finish.hidden = false
172
+ }
173
+
174
+ function winnerName(winner, summary) {
175
+ if (winner === 'left') return summary.left.display_name
176
+ if (winner === 'right') return summary.right.display_name
177
+ return 'Tie'
178
+ }
179
+
180
+ function comparisonText(comparison, summary, label) {
181
+ if (comparison.winner === 'unavailable') return 'Not comparable'
182
+ if (comparison.winner === 'tie') return 'Tie'
183
+ const name = winnerName(comparison.winner, summary)
184
+ if (comparison.unbounded || comparison.percent == null) return `${name} wins ${label}`
185
+ const suffix = label === 'speed' ? 'faster' : 'higher'
186
+ return `${name} · ${comparison.percent.toFixed(1)}% ${suffix}`
187
+ }
188
+
189
+ function showSummary(summary) {
190
+ elements.summaryLabel.textContent = summary.status === 'finished'
191
+ ? 'Final result'
192
+ : 'Race incomplete'
193
+ elements.summaryLeft.textContent = summary.left.display_name
194
+ elements.summaryRight.textContent = summary.right.display_name
195
+ elements.summarySpeed.textContent = comparisonText(summary.speed, summary, 'speed')
196
+ elements.summaryScore.textContent = comparisonText(summary.score, summary, 'score')
197
+ elements.summary.hidden = false
198
+ }
199
+
200
+ function parseEvent(event) {
201
+ try {
202
+ return JSON.parse(event.data)
203
+ } catch {
204
+ return null
205
+ }
206
+ }
207
+
208
+ function raceStatus(status) {
209
+ return {
210
+ finished: 'Race complete.',
211
+ incomplete: 'Race incomplete. One or both models did not complete.',
212
+ timed_out: 'Race timed out before both models completed.',
213
+ cancelled: 'Race stopped before completion.',
214
+ error: 'The race server could not complete this race.',
215
+ }[status] ?? 'Race ended before completion.'
216
+ }
217
+
218
+ function finishRace(status, summary, stream) {
219
+ stream.close()
220
+ if (events === stream) events = null
221
+ raceId = null
222
+ setLifecycle('idle')
223
+ setStatus(raceStatus(status))
224
+ if (summary) {
225
+ for (const side of ['left', 'right']) {
226
+ if (summary[side]) showSideResult(side, summary[side])
227
+ }
228
+ showSummary(summary)
229
+ }
230
+ }
231
+
232
+ function applySnapshot(snapshot) {
233
+ for (const side of ['left', 'right']) {
234
+ const state = snapshot.states?.[side]
235
+ if (state?.game && state.step >= seenSteps[side]) {
236
+ seenSteps[side] = state.step
237
+ latest[side] = state
238
+ scheduleRender()
239
+ }
240
+ const result = snapshot.results?.[side]
241
+ if (result) showSideResult(side, result)
242
+ }
243
+ }
244
+
245
+ function connectEventStream(url, resultUrl, expectedOperation, expectedRaceId) {
246
+ events?.close()
247
+ const stream = new EventSource(url)
248
+ events = stream
249
+ let recovering = false
250
+ const current = () => expectedOperation === operation && expectedRaceId === raceId
251
+ async function recoverSnapshot() {
252
+ if (!current() || recovering) return
253
+ recovering = true
254
+ try {
255
+ const response = await fetch(resultUrl, { headers: { Accept: 'application/json' } })
256
+ if (!current()) return
257
+ if (response.status === 404) {
258
+ finishRace('error', null, stream)
259
+ setStatus('The race result is no longer available. Start a new race.')
260
+ return
261
+ }
262
+ if (!response.ok) return
263
+ const snapshot = await response.json()
264
+ if (!current() || snapshot.id !== expectedRaceId) return
265
+ applySnapshot(snapshot)
266
+ if (!['created', 'running'].includes(snapshot.status)) {
267
+ finishRace(snapshot.status, snapshot.summary, stream)
268
+ }
269
+ } catch {
270
+ // EventSource keeps retrying transient failures within the server lease.
271
+ } finally {
272
+ recovering = false
273
+ }
274
+ }
275
+ stream.addEventListener('side.step', (event) => {
276
+ if (!current()) return
277
+ const data = parseEvent(event)
278
+ if (!data || !views[data.side] || data.step <= seenSteps[data.side]) return
279
+ seenSteps[data.side] = data.step
280
+ latest[data.side] = data
281
+ scheduleRender()
282
+ })
283
+ stream.addEventListener('side.finished', (event) => {
284
+ if (!current()) return
285
+ const data = parseEvent(event)
286
+ if (data && views[data.side]) showSideResult(data.side, data.result)
287
+ })
288
+ stream.addEventListener('side.error', (event) => {
289
+ if (!current()) return
290
+ const data = parseEvent(event)
291
+ if (data) setStatus(`${data.side === 'left' ? 'Left' : 'Right'}: ${data.message}`)
292
+ })
293
+ stream.addEventListener('race.finished', (event) => {
294
+ if (!current()) return
295
+ const summary = parseEvent(event)
296
+ if (summary) finishRace(summary.status, summary, stream)
297
+ else void recoverSnapshot()
298
+ })
299
+ stream.addEventListener('open', () => {
300
+ if (current()) void recoverSnapshot()
301
+ })
302
+ stream.addEventListener('error', () => {
303
+ if (current() && lifecycle === 'running') {
304
+ setStatus('Reconnecting to the live race…')
305
+ void recoverSnapshot()
306
+ }
307
+ })
308
+ }
309
+
310
+ function raceSettings() {
311
+ const selected = elements.raceMode.value
312
+ if (selected === 'first_failure') {
313
+ return { mode: 'first_failure', max_steps: null }
314
+ }
315
+ const raw = selected === 'custom' ? elements.customSteps.value : selected
316
+ const maximum = config?.limits.max_steps ?? 120
317
+ const maxSteps = Math.min(maximum, Math.max(1, Number.parseInt(raw, 10) || 40))
318
+ if (selected === 'custom') elements.customSteps.value = String(maxSteps)
319
+ return { mode: 'steps', max_steps: maxSteps }
320
+ }
321
+
322
+ async function startRace() {
323
+ if (lifecycle !== 'idle') return
324
+ const expectedOperation = ++operation
325
+ createController = new AbortController()
326
+ resetView()
327
+ setLifecycle('starting')
328
+ setStatus('Starting both model runners…')
329
+ const settings = raceSettings()
330
+ let response
331
+ try {
332
+ response = await fetch('/api/tetris/races', {
333
+ method: 'POST',
334
+ headers: { 'Content-Type': 'application/json' },
335
+ signal: createController.signal,
336
+ body: JSON.stringify({
337
+ left: elements.leftModel.value,
338
+ right: elements.rightModel.value,
339
+ seed: elements.seed.value || '42',
340
+ ...settings,
341
+ }),
342
+ })
343
+ } catch (error) {
344
+ if (expectedOperation !== operation) return
345
+ createController = null
346
+ setLifecycle('idle')
347
+ if (error?.name === 'AbortError') return
348
+ setStatus('The race server is unavailable.')
349
+ return
350
+ }
351
+ createController = null
352
+ if (!response.ok) {
353
+ const detail = await response.json().catch(() => ({}))
354
+ if (expectedOperation !== operation) return
355
+ setLifecycle('idle')
356
+ setStatus(detail.detail || 'The race could not start.')
357
+ return
358
+ }
359
+ const race = await response.json().catch(() => null)
360
+ if (!race?.id || !race?.events_url || !race?.result_url) {
361
+ if (expectedOperation !== operation) return
362
+ setLifecycle('idle')
363
+ setStatus('The race server returned an invalid response.')
364
+ return
365
+ }
366
+ if (expectedOperation !== operation) {
367
+ await fetch(`/api/tetris/races/${encodeURIComponent(race.id)}`, {
368
+ method: 'DELETE',
369
+ keepalive: true,
370
+ }).catch(() => {})
371
+ return
372
+ }
373
+ raceId = race.id
374
+ setLifecycle('running')
375
+ setStatus('Race in progress. Each side is running independently.')
376
+ connectEventStream(race.events_url, race.result_url, expectedOperation, race.id)
377
+ }
378
+
379
+ async function stopRace() {
380
+ if (lifecycle !== 'running' || !raceId) return
381
+ const expectedOperation = ++operation
382
+ const current = raceId
383
+ raceId = null
384
+ events?.close()
385
+ events = null
386
+ setLifecycle('stopping')
387
+ await fetch(`/api/tetris/races/${encodeURIComponent(current)}`, {
388
+ method: 'DELETE',
389
+ }).catch(() => {})
390
+ if (expectedOperation !== operation) return
391
+ setLifecycle('idle')
392
+ setStatus('Race stopped.')
393
+ }
394
+
395
+ async function loadConfig() {
396
+ try {
397
+ const response = await fetch('/api/tetris/config', { headers: { Accept: 'application/json' } })
398
+ if (!response.ok) throw new Error('config unavailable')
399
+ config = await response.json()
400
+ } catch {
401
+ setStatus('The arena configuration is unavailable.')
402
+ return
403
+ }
404
+ const competitors = config.competitors.filter((competitor) => competitor.ready)
405
+ const maximum = config.limits.max_steps
406
+ elements.customSteps.max = String(maximum)
407
+ for (const option of elements.raceMode.options) {
408
+ if (/^\d+$/.test(option.value)) option.disabled = Number(option.value) > maximum
409
+ }
410
+ const defaultSteps = Math.min(config.defaults.max_steps, maximum)
411
+ if ([...elements.raceMode.options].some((option) => option.value === String(defaultSteps))) {
412
+ elements.raceMode.value = String(defaultSteps)
413
+ } else {
414
+ elements.raceMode.value = 'custom'
415
+ elements.customSteps.value = String(defaultSteps)
416
+ elements.customWrap.hidden = false
417
+ }
418
+ for (const competitor of competitors) {
419
+ addOption(elements.leftModel, competitor)
420
+ addOption(elements.rightModel, competitor)
421
+ }
422
+ const cloud = competitors.find((item) => item.id === 'jev-cloud')
423
+ const lux = competitors.find((item) => item.id === 'lux')
424
+ elements.leftModel.value = cloud?.id ?? competitors[0]?.id ?? ''
425
+ elements.rightModel.value = lux?.id ?? competitors[1]?.id ?? competitors[0]?.id ?? ''
426
+ renderControls()
427
+ setStatus(
428
+ competitors.length > 0
429
+ ? 'Choose any two configured models.'
430
+ : 'No Tetris model endpoints are configured on this server.',
431
+ )
432
+ resetView()
433
+ }
434
+
435
+ elements.raceMode.addEventListener('change', () => {
436
+ elements.customWrap.hidden = elements.raceMode.value !== 'custom'
437
+ })
438
+ elements.leftModel.addEventListener('change', updateNames)
439
+ elements.rightModel.addEventListener('change', updateNames)
440
+ elements.start.addEventListener('click', () => {
441
+ if (lifecycle === 'running') void stopRace()
442
+ else if (lifecycle === 'idle') void startRace()
443
+ })
444
+ elements.closeSummary.addEventListener('click', () => { elements.summary.hidden = true })
445
+ elements.summary.addEventListener('click', (event) => {
446
+ if (event.target === elements.summary) elements.summary.hidden = true
447
+ })
448
+ window.addEventListener('beforeunload', () => {
449
+ operation += 1
450
+ createController?.abort()
451
+ events?.close()
452
+ if (raceId) {
453
+ void fetch(`/api/tetris/races/${encodeURIComponent(raceId)}`, {
454
+ method: 'DELETE',
455
+ keepalive: true,
456
+ }).catch(() => {})
457
+ }
458
+ })
459
+
460
+ void loadConfig()
static/tetris/huggingface-logo.svg ADDED
static/tetris/index.html ADDED
@@ -0,0 +1,138 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ <!doctype html>
2
+ <html lang="en">
3
+ <head>
4
+ <meta charset="utf-8" />
5
+ <meta name="viewport" content="width=device-width, initial-scale=1" />
6
+ <meta
7
+ name="description"
8
+ content="A server-authoritative head-to-head Tetris race for Decision 1.0 models."
9
+ />
10
+ <title>Decision 1.0 · Tetris Race</title>
11
+ <link rel="stylesheet" href="styles.css" />
12
+ </head>
13
+ <body>
14
+ <main class="page">
15
+ <header class="toolbar">
16
+ <a class="wordmark" href="/tetris/" aria-label="Decision 1.0 Tetris Race">
17
+ <span class="wordmark__eyebrow">Decision 1.0 playground</span>
18
+ <strong>Tetris Race</strong>
19
+ </a>
20
+
21
+ <div class="controls" aria-label="Race controls">
22
+ <label class="control control--model">
23
+ <span>Left model</span>
24
+ <select id="left-model" aria-label="Left model"></select>
25
+ </label>
26
+ <span class="versus" aria-hidden="true">VS</span>
27
+ <label class="control control--model">
28
+ <span>Right model</span>
29
+ <select id="right-model" aria-label="Right model"></select>
30
+ </label>
31
+ <label class="control control--steps">
32
+ <span>Steps</span>
33
+ <select id="race-mode" aria-label="Race length">
34
+ <option value="20">20</option>
35
+ <option value="40" selected>40</option>
36
+ <option value="80">80</option>
37
+ <option value="120">120</option>
38
+ <option value="custom">Custom</option>
39
+ <option value="first_failure">Until first failure</option>
40
+ </select>
41
+ </label>
42
+ <label class="control control--custom" id="custom-wrap" hidden>
43
+ <span>Custom</span>
44
+ <input id="custom-steps" type="number" min="1" max="120" value="40" />
45
+ </label>
46
+ <label class="control control--seed">
47
+ <span>Seed</span>
48
+ <input id="seed" value="42" maxlength="128" />
49
+ </label>
50
+ <button id="start" class="start" type="button" disabled>Start race</button>
51
+ </div>
52
+ </header>
53
+
54
+ <section class="arena" aria-live="polite">
55
+ <article class="player" data-side="left">
56
+ <header class="player__header">
57
+ <h2 data-name>Left model</h2>
58
+ <output data-latency>— ms</output>
59
+ </header>
60
+ <div class="board-shell">
61
+ <div class="board" data-board aria-label="Left Tetris board"></div>
62
+ <div class="finish-card" data-finish hidden>
63
+ <span data-result-status>Result</span>
64
+ <strong data-result-score>0</strong>
65
+ <small class="finish-card__score-label">Score</small>
66
+ <div class="finish-card__metrics">
67
+ <p><b data-result-lines>0</b><small>Lines</small></p>
68
+ <p><b data-result-pieces>0</b><small>Pieces</small></p>
69
+ </div>
70
+ </div>
71
+ </div>
72
+ <div class="metrics">
73
+ <p><span>Score</span><strong data-score>0</strong></p>
74
+ <p><span>Lines</span><strong data-lines>0</strong></p>
75
+ <p><span>Pieces</span><strong data-pieces>0</strong></p>
76
+ </div>
77
+ </article>
78
+
79
+ <div class="arena__divider" aria-hidden="true"><span>PK</span></div>
80
+
81
+ <article class="player" data-side="right">
82
+ <header class="player__header">
83
+ <h2 data-name>Right model</h2>
84
+ <output data-latency>— ms</output>
85
+ </header>
86
+ <div class="board-shell">
87
+ <div class="board" data-board aria-label="Right Tetris board"></div>
88
+ <div class="finish-card" data-finish hidden>
89
+ <span data-result-status>Result</span>
90
+ <strong data-result-score>0</strong>
91
+ <small class="finish-card__score-label">Score</small>
92
+ <div class="finish-card__metrics">
93
+ <p><b data-result-lines>0</b><small>Lines</small></p>
94
+ <p><b data-result-pieces>0</b><small>Pieces</small></p>
95
+ </div>
96
+ </div>
97
+ </div>
98
+ <div class="metrics">
99
+ <p><span>Score</span><strong data-score>0</strong></p>
100
+ <p><span>Lines</span><strong data-lines>0</strong></p>
101
+ <p><span>Pieces</span><strong data-pieces>0</strong></p>
102
+ </div>
103
+ </article>
104
+
105
+ <section class="pk-summary" id="pk-summary" hidden aria-labelledby="pk-title">
106
+ <div class="pk-summary__card">
107
+ <button id="close-summary" class="summary-close" aria-label="Close result">×</button>
108
+ <span class="pk-summary__label" id="summary-label">Final result</span>
109
+ <h2 id="pk-title">PK</h2>
110
+ <div class="pk-summary__models">
111
+ <strong id="summary-left">Left</strong>
112
+ <span>vs</span>
113
+ <strong id="summary-right">Right</strong>
114
+ </div>
115
+ <div class="pk-summary__comparisons">
116
+ <p><span>Speed</span><strong id="summary-speed">—</strong></p>
117
+ <p><span>Score</span><strong id="summary-score">—</strong></p>
118
+ </div>
119
+ </div>
120
+ </section>
121
+ </section>
122
+
123
+ <p class="status" id="status" role="status">Configure two models to begin.</p>
124
+
125
+ <footer class="footer">
126
+ <a
127
+ href="https://huggingface.co/collections/llm-semantic-router/decision-10"
128
+ target="_blank"
129
+ rel="noopener noreferrer"
130
+ >
131
+ <img src="huggingface-logo.svg" width="62" height="58" alt="Hugging Face" />
132
+ <span><strong>Decision 1.0</strong>Open Decision Foundation Models</span>
133
+ </a>
134
+ </footer>
135
+ </main>
136
+ <script type="module" src="app.js"></script>
137
+ </body>
138
+ </html>
static/tetris/styles.css ADDED
@@ -0,0 +1,486 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ :root {
2
+ color-scheme: dark;
3
+ --page: #05070b;
4
+ --surface: rgba(15, 19, 27, 0.88);
5
+ --line: rgba(255, 255, 255, 0.11);
6
+ --muted: #8993a3;
7
+ --ink: #f7f9fc;
8
+ --accent: #63c5ff;
9
+ --board-width: clamp(250px, min(23vw, 36vh), 360px);
10
+ font-family: Inter, ui-sans-serif, -apple-system, BlinkMacSystemFont, "Segoe UI", sans-serif;
11
+ }
12
+
13
+ * { box-sizing: border-box; }
14
+
15
+ html,
16
+ body { min-width: 320px; min-height: 100%; }
17
+
18
+ body {
19
+ margin: 0;
20
+ color: var(--ink);
21
+ background:
22
+ radial-gradient(circle at 18% 12%, rgba(35, 132, 255, 0.12), transparent 32rem),
23
+ radial-gradient(circle at 82% 14%, rgba(194, 78, 255, 0.09), transparent 30rem),
24
+ linear-gradient(rgba(255, 255, 255, 0.016) 1px, transparent 1px),
25
+ linear-gradient(90deg, rgba(255, 255, 255, 0.016) 1px, transparent 1px),
26
+ #05070b;
27
+ background-size: auto, auto, 36px 36px, 36px 36px, auto;
28
+ -webkit-font-smoothing: antialiased;
29
+ }
30
+
31
+ button,
32
+ input,
33
+ select { font: inherit; }
34
+
35
+ [hidden] { display: none !important; }
36
+
37
+ .page {
38
+ display: grid;
39
+ min-height: 100dvh;
40
+ grid-template-rows: auto minmax(0, 1fr) auto auto;
41
+ gap: clamp(0.75rem, 1.6vh, 1.35rem);
42
+ width: min(1320px, 100%);
43
+ margin: 0 auto;
44
+ padding: clamp(0.8rem, 1.7vw, 1.5rem) clamp(1rem, 3vw, 2.8rem) 1.25rem;
45
+ }
46
+
47
+ .toolbar {
48
+ position: relative;
49
+ z-index: 10;
50
+ display: flex;
51
+ align-items: center;
52
+ justify-content: space-between;
53
+ gap: 1.5rem;
54
+ min-height: 78px;
55
+ padding: 0.8rem 0.95rem 0.8rem 1.25rem;
56
+ border: 1px solid var(--line);
57
+ border-radius: 18px;
58
+ background: rgba(12, 16, 23, 0.78);
59
+ box-shadow: 0 20px 70px rgba(0, 0, 0, 0.32), inset 0 1px rgba(255, 255, 255, 0.045);
60
+ backdrop-filter: blur(22px);
61
+ }
62
+
63
+ .wordmark {
64
+ min-width: 175px;
65
+ color: inherit;
66
+ text-decoration: none;
67
+ }
68
+
69
+ .wordmark__eyebrow {
70
+ display: block;
71
+ margin-bottom: 0.24rem;
72
+ color: #7f8a9b;
73
+ font-size: 0.61rem;
74
+ font-weight: 750;
75
+ letter-spacing: 0.14em;
76
+ text-transform: uppercase;
77
+ }
78
+
79
+ .wordmark strong {
80
+ display: block;
81
+ font-size: clamp(1.45rem, 2.2vw, 1.9rem);
82
+ letter-spacing: -0.055em;
83
+ line-height: 1;
84
+ }
85
+
86
+ .controls {
87
+ display: flex;
88
+ min-width: 0;
89
+ align-items: end;
90
+ justify-content: flex-end;
91
+ gap: 0.42rem;
92
+ }
93
+
94
+ .control {
95
+ display: grid;
96
+ min-width: 0;
97
+ gap: 0.27rem;
98
+ }
99
+
100
+ .control > span {
101
+ padding-left: 0.22rem;
102
+ color: var(--muted);
103
+ font-size: 0.55rem;
104
+ font-weight: 760;
105
+ letter-spacing: 0.1em;
106
+ text-transform: uppercase;
107
+ }
108
+
109
+ .control--model { width: 10.5rem; }
110
+ .control--steps { width: 8.2rem; }
111
+ .control--custom { width: 5rem; }
112
+ .control--seed { width: 4.5rem; }
113
+
114
+ .control input,
115
+ .control select,
116
+ .start {
117
+ height: 39px;
118
+ border-radius: 10px;
119
+ }
120
+
121
+ .control input,
122
+ .control select {
123
+ width: 100%;
124
+ min-width: 0;
125
+ border: 1px solid rgba(255, 255, 255, 0.13);
126
+ outline: none;
127
+ color: #111720;
128
+ background-color: #f7f9fc;
129
+ font-size: 0.73rem;
130
+ font-weight: 650;
131
+ transition: border-color 150ms ease, box-shadow 150ms ease;
132
+ }
133
+
134
+ .control select {
135
+ appearance: none;
136
+ padding: 0 1.9rem 0 0.68rem;
137
+ background-image: url("data:image/svg+xml,%3Csvg xmlns='http://www.w3.org/2000/svg' width='12' height='8' viewBox='0 0 12 8'%3E%3Cpath d='m1.5 1.5 4.5 4.5 4.5-4.5' fill='none' stroke='%23323a46' stroke-width='1.5' stroke-linecap='round' stroke-linejoin='round'/%3E%3C/svg%3E");
138
+ background-position: right 0.7rem center;
139
+ background-repeat: no-repeat;
140
+ }
141
+
142
+ .control input { padding: 0 0.62rem; }
143
+
144
+ .control input:focus,
145
+ .control select:focus {
146
+ border-color: var(--accent);
147
+ box-shadow: 0 0 0 3px rgba(99, 197, 255, 0.16);
148
+ }
149
+
150
+ .versus {
151
+ display: grid;
152
+ height: 39px;
153
+ place-items: center;
154
+ padding-inline: 0.15rem;
155
+ color: #687282;
156
+ font-size: 0.56rem;
157
+ font-weight: 850;
158
+ }
159
+
160
+ .start {
161
+ min-width: 88px;
162
+ padding: 0 0.85rem;
163
+ border: 0;
164
+ color: #04121c;
165
+ background: linear-gradient(145deg, #9cdbff, #48b9fa 55%, #258ed0);
166
+ box-shadow: 0 10px 26px rgba(48, 166, 235, 0.22);
167
+ font-size: 0.73rem;
168
+ font-weight: 800;
169
+ cursor: pointer;
170
+ transition: transform 150ms ease, opacity 150ms ease;
171
+ }
172
+
173
+ .start:hover:not(:disabled) { transform: translateY(-1px); }
174
+ .start:active:not(:disabled) { transform: translateY(0) scale(0.98); }
175
+ .start:disabled { cursor: not-allowed; opacity: 0.38; }
176
+
177
+ .arena {
178
+ position: relative;
179
+ display: grid;
180
+ grid-template-columns: var(--board-width) clamp(42px, 6vw, 78px) var(--board-width);
181
+ align-items: center;
182
+ justify-content: center;
183
+ gap: clamp(0.75rem, 2vw, 1.75rem);
184
+ min-height: 0;
185
+ }
186
+
187
+ .player {
188
+ width: var(--board-width);
189
+ min-width: 0;
190
+ }
191
+
192
+ .player__header {
193
+ display: grid;
194
+ grid-template-columns: minmax(0, 1fr) auto;
195
+ align-items: baseline;
196
+ gap: 0.75rem;
197
+ min-height: 39px;
198
+ padding: 0 0.18rem 0.58rem;
199
+ }
200
+
201
+ .player__header h2 {
202
+ overflow: hidden;
203
+ margin: 0;
204
+ font-size: clamp(1.17rem, 2vw, 1.65rem);
205
+ font-weight: 720;
206
+ letter-spacing: -0.045em;
207
+ line-height: 1;
208
+ text-overflow: ellipsis;
209
+ white-space: nowrap;
210
+ }
211
+
212
+ .player__header output {
213
+ color: #a7b1bf;
214
+ font-size: 0.75rem;
215
+ font-variant-numeric: tabular-nums;
216
+ font-weight: 700;
217
+ }
218
+
219
+ .board-shell {
220
+ position: relative;
221
+ padding: 9px;
222
+ overflow: hidden;
223
+ border: 1px solid rgba(255, 255, 255, 0.14);
224
+ border-radius: 14px;
225
+ background: linear-gradient(145deg, #252c38, #0e1219 42%, #282e38);
226
+ box-shadow:
227
+ 0 24px 70px rgba(0, 0, 0, 0.45),
228
+ inset 0 1px rgba(255, 255, 255, 0.12),
229
+ inset 0 -1px rgba(0, 0, 0, 0.8);
230
+ }
231
+
232
+ .board-shell::before {
233
+ position: absolute;
234
+ z-index: 3;
235
+ inset: 4px;
236
+ border: 1px solid rgba(255, 255, 255, 0.07);
237
+ border-radius: 10px;
238
+ content: "";
239
+ pointer-events: none;
240
+ }
241
+
242
+ .board {
243
+ display: grid;
244
+ width: 100%;
245
+ aspect-ratio: 1 / 2;
246
+ grid-template-columns: repeat(10, 1fr);
247
+ grid-template-rows: repeat(20, 1fr);
248
+ overflow: hidden;
249
+ border: 1px solid rgba(0, 0, 0, 0.92);
250
+ background:
251
+ linear-gradient(rgba(255, 255, 255, 0.027) 1px, transparent 1px),
252
+ linear-gradient(90deg, rgba(255, 255, 255, 0.027) 1px, transparent 1px),
253
+ radial-gradient(circle at 50% 8%, #111927, #05070b 62%);
254
+ background-size: 10% 5%, 10% 5%, auto;
255
+ box-shadow: inset 0 0 38px rgba(0, 0, 0, 0.66);
256
+ }
257
+
258
+ .cell {
259
+ position: relative;
260
+ min-width: 0;
261
+ min-height: 0;
262
+ transition: background-color 80ms linear, box-shadow 80ms linear;
263
+ }
264
+
265
+ .cell[data-piece]:not([data-piece="."]) {
266
+ border: 1px solid rgba(0, 0, 0, 0.32);
267
+ border-radius: 18%;
268
+ box-shadow:
269
+ inset 2px 2px 0 rgba(255, 255, 255, 0.38),
270
+ inset -3px -3px 0 rgba(0, 0, 0, 0.24),
271
+ 0 0 9px color-mix(in srgb, var(--piece-color) 42%, transparent);
272
+ background:
273
+ linear-gradient(145deg, rgba(255, 255, 255, 0.25), transparent 35%),
274
+ linear-gradient(320deg, rgba(0, 0, 0, 0.21), transparent 45%),
275
+ var(--piece-color);
276
+ }
277
+
278
+ .cell[data-piece="I"] { --piece-color: #28cce9; }
279
+ .cell[data-piece="O"] { --piece-color: #ffca32; }
280
+ .cell[data-piece="T"] { --piece-color: #a970ff; }
281
+ .cell[data-piece="S"] { --piece-color: #44d79f; }
282
+ .cell[data-piece="Z"] { --piece-color: #ff6070; }
283
+ .cell[data-piece="J"] { --piece-color: #4f82ff; }
284
+ .cell[data-piece="L"] { --piece-color: #ff964d; }
285
+
286
+ .board.is-clearing { filter: brightness(1.38); }
287
+
288
+ .metrics {
289
+ display: grid;
290
+ grid-template-columns: repeat(3, 1fr);
291
+ gap: 0.42rem;
292
+ margin-top: 0.52rem;
293
+ }
294
+
295
+ .metrics p {
296
+ display: flex;
297
+ align-items: baseline;
298
+ justify-content: space-between;
299
+ gap: 0.35rem;
300
+ margin: 0;
301
+ padding: 0.48rem 0.58rem;
302
+ border: 1px solid rgba(255, 255, 255, 0.075);
303
+ border-radius: 9px;
304
+ background: rgba(255, 255, 255, 0.025);
305
+ }
306
+
307
+ .metrics span { color: var(--muted); font-size: 0.58rem; text-transform: uppercase; }
308
+ .metrics strong { font-size: 0.78rem; font-variant-numeric: tabular-nums; }
309
+
310
+ .arena__divider {
311
+ display: grid;
312
+ align-self: center;
313
+ place-items: center;
314
+ height: 100%;
315
+ }
316
+
317
+ .arena__divider::before {
318
+ width: 1px;
319
+ height: 34%;
320
+ background: linear-gradient(transparent, rgba(255, 255, 255, 0.14), transparent);
321
+ content: "";
322
+ }
323
+
324
+ .arena__divider span {
325
+ display: grid;
326
+ width: 42px;
327
+ height: 42px;
328
+ place-items: center;
329
+ margin-block: 0.65rem;
330
+ border: 1px solid rgba(255, 255, 255, 0.12);
331
+ border-radius: 50%;
332
+ color: #b7c0cd;
333
+ background: #0b0e14;
334
+ box-shadow: 0 10px 28px rgba(0, 0, 0, 0.4);
335
+ font-size: 0.67rem;
336
+ font-weight: 850;
337
+ }
338
+
339
+ .finish-card {
340
+ position: absolute;
341
+ z-index: 4;
342
+ inset: 9px;
343
+ display: grid;
344
+ place-content: center;
345
+ justify-items: center;
346
+ background: rgba(4, 7, 11, 0.88);
347
+ backdrop-filter: blur(11px);
348
+ animation: finish-in 280ms cubic-bezier(0.2, 0.9, 0.2, 1) both;
349
+ }
350
+
351
+ .finish-card > span {
352
+ margin-bottom: 0.28rem;
353
+ color: #9ca7b7;
354
+ font-size: 0.68rem;
355
+ font-weight: 800;
356
+ letter-spacing: 0.15em;
357
+ text-transform: uppercase;
358
+ }
359
+
360
+ .finish-card > strong {
361
+ font-size: clamp(2.8rem, 6vw, 4.6rem);
362
+ letter-spacing: -0.08em;
363
+ line-height: 1;
364
+ }
365
+
366
+ .finish-card__score-label {
367
+ margin-top: 0.2rem;
368
+ color: var(--muted);
369
+ font-size: 0.6rem;
370
+ letter-spacing: 0.11em;
371
+ text-transform: uppercase;
372
+ }
373
+
374
+ .finish-card__metrics { display: flex; gap: 2rem; margin-top: 1.1rem; }
375
+ .finish-card__metrics p { display: grid; justify-items: center; gap: 0.12rem; margin: 0; }
376
+ .finish-card__metrics b { font-size: 1.25rem; }
377
+ .finish-card__metrics small { color: var(--muted); font-size: 0.63rem; text-transform: uppercase; }
378
+
379
+ .pk-summary {
380
+ position: fixed;
381
+ z-index: 30;
382
+ inset: 0;
383
+ display: grid;
384
+ place-items: center;
385
+ padding: 1rem;
386
+ background: rgba(2, 4, 8, 0.72);
387
+ backdrop-filter: blur(13px);
388
+ }
389
+
390
+ .pk-summary__card {
391
+ position: relative;
392
+ width: min(620px, 94vw);
393
+ padding: clamp(1.5rem, 4vw, 2.7rem);
394
+ border: 1px solid rgba(255, 255, 255, 0.15);
395
+ border-radius: 24px;
396
+ background:
397
+ radial-gradient(circle at 50% 0, rgba(84, 184, 255, 0.16), transparent 55%),
398
+ #0d1119;
399
+ box-shadow: 0 36px 100px rgba(0, 0, 0, 0.65), inset 0 1px rgba(255, 255, 255, 0.06);
400
+ text-align: center;
401
+ animation: summary-in 320ms cubic-bezier(0.18, 0.9, 0.2, 1) both;
402
+ }
403
+
404
+ .pk-summary__label { color: var(--muted); font-size: 0.66rem; font-weight: 800; letter-spacing: 0.16em; text-transform: uppercase; }
405
+ .pk-summary h2 { margin: 0.35rem 0 0.75rem; font-size: 2.4rem; letter-spacing: -0.08em; }
406
+ .pk-summary__models { display: grid; grid-template-columns: 1fr auto 1fr; align-items: center; gap: 0.8rem; }
407
+ .pk-summary__models strong { overflow: hidden; font-size: clamp(1.1rem, 3vw, 1.55rem); text-overflow: ellipsis; white-space: nowrap; }
408
+ .pk-summary__models span { color: #667182; font-size: 0.7rem; font-weight: 800; text-transform: uppercase; }
409
+ .pk-summary__comparisons { display: grid; grid-template-columns: 1fr 1fr; gap: 0.8rem; margin-top: 1.4rem; }
410
+ .pk-summary__comparisons p { margin: 0; padding: 1rem; border: 1px solid var(--line); border-radius: 14px; background: rgba(255, 255, 255, 0.025); }
411
+ .pk-summary__comparisons span { display: block; margin-bottom: 0.32rem; color: var(--muted); font-size: 0.62rem; font-weight: 800; letter-spacing: 0.12em; text-transform: uppercase; }
412
+ .pk-summary__comparisons strong { display: block; font-size: clamp(1.05rem, 2.6vw, 1.5rem); line-height: 1.15; }
413
+
414
+ .summary-close {
415
+ position: absolute;
416
+ top: 0.8rem;
417
+ right: 0.9rem;
418
+ width: 34px;
419
+ height: 34px;
420
+ border: 0;
421
+ border-radius: 50%;
422
+ color: #b6bfcb;
423
+ background: rgba(255, 255, 255, 0.06);
424
+ font-size: 1.25rem;
425
+ cursor: pointer;
426
+ }
427
+
428
+ .status { min-height: 1em; margin: 0; color: #747f90; font-size: 0.68rem; text-align: center; }
429
+
430
+ .footer {
431
+ display: grid;
432
+ min-height: clamp(74px, 10vh, 112px);
433
+ place-items: end center;
434
+ padding-top: 0.7rem;
435
+ }
436
+
437
+ .footer a {
438
+ display: flex;
439
+ align-items: center;
440
+ gap: 0.85rem;
441
+ color: inherit;
442
+ text-decoration: none;
443
+ }
444
+
445
+ .footer img { width: 48px; height: auto; filter: drop-shadow(0 8px 18px rgba(255, 183, 0, 0.13)); }
446
+ .footer span { display: grid; gap: 0.14rem; color: #8a95a5; font-size: 0.72rem; }
447
+ .footer strong { color: #eff3f8; font-size: 1.05rem; letter-spacing: -0.025em; }
448
+
449
+ @keyframes finish-in {
450
+ from { opacity: 0; transform: scale(0.96); }
451
+ to { opacity: 1; transform: scale(1); }
452
+ }
453
+
454
+ @keyframes summary-in {
455
+ from { opacity: 0; transform: translateY(18px) scale(0.965); }
456
+ to { opacity: 1; transform: translateY(0) scale(1); }
457
+ }
458
+
459
+ @media (max-width: 980px) {
460
+ .toolbar { align-items: flex-start; flex-direction: column; }
461
+ .controls { width: 100%; flex-wrap: wrap; justify-content: flex-start; }
462
+ .control--model { flex: 1 1 10rem; }
463
+ .arena { grid-template-columns: var(--board-width) 35px var(--board-width); gap: 0.5rem; }
464
+ }
465
+
466
+ @media (max-width: 680px) {
467
+ :root { --board-width: min(78vw, 330px); }
468
+ .page { display: block; }
469
+ .toolbar { margin-bottom: 1.25rem; }
470
+ .controls { display: grid; grid-template-columns: 1fr 1fr; }
471
+ .versus { display: none; }
472
+ .control--model,
473
+ .control--steps,
474
+ .control--custom,
475
+ .control--seed { width: auto; }
476
+ .start { align-self: end; }
477
+ .arena { grid-template-columns: 1fr; gap: 1.4rem; }
478
+ .player { margin-inline: auto; }
479
+ .arena__divider { display: none; }
480
+ .status { margin: 1rem 0; }
481
+ .footer { padding-block: 1.5rem; }
482
+ }
483
+
484
+ @media (prefers-reduced-motion: reduce) {
485
+ *, *::before, *::after { scroll-behavior: auto !important; animation-duration: 1ms !important; transition-duration: 1ms !important; }
486
+ }
systemone_api.py CHANGED
@@ -25,15 +25,30 @@ def resolve_model(value, registry):
25
  raise HTTPException(422, "This model is not available in this Studio")
26
 
27
 
 
 
 
 
 
 
 
 
 
 
 
 
 
28
  def public_models(registry):
29
- return [dict(item, name=item['label'],
30
  description=item.get('description', 'Decision foundation model'),
31
  release_date=RELEASE_DATES[key]) for key, item in registry.items()]
32
 
33
 
34
- def sdk_response(result):
35
  """Add the SDK's required distribution statistic without altering native outputs."""
36
  result = deepcopy(result)
 
 
37
  groups = [result['answers']] if 'answers' in result else [r['answers'] for r in result['results']]
38
  for answers in groups:
39
  for answer in answers.values():
@@ -42,3 +57,34 @@ def sdk_response(result):
42
  answer['confidence'] = max(probabilities)
43
  result['profile'] = dict(result.get('profile', {}), confidence_definition=CONFIDENCE_DEFINITION)
44
  return result
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
25
  raise HTTPException(422, "This model is not available in this Studio")
26
 
27
 
28
+ def resolve_canonical_model(value, registry):
29
+ """Resolve the public SystemOne identifier without accepting wire aliases."""
30
+ if not isinstance(value, str) or not value.strip():
31
+ raise HTTPException(422, "Specify an available canonical Decision model")
32
+ for key, item in registry.items():
33
+ if value == item['repo_id']:
34
+ return key
35
+ raise HTTPException(
36
+ 422,
37
+ "Use the exact Hugging Face repository ID for an available Decision model.",
38
+ )
39
+
40
+
41
  def public_models(registry):
42
+ return [dict(item, id=item['repo_id'], wire_id=key, name=item['label'],
43
  description=item.get('description', 'Decision foundation model'),
44
  release_date=RELEASE_DATES[key]) for key, item in registry.items()]
45
 
46
 
47
+ def sdk_response(result, *, public_model=None):
48
  """Add the SDK's required distribution statistic without altering native outputs."""
49
  result = deepcopy(result)
50
+ if public_model is not None:
51
+ result['model'] = public_model
52
  groups = [result['answers']] if 'answers' in result else [r['answers'] for r in result['results']]
53
  for answers in groups:
54
  for answer in answers.values():
 
57
  answer['confidence'] = max(probabilities)
58
  result['profile'] = dict(result.get('profile', {}), confidence_definition=CONFIDENCE_DEFINITION)
59
  return result
60
+
61
+
62
+ def public_response(result, *, public_model, batch):
63
+ """Project native diagnostics onto the strict public Decision envelope."""
64
+ prepared = sdk_response(result, public_model=public_model)
65
+
66
+ def answers_only(answers):
67
+ fields = {
68
+ 'noul': ('type', 'noul'),
69
+ 'choice': ('type', 'choice', 'confidence', 'probabilities'),
70
+ 'score': ('type', 'score', 'confidence', 'legend', 'probabilities'),
71
+ }
72
+ return {
73
+ key: {field: answer[field] for field in fields[answer['type']]}
74
+ for key, answer in answers.items()
75
+ }
76
+
77
+ if batch:
78
+ return {
79
+ 'model': public_model,
80
+ 'results': [
81
+ {'id': row['id'], 'answers': answers_only(row['answers']), 'usage': row['usage']}
82
+ for row in prepared['results']
83
+ ],
84
+ 'usage': prepared['usage'],
85
+ }
86
+ return {
87
+ 'model': public_model,
88
+ 'answers': answers_only(prepared['answers']),
89
+ 'usage': prepared['usage'],
90
+ }
tests/test_public_api_contract.py ADDED
@@ -0,0 +1,104 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Public Decision endpoints keep canonical IDs separate from Studio wire IDs."""
2
+
3
+ import unittest
4
+
5
+ from fastapi.testclient import TestClient
6
+
7
+ from app import create_app
8
+ from contract import MODEL, to_response
9
+ from engine import DEFAULT_MANIFEST
10
+ from model_registry import PROFILES
11
+
12
+ CANONICAL = PROFILES[MODEL]["repo_id"]
13
+
14
+
15
+ class FakeEngine:
16
+ model = MODEL
17
+ manifest = DEFAULT_MANIFEST
18
+
19
+ def evaluate(self, payload, records):
20
+ predictions = []
21
+ for row in records:
22
+ predictions.append({
23
+ "id": row["id"],
24
+ "question_id": row["question"]["id"],
25
+ "type": "noul",
26
+ "candidate_ids": ["no", "yes"],
27
+ "probabilities": [0.25, 0.75],
28
+ "input_tokens": 8,
29
+ "state_tokens_original": 0,
30
+ "state_tokens_kept": 0,
31
+ })
32
+ result = to_response(payload, records, predictions, model=MODEL)
33
+ result["source"] = "live_native"
34
+ result["timing"] = {
35
+ "model_load_ms": 0,
36
+ "inference_ms": 1,
37
+ "server_ms": 2,
38
+ "cold_start": False,
39
+ }
40
+ return result
41
+
42
+
43
+ def questions():
44
+ return {"check": {"type": "noul", "instructions": "Is this a request?"}}
45
+
46
+
47
+ class PublicAPIContractTests(unittest.TestCase):
48
+ def setUp(self):
49
+ self.app_context = TestClient(create_app(
50
+ mode="native",
51
+ native_engine=FakeEngine(),
52
+ registry=[{
53
+ "id": MODEL,
54
+ "label": "Kai",
55
+ "version": "1.0",
56
+ "manifest_sha256": DEFAULT_MANIFEST,
57
+ }],
58
+ ))
59
+ self.client = self.app_context.__enter__()
60
+
61
+ def tearDown(self):
62
+ self.app_context.__exit__(None, None, None)
63
+
64
+ def test_single_requires_explicit_canonical_model_and_strict_envelope(self):
65
+ payload = {"model": CANONICAL, "state": "Please help", "questions": questions()}
66
+ result = self.client.post("/v1/systemone", json=payload)
67
+ self.assertEqual(result.status_code, 200)
68
+ body = result.json()
69
+ self.assertEqual(set(body), {"model", "answers", "usage"})
70
+ self.assertEqual(body["model"], CANONICAL)
71
+ self.assertEqual(set(body["answers"]["check"]), {"type", "noul"})
72
+
73
+ for invalid in (dict(payload, model=MODEL), dict(payload, model="Kai"),
74
+ {key: value for key, value in payload.items() if key != "model"},
75
+ dict(payload, states=[{"id": "one", "state": "extra"}])):
76
+ self.assertEqual(self.client.post("/v1/systemone", json=invalid).status_code, 422)
77
+
78
+ def test_batch_preserves_state_order_and_rejects_invalid_shapes(self):
79
+ payload = {
80
+ "model": CANONICAL,
81
+ "states": [
82
+ {"id": "second", "state": "Please help"},
83
+ {"id": "first", "state": "Thank you"},
84
+ ],
85
+ "questions": questions(),
86
+ }
87
+ result = self.client.post("/v1/decision/batches", json=payload)
88
+ self.assertEqual(result.status_code, 200)
89
+ body = result.json()
90
+ self.assertEqual(set(body), {"model", "results", "usage"})
91
+ self.assertEqual(body["model"], CANONICAL)
92
+ self.assertEqual([row["id"] for row in body["results"]], ["second", "first"])
93
+ self.assertEqual(body["usage"]["input_tokens"], 16)
94
+ self.assertTrue(all(set(row) == {"id", "answers", "usage"} for row in body["results"]))
95
+
96
+ duplicate = dict(payload, states=[payload["states"][0]] * 2)
97
+ self.assertEqual(self.client.post("/v1/decision/batches", json=duplicate).status_code, 422)
98
+ self.assertEqual(self.client.post("/v1/decision/batches", json=dict(payload, state="extra")).status_code, 422)
99
+ self.assertEqual(self.client.post("/v1/decision/batches", json=dict(payload, model=MODEL)).status_code, 422)
100
+ self.assertEqual(self.client.post("/v1/systemone/batch", json=payload).status_code, 404)
101
+
102
+
103
+ if __name__ == "__main__":
104
+ unittest.main()
tests/test_tetris_arena.py ADDED
@@ -0,0 +1,747 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ import asyncio
2
+ import json
3
+ import unittest
4
+
5
+ from tetris_arena import (
6
+ LOCAL_MODELS,
7
+ ArenaUpstreamError,
8
+ ArenaQuotaError,
9
+ ArenaValidationError,
10
+ Competitor,
11
+ DecisionOutcome,
12
+ Endpoint,
13
+ HTTPDecisionAdapter,
14
+ MAX_UPSTREAM_BYTES,
15
+ RaceManager,
16
+ create_piece_sequence,
17
+ )
18
+
19
+ LEFT = Competitor("left-model", "Left", "test", "left-wire", "Left")
20
+ RIGHT = Competitor("right-model", "Right", "test", "right-wire", "Right")
21
+
22
+
23
+ class RecordingAdapter:
24
+ def __init__(
25
+ self, *, delay=0, invalid_for=None, invalid_at=None,
26
+ request_ms=None, inference_ms=None,
27
+ ):
28
+ self.delay = delay
29
+ self.invalid_for = invalid_for
30
+ self.invalid_at = invalid_at
31
+ self.request_ms = request_ms
32
+ self.inference_ms = inference_ms
33
+ self.calls = []
34
+
35
+ def catalog(self):
36
+ return (LEFT, RIGHT)
37
+
38
+ async def decide(self, competitor, payload, legal_choices):
39
+ criteria = tuple(payload["questions"]["placement"]["criteria"])
40
+ self.calls.append(
41
+ {
42
+ "competitor": competitor.id,
43
+ "state": payload["state"],
44
+ "model": payload.get("model"),
45
+ "criteria": criteria,
46
+ }
47
+ )
48
+ delay = (
49
+ self.delay.get(competitor.id, 0)
50
+ if isinstance(self.delay, dict)
51
+ else self.delay
52
+ )
53
+ if delay:
54
+ await asyncio.sleep(delay)
55
+ side_calls = sum(
56
+ call["competitor"] == competitor.id for call in self.calls
57
+ )
58
+ if competitor.id == self.invalid_for or (
59
+ competitor.id == LEFT.id and side_calls == self.invalid_at
60
+ ):
61
+ choice = "not-a-legal-placement"
62
+ else:
63
+ choice = criteria[0] if competitor.id == LEFT.id else criteria[-1]
64
+ request_ms = (
65
+ self.request_ms.get(competitor.id, delay * 1_000)
66
+ if isinstance(self.request_ms, dict)
67
+ else delay * 1_000
68
+ )
69
+ inference_ms = (
70
+ self.inference_ms.get(competitor.id)
71
+ if isinstance(self.inference_ms, dict)
72
+ else self.inference_ms
73
+ )
74
+ return DecisionOutcome(
75
+ choice=choice, request_ms=request_ms, inference_ms=inference_ms
76
+ )
77
+
78
+
79
+ class FakeResponse:
80
+ def __init__(
81
+ self, model, choice="r0-x0", *, chunks=None, encoding="identity",
82
+ inference_ms=None,
83
+ ):
84
+ self.is_redirect = False
85
+ self.status_code = 200
86
+ self.headers = {"content-encoding": encoding}
87
+ self.closed = False
88
+ self._document = {
89
+ "model": model,
90
+ "answers": {"placement": {"choice": choice}},
91
+ }
92
+ if inference_ms is not None:
93
+ self._document["timing"] = {"inference_ms": inference_ms}
94
+ self.chunks = chunks
95
+ self.yielded = 0
96
+
97
+ async def __aenter__(self):
98
+ return self
99
+
100
+ async def __aexit__(self, *_args):
101
+ self.closed = True
102
+
103
+ async def aiter_raw(self, *, chunk_size):
104
+ for chunk in self.chunks or (json.dumps(self._document).encode(),):
105
+ self.yielded += 1
106
+ yield chunk
107
+
108
+
109
+ class FakeHTTPClient:
110
+ def __init__(self, model, *, chunks=None, encoding="identity", inference_ms=None):
111
+ self.model = model
112
+ self.chunks = chunks
113
+ self.encoding = encoding
114
+ self.inference_ms = inference_ms
115
+ self.response = None
116
+
117
+ def stream(self, *args, **kwargs):
118
+ self.response = FakeResponse(
119
+ self.model, chunks=self.chunks, encoding=self.encoding,
120
+ inference_ms=self.inference_ms,
121
+ )
122
+ return self.response
123
+
124
+
125
+ class BlockingAdapter(RecordingAdapter):
126
+ def __init__(self):
127
+ super().__init__()
128
+ self.entered = asyncio.Event()
129
+
130
+ async def decide(self, competitor, payload, legal_choices):
131
+ self.entered.set()
132
+ await asyncio.sleep(60)
133
+ return DecisionOutcome(choice=next(iter(legal_choices)), request_ms=60_000)
134
+
135
+
136
+ class TetrisArenaTests(unittest.IsolatedAsyncioTestCase):
137
+ async def asyncSetUp(self):
138
+ self.managers = []
139
+
140
+ async def asyncTearDown(self):
141
+ await asyncio.gather(
142
+ *(manager.aclose() for manager in self.managers),
143
+ return_exceptions=True,
144
+ )
145
+
146
+ def manager(self, adapter, **kwargs):
147
+ manager = RaceManager(adapter, **kwargs)
148
+ self.managers.append(manager)
149
+ return manager
150
+
151
+ async def test_independent_requests_boards_actions_and_traces(self):
152
+ adapter = RecordingAdapter()
153
+ manager = self.manager(adapter)
154
+ race = await manager.create(
155
+ {
156
+ "left": LEFT.id,
157
+ "right": RIGHT.id,
158
+ "seed": 42,
159
+ "mode": "steps",
160
+ "max_steps": 6,
161
+ }
162
+ )
163
+
164
+ result = await asyncio.wait_for(race.wait(), 1)
165
+
166
+ self.assertEqual(result["results"]["left"]["pieces"], 6)
167
+ self.assertEqual(result["results"]["right"]["pieces"], 6)
168
+ self.assertTrue(result["summary"]["speed"]["comparable"])
169
+ self.assertEqual(result["states"]["left"]["step"], 6)
170
+ self.assertEqual(
171
+ result["states"]["left"]["game"]["board"],
172
+ race.traces["left"][-1]["board_after"],
173
+ )
174
+ left_calls = [call for call in adapter.calls if call["competitor"] == LEFT.id]
175
+ right_calls = [call for call in adapter.calls if call["competitor"] == RIGHT.id]
176
+ self.assertEqual(len(left_calls), 6)
177
+ self.assertEqual(len(right_calls), 6)
178
+ self.assertEqual(left_calls[0]["state"], right_calls[0]["state"])
179
+ self.assertNotEqual(left_calls[1]["state"], right_calls[1]["state"])
180
+ self.assertEqual(left_calls[0]["model"], "left-wire")
181
+ self.assertEqual(right_calls[0]["model"], "right-wire")
182
+
183
+ left_trace = race.traces["left"]
184
+ right_trace = race.traces["right"]
185
+ self.assertEqual(
186
+ [step["piece"] for step in left_trace],
187
+ [step["piece"] for step in right_trace],
188
+ "only the immutable piece sequence should be shared",
189
+ )
190
+ self.assertNotEqual(
191
+ [step["action"] for step in left_trace],
192
+ [step["action"] for step in right_trace],
193
+ )
194
+ self.assertNotEqual(left_trace[0]["board_after"], right_trace[0]["board_after"])
195
+ self.assertIsNot(left_trace, right_trace)
196
+ self.assertIsNone(race._orphan_task)
197
+
198
+ async def test_slow_renderer_does_not_gate_model_scheduling(self):
199
+ adapter = RecordingAdapter(delay=0.001)
200
+ manager = self.manager(adapter)
201
+ race = await manager.create(
202
+ {
203
+ "left": LEFT.id,
204
+ "right": RIGHT.id,
205
+ "seed": "render-independent",
206
+ "mode": "steps",
207
+ "max_steps": 8,
208
+ }
209
+ )
210
+
211
+ async def deliberately_slow_consumer():
212
+ async for event in race.iter_events():
213
+ if event is not None:
214
+ await asyncio.sleep(0.05)
215
+
216
+ consumer = asyncio.create_task(deliberately_slow_consumer())
217
+ try:
218
+ result = await asyncio.wait_for(race.wait(), 0.25)
219
+ finally:
220
+ consumer.cancel()
221
+ await asyncio.gather(consumer, return_exceptions=True)
222
+
223
+ self.assertEqual(len(adapter.calls), 16)
224
+ self.assertEqual(result["status"], "finished")
225
+ self.assertEqual(result["results"]["left"]["pieces"], 8)
226
+ self.assertEqual(result["results"]["right"]["pieces"], 8)
227
+
228
+ async def test_invalid_model_choice_is_not_rewritten_or_retried(self):
229
+ adapter = RecordingAdapter(invalid_for=LEFT.id)
230
+ manager = self.manager(adapter)
231
+ race = await manager.create(
232
+ {
233
+ "left": LEFT.id,
234
+ "right": RIGHT.id,
235
+ "seed": 7,
236
+ "mode": "steps",
237
+ "max_steps": 3,
238
+ }
239
+ )
240
+
241
+ result = await asyncio.wait_for(race.wait(), 1)
242
+
243
+ self.assertEqual(result["status"], "incomplete")
244
+ self.assertEqual(result["results"]["left"]["status"], "error")
245
+ self.assertEqual(result["results"]["left"]["reason"], "invalid_choice")
246
+ self.assertEqual(result["results"]["left"]["pieces"], 0)
247
+ self.assertEqual(race.traces["left"][0]["action"], "not-a-legal-placement")
248
+ self.assertEqual(
249
+ len([call for call in adapter.calls if call["competitor"] == LEFT.id]),
250
+ 1,
251
+ "an invalid choice must not trigger a hidden fallback request",
252
+ )
253
+ self.assertEqual(result["results"]["right"]["pieces"], 3)
254
+ self.assertEqual(result["summary"]["score"]["winner"], "unavailable")
255
+ self.assertFalse(result["summary"]["score"]["comparable"])
256
+ self.assertEqual(result["summary"]["speed"]["winner"], "unavailable")
257
+
258
+ async def test_partial_provider_failure_never_awards_a_score_winner(self):
259
+ race = await self.manager(RecordingAdapter(invalid_at=2)).create(
260
+ {
261
+ "left": LEFT.id,
262
+ "right": RIGHT.id,
263
+ "mode": "steps",
264
+ "max_steps": 3,
265
+ }
266
+ )
267
+
268
+ result = await asyncio.wait_for(race.wait(), 1)
269
+
270
+ self.assertEqual(result["status"], "incomplete")
271
+ self.assertEqual(result["results"]["left"]["pieces"], 1)
272
+ self.assertEqual(result["results"]["left"]["status"], "error")
273
+ self.assertEqual(result["results"]["right"]["status"], "finished")
274
+ self.assertEqual(result["summary"]["score"]["winner"], "unavailable")
275
+ self.assertEqual(result["summary"]["speed"]["winner"], "unavailable")
276
+ self.assertEqual(
277
+ [event["data"]["result"]["status"] for event in race.events
278
+ if event["type"] == "side.finished" and event["data"]["side"] == "left"],
279
+ ["error"],
280
+ )
281
+
282
+ async def test_first_failure_stops_a_pending_opponent(self):
283
+ adapter = RecordingAdapter(
284
+ delay={RIGHT.id: 60},
285
+ invalid_for=LEFT.id,
286
+ )
287
+ manager = self.manager(adapter, stream_grace_seconds=1)
288
+ race = await manager.create(
289
+ {
290
+ "left": LEFT.id,
291
+ "right": RIGHT.id,
292
+ "seed": 7,
293
+ "mode": "first_failure",
294
+ "max_steps": None,
295
+ }
296
+ )
297
+
298
+ result = await asyncio.wait_for(race.wait(), 1)
299
+
300
+ self.assertEqual(result["status"], "incomplete")
301
+ self.assertEqual(result["results"]["left"]["reason"], "invalid_choice")
302
+ self.assertEqual(
303
+ result["results"]["right"]["reason"],
304
+ "opponent_finished",
305
+ )
306
+ self.assertEqual(result["summary"]["score"]["winner"], "unavailable")
307
+
308
+ async def test_first_failure_guard_is_bounded_and_not_scored_as_a_win(self):
309
+ manager = self.manager(
310
+ RecordingAdapter(), max_steps_limit=2, first_failure_guard=2
311
+ )
312
+ race = await manager.create(
313
+ {"left": LEFT.id, "right": RIGHT.id, "mode": "first_failure"}
314
+ )
315
+
316
+ result = await asyncio.wait_for(race.wait(), 1)
317
+
318
+ self.assertEqual(result["status"], "incomplete")
319
+ self.assertEqual(
320
+ {side["reason"] for side in result["results"].values()},
321
+ {"operational_guard"},
322
+ )
323
+ self.assertEqual(race.provider_calls, 4)
324
+ self.assertEqual(result["summary"]["score"]["winner"], "unavailable")
325
+
326
+ async def test_speed_compares_server_request_e2e_even_when_runner_wall_time_differs(self):
327
+ adapter = RecordingAdapter(
328
+ delay={LEFT.id: 0.015, RIGHT.id: 0.001},
329
+ request_ms={LEFT.id: 2, RIGHT.id: 20},
330
+ inference_ms={LEFT.id: 100, RIGHT.id: 1},
331
+ )
332
+ race = await self.manager(adapter).create(
333
+ {
334
+ "left": LEFT.id,
335
+ "right": RIGHT.id,
336
+ "seed": 42,
337
+ "mode": "steps",
338
+ "max_steps": 4,
339
+ }
340
+ )
341
+
342
+ result = await asyncio.wait_for(race.wait(), 1)
343
+
344
+ self.assertEqual(result["summary"]["speed"]["winner"], "left")
345
+ self.assertEqual(result["summary"]["speed"]["percent"], 900.0)
346
+ self.assertEqual(
347
+ result["summary"]["speed"]["basis"],
348
+ "server_observed_request_to_response_e2e_total",
349
+ )
350
+ self.assertEqual(result["results"]["left"]["total_request_ms"], 8)
351
+ self.assertEqual(result["results"]["right"]["total_request_ms"], 80)
352
+ self.assertEqual(race.traces["left"][0]["provider_inference_ms"], 100)
353
+ self.assertGreater(
354
+ result["results"]["left"]["duration_ms"],
355
+ result["results"]["right"]["duration_ms"],
356
+ )
357
+
358
+ async def test_piece_sequence_is_seeded_bags(self):
359
+ first = create_piece_sequence("shared", 21)
360
+ second = create_piece_sequence("shared", 21)
361
+ different = create_piece_sequence("different", 21)
362
+ self.assertEqual(first, second)
363
+ self.assertNotEqual(first, different)
364
+ for offset in range(0, 21, 7):
365
+ self.assertEqual(set(first[offset : offset + 7]), set("IOTSZJL"))
366
+
367
+ async def test_local_catalog_contains_all_exact_models_with_short_display_names(
368
+ self,
369
+ ):
370
+ self.assertEqual(
371
+ [model.id for model in LOCAL_MODELS],
372
+ ["lux", "nox", "sol", "eos", "kai", "lex"],
373
+ )
374
+ self.assertTrue(
375
+ all(
376
+ model.request_model.startswith("llm-semantic-router/Decision-1.0-")
377
+ for model in LOCAL_MODELS
378
+ )
379
+ )
380
+ self.assertTrue(
381
+ all("Decision-1.0-" not in model.display_name for model in LOCAL_MODELS)
382
+ )
383
+
384
+ async def test_public_config_never_exposes_server_endpoints_or_credentials(self):
385
+ secret = "server-secret-canary"
386
+ private_url = "https://private-upstream.example.test/v1/systemone"
387
+ adapter = HTTPDecisionAdapter.from_environment(
388
+ {
389
+ "TETRIS_LOCAL_API_URL": private_url,
390
+ "TETRIS_JEV_API_URL": "https://jev.example.test/v1/systemone",
391
+ "TETRIS_JEV_API_KEY": secret,
392
+ "TETRIS_SYSTEMONE_API_URL": "https://systemone.example.test/api/v1/decide",
393
+ "TETRIS_SYSTEMONE_API_KEY": f"{secret}-mirror",
394
+ }
395
+ )
396
+ config = self.manager(adapter).public_config()
397
+ encoded = repr(config)
398
+
399
+ self.assertEqual(adapter._timeout_seconds, 45.0)
400
+ self.assertNotIn(secret, encoded)
401
+ self.assertNotIn(private_url, encoded)
402
+ self.assertTrue(all(item["ready"] for item in config["competitors"]))
403
+ self.assertEqual(len(config["competitors"]), 8)
404
+ by_id = {item["id"]: item for item in config["competitors"]}
405
+ self.assertEqual(by_id["jev-cloud-mirror"]["option_label"], "Jev Cloud")
406
+ self.assertEqual(by_id["jev-cloud-mirror"]["display_name"], "Jev Cloud")
407
+
408
+ async def test_decision_response_requires_exact_canonical_model_identity(self):
409
+ canonical = "llm-semantic-router/Decision-1.0-Lux-9B"
410
+ competitor = next(model for model in LOCAL_MODELS if model.id == "lux")
411
+ payload = {"state": "state", "questions": {}}
412
+ endpoint = Endpoint("https://local.example.test/v1/systemone", "", canonical)
413
+
414
+ accepted = HTTPDecisionAdapter(
415
+ endpoints={"lux": endpoint},
416
+ client=FakeHTTPClient(canonical),
417
+ )
418
+ outcome = await accepted.decide(competitor, payload, frozenset({"r0-x0"}))
419
+ self.assertEqual(outcome.provider_model, canonical)
420
+
421
+ mismatched = HTTPDecisionAdapter(
422
+ endpoints={"lux": endpoint},
423
+ client=FakeHTTPClient("llm-semantic-router/Decision-1.0-Kai-0.6B"),
424
+ )
425
+ with self.assertRaisesRegex(ArenaUpstreamError, "different model identity"):
426
+ await mismatched.decide(competitor, payload, frozenset({"r0-x0"}))
427
+
428
+ legacy_alias = HTTPDecisionAdapter(
429
+ endpoints={"lux": endpoint},
430
+ client=FakeHTTPClient("decision-lux"),
431
+ )
432
+ with self.assertRaisesRegex(ArenaUpstreamError, "different model identity"):
433
+ await legacy_alias.decide(competitor, payload, frozenset({"r0-x0"}))
434
+
435
+ async def test_request_latency_uses_server_clock_not_provider_inference_metric(self):
436
+ competitor = LOCAL_MODELS[0]
437
+ clock_ticks = iter((10.0, 10.037))
438
+ adapter = HTTPDecisionAdapter(
439
+ endpoints={
440
+ competitor.id: Endpoint(
441
+ "https://local.example.test/v1/systemone", "", competitor.request_model
442
+ )
443
+ },
444
+ client=FakeHTTPClient(competitor.request_model, inference_ms=1),
445
+ clock=lambda: next(clock_ticks),
446
+ )
447
+
448
+ outcome = await adapter.decide(
449
+ competitor, {"state": "state"}, frozenset({"r0-x0"})
450
+ )
451
+
452
+ self.assertAlmostEqual(outcome.request_ms, 37.0)
453
+ self.assertEqual(outcome.inference_ms, 1)
454
+
455
+ async def test_jev_latest_accepts_versioned_identity_but_not_another_family(self):
456
+ competitor = Competitor(
457
+ "jev-cloud", "Jev Cloud", "cloud", "jev-latest", "Jev Cloud"
458
+ )
459
+ endpoint = Endpoint(
460
+ "https://api.typesafe.ai/v1/systemone",
461
+ "secret",
462
+ "jev-latest",
463
+ "jev_alias_or_version",
464
+ )
465
+ payload = {"state": "state", "questions": {}}
466
+ accepted = HTTPDecisionAdapter(
467
+ endpoints={"jev-cloud": endpoint},
468
+ client=FakeHTTPClient("jev-1.13.0"),
469
+ )
470
+ outcome = await accepted.decide(competitor, payload, frozenset({"r0-x0"}))
471
+ self.assertEqual(outcome.provider_model, "jev-1.13.0")
472
+
473
+ rejected = HTTPDecisionAdapter(
474
+ endpoints={"jev-cloud": endpoint},
475
+ client=FakeHTTPClient("unrelated-1.0.0"),
476
+ )
477
+ with self.assertRaisesRegex(ArenaUpstreamError, "unsupported model identity"):
478
+ await rejected.decide(competitor, payload, frozenset({"r0-x0"}))
479
+
480
+ pinned_endpoint = Endpoint(
481
+ "https://api.typesafe.ai/v1/systemone",
482
+ "secret",
483
+ "jev-1.12.0",
484
+ "jev_alias_or_version",
485
+ )
486
+ pinned = HTTPDecisionAdapter(
487
+ endpoints={"jev-cloud": pinned_endpoint},
488
+ client=FakeHTTPClient("jev-1.13.0"),
489
+ )
490
+ with self.assertRaisesRegex(ArenaUpstreamError, "different pinned model"):
491
+ await pinned.decide(competitor, payload, frozenset({"r0-x0"}))
492
+
493
+ async def test_upstream_stream_stops_at_byte_limit_and_closes(self):
494
+ canonical = LOCAL_MODELS[0].request_model
495
+ chunks = [b"x" * (64 * 1024)] * (MAX_UPSTREAM_BYTES // (64 * 1024) + 2)
496
+ client = FakeHTTPClient(canonical, chunks=chunks)
497
+ adapter = HTTPDecisionAdapter(
498
+ endpoints={"lux": Endpoint("https://local.example.test/v1/systemone", "", canonical)},
499
+ client=client,
500
+ )
501
+ with self.assertRaisesRegex(ArenaUpstreamError, "oversized response"):
502
+ await adapter.decide(LOCAL_MODELS[0], {"state": "state"}, frozenset({"r0-x0"}))
503
+ self.assertTrue(client.response.closed)
504
+ self.assertEqual(client.response.yielded, MAX_UPSTREAM_BYTES // (64 * 1024) + 1)
505
+
506
+ async def test_client_active_rate_and_provider_call_budgets(self):
507
+ adapter = BlockingAdapter()
508
+ manager = self.manager(
509
+ adapter,
510
+ client_races_per_minute=2,
511
+ client_calls_per_window=240,
512
+ stream_grace_seconds=1,
513
+ )
514
+ payload = {"left": LEFT.id, "right": RIGHT.id, "mode": "steps", "max_steps": 120}
515
+ first = await manager.create(payload, client_id="client-a")
516
+ await adapter.entered.wait()
517
+ with self.assertRaisesRegex(ArenaQuotaError, "already active"):
518
+ await manager.create(payload, client_id="client-a")
519
+ await first.cancel()
520
+ self.assertEqual(first.provider_calls, 2)
521
+ with self.assertRaisesRegex(ArenaQuotaError, "Provider call budget"):
522
+ await manager.create(payload, client_id="client-a")
523
+ second = await manager.create(
524
+ {**payload, "max_steps": 119}, client_id="client-a"
525
+ )
526
+ with self.assertRaisesRegex(ArenaQuotaError, "rate limit"):
527
+ await manager.create({**payload, "max_steps": 1}, client_id="client-a")
528
+ other = await manager.create(
529
+ {**payload, "max_steps": 1}, client_id="client-b"
530
+ )
531
+ await asyncio.gather(second.cancel(), other.cancel())
532
+
533
+ async def test_configured_long_run_limit_scales_call_budget_and_releases_history(self):
534
+ adapter = BlockingAdapter()
535
+ manager = RaceManager.from_environment(
536
+ adapter,
537
+ {
538
+ "TETRIS_MAX_STEPS": "2000",
539
+ "TETRIS_FIRST_FAILURE_GUARD": "1800",
540
+ },
541
+ )
542
+ self.managers.append(manager)
543
+
544
+ self.assertEqual(manager.public_config()["limits"], {
545
+ "max_steps": 2000,
546
+ "first_failure_guard": 1800,
547
+ })
548
+ self.assertEqual(manager.client_calls_per_window, 4000)
549
+ race = await manager.create(
550
+ {"left": LEFT.id, "right": RIGHT.id, "mode": "steps", "max_steps": 2000},
551
+ client_id="long-run",
552
+ )
553
+ await adapter.entered.wait()
554
+ self.assertEqual(race.call_budget, 4000)
555
+ self.assertEqual(race.history_limit, 4008)
556
+ await race.cancel()
557
+ manager._prune_admissions_locked()
558
+ admission = manager._client_history["long-run"][0]
559
+ self.assertIsNone(admission.session)
560
+ self.assertGreaterEqual(admission.provider_calls, 1)
561
+ self.assertLessEqual(admission.provider_calls, 2)
562
+
563
+ with self.assertRaisesRegex(ArenaValidationError, "between 1 and 2000"):
564
+ await manager.create(
565
+ {"left": LEFT.id, "right": RIGHT.id, "mode": "steps", "max_steps": 2001}
566
+ )
567
+ with self.assertRaisesRegex(ValueError, "TETRIS_CLIENT_CALLS_PER_WINDOW"):
568
+ RaceManager.from_environment(
569
+ RecordingAdapter(),
570
+ {"TETRIS_MAX_STEPS": "2000", "TETRIS_CLIENT_CALLS_PER_WINDOW": "3999"},
571
+ )
572
+
573
+ async def test_immediate_cancel_finalizes_a_not_yet_started_race(self):
574
+ manager = self.manager(RecordingAdapter())
575
+ race = await manager.create(
576
+ {
577
+ "left": LEFT.id,
578
+ "right": RIGHT.id,
579
+ "mode": "steps",
580
+ "max_steps": 3,
581
+ }
582
+ )
583
+
584
+ await race.cancel("user_cancelled")
585
+
586
+ self.assertTrue(race._done.is_set())
587
+ self.assertEqual(race.status, "cancelled")
588
+ self.assertIsNotNone(race.completed_at)
589
+ self.assertEqual((await race.wait())["status"], "cancelled")
590
+
591
+ async def test_race_deadline_stops_both_model_runners(self):
592
+ adapter = BlockingAdapter()
593
+ manager = self.manager(
594
+ adapter,
595
+ race_timeout_seconds=0.02,
596
+ stream_grace_seconds=1,
597
+ )
598
+ race = await manager.create(
599
+ {
600
+ "left": LEFT.id,
601
+ "right": RIGHT.id,
602
+ "mode": "steps",
603
+ "max_steps": 20,
604
+ }
605
+ )
606
+
607
+ result = await asyncio.wait_for(race.wait(), 1)
608
+
609
+ self.assertEqual(result["status"], "timed_out")
610
+ self.assertEqual(result["summary"]["status"], "timed_out")
611
+ self.assertEqual(result["summary"]["score"]["winner"], "unavailable")
612
+ self.assertEqual(
613
+ {item["reason"] for item in result["results"].values()},
614
+ {"race_timeout"},
615
+ )
616
+
617
+ async def test_stream_grace_cancels_an_abandoned_race(self):
618
+ adapter = BlockingAdapter()
619
+ manager = self.manager(
620
+ adapter,
621
+ race_timeout_seconds=1,
622
+ stream_grace_seconds=0.02,
623
+ )
624
+ race = await manager.create(
625
+ {
626
+ "left": LEFT.id,
627
+ "right": RIGHT.id,
628
+ "mode": "steps",
629
+ "max_steps": 20,
630
+ }
631
+ )
632
+
633
+ result = await asyncio.wait_for(race.wait(), 1)
634
+
635
+ self.assertEqual(result["status"], "cancelled")
636
+ self.assertEqual(
637
+ {item["reason"] for item in result["results"].values()},
638
+ {"stream_disconnected"},
639
+ )
640
+
641
+ async def test_connected_event_stream_holds_the_orphan_lease(self):
642
+ adapter = BlockingAdapter()
643
+ manager = self.manager(
644
+ adapter,
645
+ race_timeout_seconds=1,
646
+ stream_grace_seconds=0.02,
647
+ )
648
+ race = await manager.create(
649
+ {
650
+ "left": LEFT.id,
651
+ "right": RIGHT.id,
652
+ "mode": "steps",
653
+ "max_steps": 20,
654
+ }
655
+ )
656
+ stream = race.iter_events()
657
+ await anext(stream)
658
+
659
+ await asyncio.sleep(0.05)
660
+ self.assertFalse(race._done.is_set())
661
+
662
+ await stream.aclose()
663
+ result = await asyncio.wait_for(race.wait(), 1)
664
+ self.assertEqual(result["status"], "cancelled")
665
+
666
+ async def test_manager_close_cancels_and_awaits_active_sessions(self):
667
+ adapter = BlockingAdapter()
668
+ manager = self.manager(adapter, stream_grace_seconds=1)
669
+ race = await manager.create(
670
+ {
671
+ "left": LEFT.id,
672
+ "right": RIGHT.id,
673
+ "mode": "steps",
674
+ "max_steps": 20,
675
+ }
676
+ )
677
+ await adapter.entered.wait()
678
+
679
+ await manager.aclose()
680
+
681
+ self.assertEqual(race.status, "cancelled")
682
+ self.assertTrue(race._done.is_set())
683
+
684
+ async def test_concurrent_manager_close_waits_for_the_same_shutdown(self):
685
+ adapter = BlockingAdapter()
686
+ manager = self.manager(adapter, stream_grace_seconds=1)
687
+ race = await manager.create(
688
+ {
689
+ "left": LEFT.id,
690
+ "right": RIGHT.id,
691
+ "mode": "steps",
692
+ "max_steps": 20,
693
+ }
694
+ )
695
+ await adapter.entered.wait()
696
+
697
+ first = asyncio.create_task(manager.aclose())
698
+ second = asyncio.create_task(manager.aclose())
699
+ await asyncio.gather(first, second)
700
+
701
+ self.assertTrue(race._done.is_set())
702
+ self.assertEqual(race.status, "cancelled")
703
+ self.assertEqual(manager.sessions, {})
704
+
705
+ async def test_completed_sessions_expire_on_access(self):
706
+ manager = self.manager(
707
+ RecordingAdapter(),
708
+ completed_ttl_seconds=0.01,
709
+ )
710
+ race = await manager.create(
711
+ {
712
+ "left": LEFT.id,
713
+ "right": RIGHT.id,
714
+ "mode": "steps",
715
+ "max_steps": 1,
716
+ }
717
+ )
718
+ await race.wait()
719
+ await asyncio.sleep(0.02)
720
+
721
+ with self.assertRaises(KeyError):
722
+ manager.get(race.id)
723
+
724
+ async def test_race_lifecycle_limits_are_explicit_environment_settings(self):
725
+ manager = RaceManager.from_environment(
726
+ RecordingAdapter(),
727
+ {
728
+ "TETRIS_RACE_TIMEOUT_SECONDS": "120",
729
+ "TETRIS_STREAM_GRACE_SECONDS": "12",
730
+ "TETRIS_COMPLETED_TTL_SECONDS": "60",
731
+ },
732
+ )
733
+ self.managers.append(manager)
734
+
735
+ self.assertEqual(manager.race_timeout_seconds, 120)
736
+ self.assertEqual(manager.stream_grace_seconds, 12)
737
+ self.assertEqual(manager.completed_ttl_seconds, 60)
738
+
739
+ with self.assertRaisesRegex(ValueError, "TETRIS_RACE_TIMEOUT_SECONDS"):
740
+ RaceManager.from_environment(
741
+ RecordingAdapter(),
742
+ {"TETRIS_RACE_TIMEOUT_SECONDS": "unbounded"},
743
+ )
744
+
745
+
746
+ if __name__ == "__main__":
747
+ unittest.main()
tests/test_tetris_assets.py ADDED
@@ -0,0 +1,42 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ import unittest
2
+ from pathlib import Path
3
+
4
+ ROOT = Path(__file__).resolve().parents[1]
5
+
6
+
7
+ class TetrisAssetTests(unittest.TestCase):
8
+ def test_docker_context_contains_tetris_server_module(self):
9
+ dockerignore = (ROOT / ".dockerignore").read_text().splitlines()
10
+ dockerfile = (ROOT / "Dockerfile").read_text()
11
+
12
+ self.assertIn("!tetris_arena.py", dockerignore)
13
+ self.assertIn("tetris_arena.py", dockerfile)
14
+
15
+ def test_frontend_has_a_single_flight_race_lifecycle(self):
16
+ source = (ROOT / "static" / "tetris" / "app.js").read_text()
17
+
18
+ for state in ("idle", "starting", "running", "stopping"):
19
+ self.assertIn(state, source)
20
+ self.assertIn("if (lifecycle !== 'idle') return", source)
21
+ self.assertIn("lifecycle === 'starting' || lifecycle === 'stopping'", source)
22
+ self.assertIn("expectedOperation !== operation", source)
23
+ self.assertIn("createController.signal", source)
24
+
25
+ def test_frontend_uses_server_limit_and_recovers_terminal_snapshots(self):
26
+ source = (ROOT / "static" / "tetris" / "app.js").read_text()
27
+ page = (ROOT / "static" / "tetris" / "index.html").read_text()
28
+
29
+ self.assertIn('elements.customSteps.max = String(maximum)', source)
30
+ self.assertIn('option.disabled = Number(option.value) > maximum', source)
31
+ self.assertIn('connectEventStream(race.events_url, race.result_url', source)
32
+ self.assertIn("async function recoverSnapshot()", source)
33
+ self.assertIn("applySnapshot(snapshot)", source)
34
+ self.assertIn("data.step <= seenSteps[data.side]", source)
35
+ self.assertIn("finishRace(snapshot.status, snapshot.summary, stream)", source)
36
+ self.assertIn("error: 'Failed'", source)
37
+ self.assertIn('max="120"', page)
38
+ self.assertEqual(page.count('data-result-status'), 2)
39
+
40
+
41
+ if __name__ == "__main__":
42
+ unittest.main()
tests/test_tetris_routes.py ADDED
@@ -0,0 +1,138 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ import asyncio
2
+ import json
3
+ import unittest
4
+
5
+ from fastapi.testclient import TestClient
6
+
7
+ from app import create_app
8
+ from tetris_arena import Competitor, DecisionOutcome, RaceManager
9
+
10
+ LEFT = Competitor("route-left", "Route Left", "test", "left-wire", "Route Left")
11
+ RIGHT = Competitor("route-right", "Route Right", "test", "right-wire", "Route Right")
12
+
13
+
14
+ class RouteAdapter:
15
+ def catalog(self):
16
+ return (LEFT, RIGHT)
17
+
18
+ async def decide(self, competitor, payload, legal_choices):
19
+ choices = tuple(payload["questions"]["placement"]["criteria"])
20
+ choice = choices[0] if competitor.id == LEFT.id else choices[-1]
21
+ return DecisionOutcome(choice, 1.25, competitor.request_model)
22
+
23
+
24
+ class SlowRouteAdapter(RouteAdapter):
25
+ async def decide(self, competitor, payload, legal_choices):
26
+ await asyncio.sleep(60)
27
+ return DecisionOutcome(
28
+ next(iter(legal_choices)), 60_000, competitor.request_model
29
+ )
30
+
31
+
32
+ class TetrisRouteTests(unittest.TestCase):
33
+ def setUp(self):
34
+ self.manager = RaceManager(RouteAdapter())
35
+ self.client_context = TestClient(create_app(tetris_manager=self.manager))
36
+ self.client = self.client_context.__enter__()
37
+
38
+ def tearDown(self):
39
+ self.client_context.portal.call(self.manager.aclose)
40
+ self.client_context.__exit__(None, None, None)
41
+
42
+ def test_static_page_config_sse_and_trace_are_integrated(self):
43
+ page = self.client.get("/tetris/")
44
+ self.assertEqual(page.status_code, 200)
45
+ self.assertIn("Decision 1.0 · Tetris Race", page.text)
46
+
47
+ config = self.client.get("/api/tetris/config")
48
+ self.assertEqual(config.status_code, 200)
49
+ self.assertEqual(
50
+ [item["id"] for item in config.json()["competitors"]],
51
+ [LEFT.id, RIGHT.id],
52
+ )
53
+
54
+ created = self.client.post(
55
+ "/api/tetris/races",
56
+ json={
57
+ "left": LEFT.id,
58
+ "right": RIGHT.id,
59
+ "seed": 42,
60
+ "mode": "steps",
61
+ "max_steps": 3,
62
+ },
63
+ )
64
+ self.assertEqual(created.status_code, 201)
65
+ document = created.json()
66
+ with self.client.stream("GET", document["events_url"]) as response:
67
+ self.assertEqual(response.status_code, 200)
68
+ stream = "".join(response.iter_text())
69
+ self.assertIn("event: side.step", stream)
70
+ self.assertIn("event: race.finished", stream)
71
+ event_ids = [
72
+ int(line.removeprefix("id: "))
73
+ for line in stream.splitlines()
74
+ if line.startswith("id: ")
75
+ ]
76
+ resumed = self.client.get(
77
+ document["events_url"],
78
+ headers={"Last-Event-ID": str(event_ids[-2])},
79
+ )
80
+ self.assertEqual(resumed.status_code, 200)
81
+ self.assertEqual(
82
+ [line for line in resumed.text.splitlines() if line.startswith("id: ")],
83
+ [f"id: {event_ids[-1]}"],
84
+ )
85
+
86
+ trace = self.client.get(f"/api/tetris/races/{document['id']}/trace")
87
+ self.assertEqual(trace.status_code, 200)
88
+ payload = trace.json()
89
+ self.assertEqual(payload["status"], "finished")
90
+ self.assertEqual(len(payload["traces"]["left"]), 3)
91
+ self.assertEqual(len(payload["traces"]["right"]), 3)
92
+ self.assertEqual(payload["states"]["left"]["step"], 3)
93
+ self.assertEqual(
94
+ payload["states"]["left"]["game"]["board"],
95
+ payload["traces"]["left"][-1]["board_after"],
96
+ )
97
+ self.assertEqual(payload["summary"]["status"], "finished")
98
+ self.assertNotEqual(
99
+ payload["traces"]["left"][0]["action"],
100
+ payload["traces"]["right"][0]["action"],
101
+ )
102
+ self.assertNotIn("Authorization", json.dumps(payload))
103
+
104
+ def test_unconfigured_or_invalid_race_is_rejected(self):
105
+ response = self.client.post(
106
+ "/api/tetris/races",
107
+ json={
108
+ "left": "missing",
109
+ "right": RIGHT.id,
110
+ "mode": "steps",
111
+ "max_steps": 3,
112
+ },
113
+ )
114
+ self.assertEqual(response.status_code, 422)
115
+
116
+ def test_app_shutdown_closes_its_manager_and_active_races(self):
117
+ api = create_app(tetris_adapter=SlowRouteAdapter())
118
+ with TestClient(api) as client:
119
+ created = client.post(
120
+ "/api/tetris/races",
121
+ json={
122
+ "left": LEFT.id,
123
+ "right": RIGHT.id,
124
+ "mode": "steps",
125
+ "max_steps": 20,
126
+ },
127
+ )
128
+ self.assertEqual(created.status_code, 201)
129
+ race = api.state.tetris.get(created.json()["id"])
130
+
131
+ self.assertTrue(api.state.tetris._closing)
132
+ self.assertTrue(race._done.is_set())
133
+ self.assertEqual(race.status, "cancelled")
134
+ self.assertEqual(api.state.tetris.sessions, {})
135
+
136
+
137
+ if __name__ == "__main__":
138
+ unittest.main()
tetris_arena.py ADDED
@@ -0,0 +1,1673 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Server-authoritative Decision 1.0 Tetris races.
2
+
3
+ The browser is deliberately a passive renderer. Each competitor owns an
4
+ independent game, request loop, and trace. The only shared race input is the
5
+ immutable tetromino sequence generated from the requested seed.
6
+ """
7
+
8
+ from __future__ import annotations
9
+
10
+ import asyncio
11
+ import copy
12
+ import hashlib
13
+ import json
14
+ import math
15
+ import os
16
+ import random
17
+ import re
18
+ import secrets
19
+ import time
20
+ from collections import deque
21
+ from collections.abc import AsyncIterator, Mapping
22
+ from dataclasses import dataclass, field
23
+ from typing import Any
24
+ from urllib.parse import urlsplit
25
+
26
+ BOARD_WIDTH = 10
27
+ BOARD_HEIGHT = 20
28
+ DEFAULT_MAX_STEPS = 40
29
+ DEFAULT_MAX_CONFIGURED_STEPS = 120
30
+ MAX_CONFIGURED_STEPS = 2_000
31
+ MAX_PROVIDER_CALLS_PER_RACE = 2 * MAX_CONFIGURED_STEPS
32
+ MAX_CONCURRENT_RACES = 4
33
+ DEFAULT_CLIENT_ACTIVE_RACES = 1
34
+ DEFAULT_CLIENT_RACES_PER_MINUTE = 6
35
+ MAX_CLIENT_RACES_PER_MINUTE = 30
36
+ DEFAULT_CLIENT_WINDOW_SECONDS = 60
37
+ DEFAULT_CLIENT_CALLS_PER_WINDOW = 480
38
+ DEFAULT_CLIENT_CALL_WINDOW_SECONDS = 10 * 60
39
+ MAX_ADMISSION_HISTORY = 4096
40
+ MAX_RACE_HISTORY = MAX_PROVIDER_CALLS_PER_RACE + 8
41
+ MAX_UPSTREAM_BYTES = 1_048_576
42
+ MAX_RETAINED_RACES = 16
43
+ DEFAULT_RACE_TIMEOUT_SECONDS = 15 * 60
44
+ DEFAULT_STREAM_GRACE_SECONDS = 30
45
+ DEFAULT_COMPLETED_TTL_SECONDS = 5 * 60
46
+
47
+ _JEV_VERSION = re.compile(r"jev-\d+\.\d+\.\d+(?:[-+][0-9A-Za-z.-]+)?")
48
+
49
+ PIECE_NAMES = ("I", "O", "T", "S", "Z", "J", "L")
50
+ PIECES: dict[str, tuple[tuple[tuple[int, int], ...], ...]] = {
51
+ "I": (
52
+ ((0, 1), (1, 1), (2, 1), (3, 1)),
53
+ ((2, 0), (2, 1), (2, 2), (2, 3)),
54
+ ((0, 2), (1, 2), (2, 2), (3, 2)),
55
+ ((1, 0), (1, 1), (1, 2), (1, 3)),
56
+ ),
57
+ "O": (((1, 0), (2, 0), (1, 1), (2, 1)),),
58
+ "T": (
59
+ ((1, 0), (0, 1), (1, 1), (2, 1)),
60
+ ((1, 0), (1, 1), (2, 1), (1, 2)),
61
+ ((0, 1), (1, 1), (2, 1), (1, 2)),
62
+ ((1, 0), (0, 1), (1, 1), (1, 2)),
63
+ ),
64
+ "S": (
65
+ ((1, 0), (2, 0), (0, 1), (1, 1)),
66
+ ((1, 0), (1, 1), (2, 1), (2, 2)),
67
+ ((1, 1), (2, 1), (0, 2), (1, 2)),
68
+ ((0, 0), (0, 1), (1, 1), (1, 2)),
69
+ ),
70
+ "Z": (
71
+ ((0, 0), (1, 0), (1, 1), (2, 1)),
72
+ ((2, 0), (1, 1), (2, 1), (1, 2)),
73
+ ((0, 1), (1, 1), (1, 2), (2, 2)),
74
+ ((1, 0), (0, 1), (1, 1), (0, 2)),
75
+ ),
76
+ "J": (
77
+ ((0, 0), (0, 1), (1, 1), (2, 1)),
78
+ ((1, 0), (2, 0), (1, 1), (1, 2)),
79
+ ((0, 1), (1, 1), (2, 1), (2, 2)),
80
+ ((1, 0), (1, 1), (0, 2), (1, 2)),
81
+ ),
82
+ "L": (
83
+ ((2, 0), (0, 1), (1, 1), (2, 1)),
84
+ ((1, 0), (1, 1), (1, 2), (2, 2)),
85
+ ((0, 1), (1, 1), (2, 1), (0, 2)),
86
+ ((0, 0), (1, 0), (1, 1), (1, 2)),
87
+ ),
88
+ }
89
+ LINE_SCORES = (0, 100, 300, 500, 800)
90
+
91
+
92
+ @dataclass(frozen=True)
93
+ class Competitor:
94
+ id: str
95
+ display_name: str
96
+ family: str
97
+ request_model: str | None
98
+ option_label: str
99
+ ready: bool = True
100
+
101
+ def public(self) -> dict[str, Any]:
102
+ return {
103
+ "id": self.id,
104
+ "display_name": self.display_name,
105
+ "option_label": self.option_label,
106
+ "family": self.family,
107
+ "ready": self.ready,
108
+ }
109
+
110
+
111
+ LOCAL_MODELS: tuple[Competitor, ...] = (
112
+ Competitor(
113
+ "lux",
114
+ "Lux-9B",
115
+ "decision",
116
+ "llm-semantic-router/Decision-1.0-Lux-9B",
117
+ "Lux-9B",
118
+ ),
119
+ Competitor(
120
+ "nox",
121
+ "Nox-4B",
122
+ "decision",
123
+ "llm-semantic-router/Decision-1.0-Nox-4B",
124
+ "Nox-4B",
125
+ ),
126
+ Competitor(
127
+ "sol",
128
+ "Sol-2B",
129
+ "decision",
130
+ "llm-semantic-router/Decision-1.0-Sol-2B",
131
+ "Sol-2B",
132
+ ),
133
+ Competitor(
134
+ "eos",
135
+ "Eos-0.8B",
136
+ "decision",
137
+ "llm-semantic-router/Decision-1.0-Eos-0.8B",
138
+ "Eos-0.8B",
139
+ ),
140
+ Competitor(
141
+ "kai",
142
+ "Kai-0.6B",
143
+ "decision",
144
+ "llm-semantic-router/Decision-1.0-Kai-0.6B",
145
+ "Kai-0.6B",
146
+ ),
147
+ Competitor(
148
+ "lex",
149
+ "Lex-0.6B",
150
+ "decision",
151
+ "llm-semantic-router/Decision-1.0-Lex-0.6B",
152
+ "Lex-0.6B",
153
+ ),
154
+ )
155
+
156
+
157
+ class ArenaValidationError(ValueError):
158
+ """A user-visible race request error."""
159
+
160
+ status_code = 422
161
+ retry_after: int | None = None
162
+
163
+
164
+ class ArenaQuotaError(ArenaValidationError):
165
+ """A bounded public-demo admission limit was reached."""
166
+
167
+ status_code = 429
168
+
169
+ def __init__(self, message: str, *, retry_after: int):
170
+ super().__init__(message)
171
+ self.retry_after = retry_after
172
+
173
+
174
+ class ArenaUpstreamError(RuntimeError):
175
+ """A sanitized decision-provider error."""
176
+
177
+
178
+ @dataclass(frozen=True)
179
+ class Placement:
180
+ id: str
181
+ piece: str
182
+ rotation: int
183
+ x: int
184
+ y: int
185
+ cells: tuple[tuple[int, int], ...]
186
+ board: tuple[str, ...]
187
+ cleared_rows: tuple[int, ...]
188
+ top_out: bool
189
+ holes: int
190
+ max_height: int
191
+
192
+
193
+ @dataclass
194
+ class Game:
195
+ board: list[list[str | None]] = field(
196
+ default_factory=lambda: [
197
+ [None for _ in range(BOARD_WIDTH)] for _ in range(BOARD_HEIGHT)
198
+ ]
199
+ )
200
+ score: int = 0
201
+ lines: int = 0
202
+ pieces: int = 0
203
+ game_over: bool = False
204
+
205
+ def rows(self) -> list[str]:
206
+ return ["".join(cell or "." for cell in row) for row in self.board]
207
+
208
+ def public(self) -> dict[str, Any]:
209
+ return {
210
+ "board": self.rows(),
211
+ "score": self.score,
212
+ "lines": self.lines,
213
+ "pieces": self.pieces,
214
+ "game_over": self.game_over,
215
+ }
216
+
217
+
218
+ @dataclass(frozen=True)
219
+ class DecisionOutcome:
220
+ choice: str
221
+ request_ms: float
222
+ provider_model: str | None = None
223
+ inference_ms: float | None = None
224
+
225
+
226
+ @dataclass(frozen=True)
227
+ class Endpoint:
228
+ url: str
229
+ api_key: str = field(repr=False)
230
+ model: str | None = None
231
+ identity_policy: str = "exact"
232
+
233
+
234
+ def _seed_number(seed: str | int) -> int:
235
+ if isinstance(seed, int) and not isinstance(seed, bool):
236
+ return seed
237
+ digest = hashlib.sha256(str(seed).encode("utf-8")).digest()
238
+ return int.from_bytes(digest[:8], "big")
239
+
240
+
241
+ def create_piece_sequence(seed: str | int, count: int) -> tuple[str, ...]:
242
+ """Create deterministic shuffled seven-piece bags."""
243
+ if not isinstance(count, int) or isinstance(count, bool) or count < 0:
244
+ raise ValueError("count must be a non-negative integer")
245
+ generator = random.Random(_seed_number(seed))
246
+ result: list[str] = []
247
+ while len(result) < count:
248
+ bag = list(PIECE_NAMES)
249
+ generator.shuffle(bag)
250
+ result.extend(bag)
251
+ return tuple(result[:count])
252
+
253
+
254
+ def _can_place(
255
+ board: list[list[str | None]],
256
+ shape: tuple[tuple[int, int], ...],
257
+ x: int,
258
+ y: int,
259
+ ) -> bool:
260
+ for cell_x, cell_y in shape:
261
+ absolute_x = x + cell_x
262
+ absolute_y = y + cell_y
263
+ if absolute_x < 0 or absolute_x >= BOARD_WIDTH or absolute_y >= BOARD_HEIGHT:
264
+ return False
265
+ if absolute_y >= 0 and board[absolute_y][absolute_x] is not None:
266
+ return False
267
+ return True
268
+
269
+
270
+ def _board_stats(board: list[list[str | None]]) -> tuple[int, int]:
271
+ holes = 0
272
+ heights: list[int] = []
273
+ for x in range(BOARD_WIDTH):
274
+ first = next((y for y in range(BOARD_HEIGHT) if board[y][x] is not None), None)
275
+ if first is None:
276
+ heights.append(0)
277
+ continue
278
+ heights.append(BOARD_HEIGHT - first)
279
+ holes += sum(1 for y in range(first + 1, BOARD_HEIGHT) if board[y][x] is None)
280
+ return holes, max(heights, default=0)
281
+
282
+
283
+ def _preview_placement(
284
+ game: Game,
285
+ piece: str,
286
+ rotation: int,
287
+ x: int,
288
+ y: int,
289
+ ) -> Placement:
290
+ shape = PIECES[piece][rotation]
291
+ cells = tuple(sorted((x + cell_x, y + cell_y) for cell_x, cell_y in shape))
292
+ board = [row[:] for row in game.board]
293
+ top_out = any(cell_y < 0 for _, cell_y in cells)
294
+ for cell_x, cell_y in cells:
295
+ if cell_y >= 0:
296
+ board[cell_y][cell_x] = piece
297
+ cleared_rows = tuple(
298
+ row_index
299
+ for row_index, row in enumerate(board)
300
+ if all(cell is not None for cell in row)
301
+ )
302
+ if cleared_rows:
303
+ kept = [row for index, row in enumerate(board) if index not in cleared_rows]
304
+ board = [[None] * BOARD_WIDTH for _ in cleared_rows] + kept
305
+ holes, max_height = _board_stats(board)
306
+ return Placement(
307
+ id=f"r{rotation}-x{x}",
308
+ piece=piece,
309
+ rotation=rotation,
310
+ x=x,
311
+ y=y,
312
+ cells=cells,
313
+ board=tuple("".join(cell or "." for cell in row) for row in board),
314
+ cleared_rows=cleared_rows,
315
+ top_out=top_out,
316
+ holes=holes,
317
+ max_height=max_height,
318
+ )
319
+
320
+
321
+ def enumerate_placements(game: Game, piece: str) -> tuple[Placement, ...]:
322
+ if piece not in PIECES:
323
+ raise ValueError("unknown tetromino")
324
+ placements: list[Placement] = []
325
+ seen: set[tuple[tuple[int, int], ...]] = set()
326
+ for rotation, shape in enumerate(PIECES[piece]):
327
+ minimum_x = min(x for x, _ in shape)
328
+ maximum_x = max(x for x, _ in shape)
329
+ for x in range(-minimum_x, BOARD_WIDTH - maximum_x):
330
+ y = -4
331
+ while _can_place(game.board, shape, x, y + 1):
332
+ y += 1
333
+ if not _can_place(game.board, shape, x, y):
334
+ continue
335
+ absolute_cells = tuple(
336
+ sorted((x + cell_x, y + cell_y) for cell_x, cell_y in shape)
337
+ )
338
+ if absolute_cells in seen:
339
+ continue
340
+ seen.add(absolute_cells)
341
+ placements.append(_preview_placement(game, piece, rotation, x, y))
342
+ return tuple(placements)
343
+
344
+
345
+ def apply_placement(game: Game, placement: Placement) -> Game:
346
+ board = [[None if cell == "." else cell for cell in row] for row in placement.board]
347
+ cleared = len(placement.cleared_rows)
348
+ return Game(
349
+ board=board,
350
+ score=game.score + LINE_SCORES[min(cleared, 4)],
351
+ lines=game.lines + cleared,
352
+ pieces=game.pieces + 1,
353
+ game_over=placement.top_out,
354
+ )
355
+
356
+
357
+ def build_decision_request(
358
+ game: Game,
359
+ piece: str,
360
+ next_piece: str | None,
361
+ placements: tuple[Placement, ...],
362
+ model: str | None,
363
+ ) -> tuple[dict[str, Any], str]:
364
+ rows = game.rows()
365
+ state = "\n".join(
366
+ (
367
+ (
368
+ f"piece={piece}; next={next_piece or 'none'}; score={game.score}; "
369
+ f"lines={game.lines}; pieces={game.pieces}"
370
+ ),
371
+ "board (top to bottom, .=empty):",
372
+ *rows,
373
+ )
374
+ )
375
+ criteria = {
376
+ placement.id: (
377
+ f"rotation {placement.rotation}, column {placement.x}; "
378
+ f"clears {len(placement.cleared_rows)}; holes {placement.holes}; "
379
+ f"height {placement.max_height}; top-out {'yes' if placement.top_out else 'no'}"
380
+ )
381
+ for placement in placements
382
+ }
383
+ payload: dict[str, Any] = {
384
+ "state": state,
385
+ "questions": {
386
+ "placement": {
387
+ "type": "choice",
388
+ "instructions": (
389
+ "Play Tetris to maximize long-term survival and score. Choose exactly "
390
+ "one listed legal placement for the current tetromino. Prefer line "
391
+ "clears, fewer holes, and a lower stack; avoid top-out. Return its "
392
+ "choice id."
393
+ ),
394
+ "criteria": criteria,
395
+ }
396
+ },
397
+ }
398
+ if model:
399
+ payload["model"] = model
400
+ return payload, state
401
+
402
+
403
+ def _valid_url(value: str, name: str) -> str:
404
+ try:
405
+ parsed = urlsplit(value.strip())
406
+ except ValueError as exc:
407
+ raise ValueError(f"{name} must be a valid HTTP or HTTPS URL") from exc
408
+ if (
409
+ parsed.scheme not in {"http", "https"}
410
+ or not parsed.netloc
411
+ or parsed.username
412
+ or parsed.password
413
+ or parsed.query
414
+ or parsed.fragment
415
+ ):
416
+ raise ValueError(
417
+ f"{name} must be an HTTP or HTTPS URL without credentials, query, or fragment"
418
+ )
419
+ return value.strip()
420
+
421
+
422
+ def _first_environment(env: Mapping[str, str], *names: str) -> str:
423
+ for name in names:
424
+ value = env.get(name, "").strip()
425
+ if value:
426
+ return value
427
+ return ""
428
+
429
+
430
+ def _environment_float(
431
+ env: Mapping[str, str],
432
+ name: str,
433
+ default: float,
434
+ *,
435
+ minimum: float,
436
+ maximum: float,
437
+ ) -> float:
438
+ raw = env.get(name, "").strip()
439
+ if not raw:
440
+ return float(default)
441
+ try:
442
+ value = float(raw)
443
+ except ValueError as exc:
444
+ raise ValueError(f"{name} must be numeric") from exc
445
+ if not math.isfinite(value) or not minimum <= value <= maximum:
446
+ raise ValueError(f"{name} must be between {minimum:g} and {maximum:g}")
447
+ return value
448
+
449
+
450
+ def _environment_int(
451
+ env: Mapping[str, str],
452
+ name: str,
453
+ default: int,
454
+ *,
455
+ minimum: int,
456
+ maximum: int,
457
+ ) -> int:
458
+ raw = env.get(name, "").strip()
459
+ if not raw:
460
+ return default
461
+ try:
462
+ value = int(raw)
463
+ except ValueError as exc:
464
+ raise ValueError(f"{name} must be an integer") from exc
465
+ if not minimum <= value <= maximum:
466
+ raise ValueError(f"{name} must be between {minimum} and {maximum}")
467
+ return value
468
+
469
+
470
+ class HTTPDecisionAdapter:
471
+ """Server-only adapters for local Decision, Jev Cloud, and SystemOne upstream."""
472
+
473
+ def __init__(
474
+ self,
475
+ *,
476
+ endpoints: Mapping[str, Endpoint],
477
+ timeout_seconds: float = 45.0,
478
+ client: Any = None,
479
+ clock=time.perf_counter,
480
+ ):
481
+ self._endpoints = dict(endpoints)
482
+ self._timeout_seconds = timeout_seconds
483
+ self._client = client
484
+ self._owns_client = client is None
485
+ self._clock = clock
486
+
487
+ @classmethod
488
+ def from_environment(
489
+ cls, env: Mapping[str, str] | None = None
490
+ ) -> HTTPDecisionAdapter:
491
+ env = os.environ if env is None else env
492
+ shared_local_key = _first_environment(env, "TETRIS_LOCAL_API_KEY")
493
+ endpoints: dict[str, Endpoint] = {}
494
+ for competitor in LOCAL_MODELS:
495
+ model_key = competitor.id.upper()
496
+ raw_url = _first_environment(
497
+ env,
498
+ f"TETRIS_LOCAL_{model_key}_API_URL",
499
+ "TETRIS_LOCAL_API_URL",
500
+ "LOCAL_API_URL",
501
+ )
502
+ if raw_url:
503
+ endpoints[competitor.id] = Endpoint(
504
+ _valid_url(raw_url, f"TETRIS_LOCAL_{model_key}_API_URL"),
505
+ _first_environment(env, f"TETRIS_LOCAL_{model_key}_API_KEY")
506
+ or shared_local_key,
507
+ competitor.request_model,
508
+ )
509
+
510
+ jev_url = _first_environment(env, "TETRIS_JEV_API_URL", "JEV_API_URL")
511
+ jev_key = _first_environment(env, "TETRIS_JEV_API_KEY", "JEV_API_KEY")
512
+ if jev_url and jev_key:
513
+ endpoints["jev-cloud"] = Endpoint(
514
+ _valid_url(jev_url, "TETRIS_JEV_API_URL"),
515
+ jev_key,
516
+ _first_environment(env, "TETRIS_JEV_MODEL") or "jev-latest",
517
+ "jev_alias_or_version",
518
+ )
519
+
520
+ systemone_url = _first_environment(
521
+ env, "TETRIS_SYSTEMONE_API_URL", "JEV_MIRROR_API_URL"
522
+ )
523
+ systemone_key = _first_environment(
524
+ env, "TETRIS_SYSTEMONE_API_KEY", "JEV_MIRROR_API_KEY"
525
+ )
526
+ if systemone_url and systemone_key:
527
+ endpoints["jev-cloud-mirror"] = Endpoint(
528
+ _valid_url(systemone_url, "TETRIS_SYSTEMONE_API_URL"),
529
+ systemone_key,
530
+ _first_environment(env, "TETRIS_SYSTEMONE_MODEL") or None,
531
+ "jev_alias_or_version",
532
+ )
533
+
534
+ timeout_raw = _first_environment(env, "TETRIS_REQUEST_TIMEOUT_SECONDS") or "45"
535
+ try:
536
+ timeout = float(timeout_raw)
537
+ except ValueError as exc:
538
+ raise ValueError("TETRIS_REQUEST_TIMEOUT_SECONDS must be numeric") from exc
539
+ if not math.isfinite(timeout) or not 0.1 <= timeout <= 300:
540
+ raise ValueError(
541
+ "TETRIS_REQUEST_TIMEOUT_SECONDS must be between 0.1 and 300"
542
+ )
543
+ return cls(endpoints=endpoints, timeout_seconds=timeout)
544
+
545
+ def catalog(self) -> tuple[Competitor, ...]:
546
+ cloud = (
547
+ Competitor(
548
+ "jev-cloud",
549
+ "Jev Cloud",
550
+ "cloud",
551
+ self._endpoints.get("jev-cloud", Endpoint("", "")).model,
552
+ "Jev Cloud",
553
+ "jev-cloud" in self._endpoints,
554
+ ),
555
+ Competitor(
556
+ "jev-cloud-mirror",
557
+ "Jev Cloud",
558
+ "cloud",
559
+ self._endpoints.get("jev-cloud-mirror", Endpoint("", "")).model,
560
+ "Jev Cloud",
561
+ "jev-cloud-mirror" in self._endpoints,
562
+ ),
563
+ )
564
+ return cloud + tuple(
565
+ Competitor(
566
+ model.id,
567
+ model.display_name,
568
+ model.family,
569
+ model.request_model,
570
+ model.option_label,
571
+ model.id in self._endpoints,
572
+ )
573
+ for model in LOCAL_MODELS
574
+ )
575
+
576
+ async def close(self) -> None:
577
+ if self._owns_client and self._client is not None:
578
+ await self._client.aclose()
579
+ self._client = None
580
+
581
+ async def decide(
582
+ self,
583
+ competitor: Competitor,
584
+ payload: Mapping[str, Any],
585
+ legal_choices: frozenset[str],
586
+ ) -> DecisionOutcome:
587
+ try:
588
+ import httpx
589
+ except ModuleNotFoundError as exc:
590
+ raise ArenaUpstreamError(
591
+ "The decision transport is not installed."
592
+ ) from exc
593
+ endpoint = self._endpoints.get(competitor.id)
594
+ if endpoint is None:
595
+ raise ArenaUpstreamError(
596
+ "The selected model is not configured on this server."
597
+ )
598
+ if self._client is None:
599
+ self._client = httpx.AsyncClient(
600
+ follow_redirects=False,
601
+ timeout=self._timeout_seconds,
602
+ trust_env=False,
603
+ )
604
+
605
+ upstream_payload = copy.deepcopy(dict(payload))
606
+ if endpoint.model:
607
+ upstream_payload["model"] = endpoint.model
608
+ else:
609
+ upstream_payload.pop("model", None)
610
+ headers = {
611
+ "Accept": "application/json",
612
+ "Accept-Encoding": "identity",
613
+ "Content-Type": "application/json",
614
+ }
615
+ if endpoint.api_key:
616
+ headers["Authorization"] = f"Bearer {endpoint.api_key}"
617
+
618
+ started = self._clock()
619
+ try:
620
+ async with self._client.stream(
621
+ "POST",
622
+ endpoint.url,
623
+ headers=headers,
624
+ json=upstream_payload,
625
+ ) as response:
626
+ if response.is_redirect:
627
+ raise ArenaUpstreamError(
628
+ "The selected model returned an unsupported redirect."
629
+ )
630
+ if response.status_code >= 400:
631
+ raise ArenaUpstreamError(
632
+ "The selected model could not complete this turn."
633
+ )
634
+ if response.headers.get("content-encoding", "identity").lower() != "identity":
635
+ raise ArenaUpstreamError(
636
+ "The selected model returned an unsupported encoding."
637
+ )
638
+ body = await self._read_response_body(response)
639
+ except httpx.HTTPError as exc:
640
+ # httpx exceptions are intentionally not reflected to the browser;
641
+ # messages can contain upstream topology.
642
+ raise ArenaUpstreamError(
643
+ "The selected model did not respond in time."
644
+ ) from exc
645
+ request_ms = max(0.0, (self._clock() - started) * 1_000)
646
+ try:
647
+ document = json.loads(body)
648
+ except (TypeError, ValueError, UnicodeError) as exc:
649
+ raise ArenaUpstreamError(
650
+ "The selected model returned invalid JSON."
651
+ ) from exc
652
+ if not isinstance(document, dict):
653
+ raise ArenaUpstreamError(
654
+ "The selected model returned an invalid response envelope."
655
+ )
656
+ answers = document.get("answers")
657
+ placement = answers.get("placement") if isinstance(answers, dict) else None
658
+ choice = placement.get("choice") if isinstance(placement, dict) else None
659
+ if not isinstance(choice, str) or choice not in legal_choices:
660
+ raise ArenaUpstreamError(
661
+ "The selected model did not return a legal placement."
662
+ )
663
+ provider_model = document.get("model")
664
+ self._validate_provider_model(endpoint, provider_model)
665
+ timing = document.get("timing")
666
+ inference_ms = timing.get("inference_ms") if isinstance(timing, dict) else None
667
+ if (
668
+ competitor.family != "decision"
669
+ or isinstance(inference_ms, bool)
670
+ or not isinstance(inference_ms, (int, float))
671
+ or not math.isfinite(inference_ms)
672
+ or inference_ms < 0
673
+ ):
674
+ inference_ms = None
675
+ return DecisionOutcome(
676
+ choice=choice,
677
+ request_ms=request_ms,
678
+ provider_model=provider_model,
679
+ inference_ms=float(inference_ms) if inference_ms is not None else None,
680
+ )
681
+
682
+ @staticmethod
683
+ async def _read_response_body(response: Any) -> bytes:
684
+ body = bytearray()
685
+ async for chunk in response.aiter_raw(chunk_size=64 * 1024):
686
+ body.extend(chunk)
687
+ if len(body) > MAX_UPSTREAM_BYTES:
688
+ raise ArenaUpstreamError(
689
+ "The selected model returned an oversized response."
690
+ )
691
+ return bytes(body)
692
+
693
+ @staticmethod
694
+ def _validate_provider_model(endpoint: Endpoint, provider_model: Any) -> None:
695
+ if not isinstance(provider_model, str) or not provider_model:
696
+ raise ArenaUpstreamError("The selected model returned no model identity.")
697
+ if endpoint.identity_policy == "exact":
698
+ if endpoint.model is None or provider_model != endpoint.model:
699
+ raise ArenaUpstreamError(
700
+ "The selected endpoint returned a different model identity."
701
+ )
702
+ return
703
+ if endpoint.identity_policy == "jev_alias_or_version":
704
+ requested = endpoint.model
705
+ if requested not in {None, "jev-latest"}:
706
+ if provider_model != requested:
707
+ raise ArenaUpstreamError(
708
+ "Jev returned a different pinned model identity."
709
+ )
710
+ return
711
+ if provider_model == "jev-latest" or _JEV_VERSION.fullmatch(provider_model):
712
+ return
713
+ raise ArenaUpstreamError("Jev returned an unsupported model identity.")
714
+ raise RuntimeError("unknown endpoint identity policy")
715
+
716
+
717
+ @dataclass
718
+ class SideResult:
719
+ side: str
720
+ competitor_id: str
721
+ display_name: str
722
+ status: str
723
+ reason: str
724
+ score: int
725
+ lines: int
726
+ pieces: int
727
+ duration_ms: float
728
+ total_request_ms: float
729
+
730
+ def public(self, *, include_duration: bool = True) -> dict[str, Any]:
731
+ result = {
732
+ "side": self.side,
733
+ "competitor_id": self.competitor_id,
734
+ "display_name": self.display_name,
735
+ "status": self.status,
736
+ "reason": self.reason,
737
+ "score": self.score,
738
+ "lines": self.lines,
739
+ "pieces": self.pieces,
740
+ }
741
+ if include_duration:
742
+ result["duration_ms"] = round(self.duration_ms, 3)
743
+ result["total_request_ms"] = round(self.total_request_ms, 3)
744
+ return result
745
+
746
+
747
+ def _advantage(
748
+ left: float,
749
+ right: float,
750
+ *,
751
+ lower_wins: bool,
752
+ ) -> dict[str, Any]:
753
+ if math.isclose(left, right, rel_tol=1e-9, abs_tol=1e-9):
754
+ return {"winner": "tie", "percent": 0.0, "unbounded": False}
755
+ if lower_wins:
756
+ winner = "left" if left < right else "right"
757
+ better, worse = sorted((left, right))
758
+ percent = None if better <= 0 else ((worse / better) - 1) * 100
759
+ else:
760
+ winner = "left" if left > right else "right"
761
+ worse, better = sorted((left, right))
762
+ percent = None if worse <= 0 else ((better / worse) - 1) * 100
763
+ return {
764
+ "winner": winner,
765
+ "percent": None if percent is None else round(percent, 1),
766
+ "unbounded": percent is None,
767
+ }
768
+
769
+
770
+ class RaceSession:
771
+ def __init__(
772
+ self,
773
+ *,
774
+ race_id: str,
775
+ adapter: Any,
776
+ left: Competitor,
777
+ right: Competitor,
778
+ client_id: str,
779
+ seed: str | int,
780
+ mode: str,
781
+ max_steps: int | None,
782
+ first_failure_guard: int,
783
+ race_timeout_seconds: float = DEFAULT_RACE_TIMEOUT_SECONDS,
784
+ stream_grace_seconds: float = DEFAULT_STREAM_GRACE_SECONDS,
785
+ clock=time.perf_counter,
786
+ ):
787
+ self.id = race_id
788
+ self.adapter = adapter
789
+ self.competitors = {"left": left, "right": right}
790
+ self.client_id = client_id
791
+ self.seed = seed
792
+ self.mode = mode
793
+ self.max_steps = max_steps
794
+ self.call_budget = 2 * (max_steps or first_failure_guard)
795
+ self.history_limit = min(MAX_RACE_HISTORY, self.call_budget + 8)
796
+ self.provider_calls = 0
797
+ self.first_failure_guard = first_failure_guard
798
+ self.race_timeout_seconds = race_timeout_seconds
799
+ self.stream_grace_seconds = stream_grace_seconds
800
+ sequence_length = (max_steps or first_failure_guard) + 1
801
+ self.sequence = create_piece_sequence(seed, sequence_length)
802
+ self.clock = clock
803
+ self.created_at = self.clock()
804
+ self.completed_at: float | None = None
805
+ self.status = "created"
806
+ self.events: list[dict[str, Any]] = []
807
+ self.states: dict[str, dict[str, Any]] = {
808
+ side: {"side": side, "step": 0, "request_ms": None, "game": Game().public()}
809
+ for side in ("left", "right")
810
+ }
811
+ self.traces: dict[str, list[dict[str, Any]]] = {"left": [], "right": []}
812
+ self.results: dict[str, SideResult] = {}
813
+ self.summary: dict[str, Any] | None = None
814
+ self._condition = asyncio.Condition()
815
+ self._done = asyncio.Event()
816
+ self._task: asyncio.Task[None] | None = None
817
+ self._orphan_task: asyncio.Task[None] | None = None
818
+ self._subscribers = 0
819
+ self._stop_reason = "cancelled"
820
+
821
+ def start(self) -> None:
822
+ if self._task is not None:
823
+ raise RuntimeError("race already started")
824
+ self._task = asyncio.create_task(self._run(), name=f"tetris-race-{self.id}")
825
+ self._schedule_orphan_watch()
826
+
827
+ async def cancel(self, reason: str = "cancelled") -> None:
828
+ self._stop_reason = reason
829
+ current = asyncio.current_task()
830
+ orphan_to_wait: asyncio.Task[None] | None = None
831
+ if self._orphan_task is not None and self._orphan_task is not current:
832
+ orphan_to_wait = self._orphan_task
833
+ self._orphan_task = None
834
+ orphan_to_wait.cancel()
835
+ if self._task is not None and not self._task.done():
836
+ self._task.cancel()
837
+ try:
838
+ await self._task
839
+ except asyncio.CancelledError:
840
+ pass
841
+ if orphan_to_wait is not None:
842
+ await asyncio.gather(orphan_to_wait, return_exceptions=True)
843
+ # A task cancelled before its coroutine first runs never reaches _run's
844
+ # finally block. Finalize it here so waiters and capacity accounting do
845
+ # not retain a permanently "created" race.
846
+ if not self._done.is_set():
847
+ self.status = "cancelled"
848
+ self.completed_at = self.clock()
849
+ self._done.set()
850
+ async with self._condition:
851
+ self._condition.notify_all()
852
+
853
+ async def wait(self) -> dict[str, Any]:
854
+ await self._done.wait()
855
+ return self.snapshot()
856
+
857
+ async def _publish(self, event_type: str, data: Mapping[str, Any]) -> None:
858
+ event = {
859
+ "id": len(self.events) + 1,
860
+ "type": event_type,
861
+ "data": copy.deepcopy(dict(data)),
862
+ }
863
+ async with self._condition:
864
+ if len(self.events) >= self.history_limit:
865
+ raise RuntimeError("race event history limit exceeded")
866
+ self.events.append(event)
867
+ self._condition.notify_all()
868
+
869
+ def _schedule_orphan_watch(self) -> None:
870
+ if self._done.is_set() or self._subscribers or self.stream_grace_seconds <= 0:
871
+ return
872
+ if self._orphan_task is not None and not self._orphan_task.done():
873
+ return
874
+ self._orphan_task = asyncio.create_task(
875
+ self._cancel_if_orphaned(),
876
+ name=f"tetris-orphan-{self.id}",
877
+ )
878
+
879
+ async def _cancel_if_orphaned(self) -> None:
880
+ try:
881
+ await asyncio.sleep(self.stream_grace_seconds)
882
+ async with self._condition:
883
+ should_cancel = (
884
+ self._subscribers == 0
885
+ and not self._done.is_set()
886
+ and self.status in {"created", "running"}
887
+ )
888
+ if should_cancel:
889
+ await self.cancel("stream_disconnected")
890
+ except asyncio.CancelledError:
891
+ return
892
+ finally:
893
+ if self._orphan_task is asyncio.current_task():
894
+ self._orphan_task = None
895
+
896
+ async def iter_events(self, after: int = 0) -> AsyncIterator[dict[str, Any] | None]:
897
+ cursor = max(0, after)
898
+ async with self._condition:
899
+ self._subscribers += 1
900
+ orphan_task = self._orphan_task
901
+ self._orphan_task = None
902
+ if orphan_task is not None:
903
+ orphan_task.cancel()
904
+ await asyncio.gather(orphan_task, return_exceptions=True)
905
+ try:
906
+ while True:
907
+ heartbeat = False
908
+ async with self._condition:
909
+ if cursor < len(self.events):
910
+ pending = copy.deepcopy(self.events[cursor:])
911
+ elif self._done.is_set():
912
+ return
913
+ else:
914
+ try:
915
+ await asyncio.wait_for(self._condition.wait(), timeout=10)
916
+ except asyncio.TimeoutError:
917
+ heartbeat = True
918
+ pending = []
919
+ if heartbeat:
920
+ # Yield outside the condition lock: a stalled renderer must never
921
+ # prevent either model loop from publishing or scheduling work.
922
+ yield None
923
+ if not pending:
924
+ continue
925
+ for event in pending:
926
+ cursor = event["id"]
927
+ yield event
928
+ finally:
929
+ async with self._condition:
930
+ self._subscribers = max(0, self._subscribers - 1)
931
+ self._schedule_orphan_watch()
932
+
933
+ def snapshot(self, *, include_traces: bool = False) -> dict[str, Any]:
934
+ result: dict[str, Any] = {
935
+ "id": self.id,
936
+ "status": self.status,
937
+ "seed": self.seed,
938
+ "mode": self.mode,
939
+ "max_steps": self.max_steps,
940
+ "competitors": {
941
+ side: competitor.public()
942
+ for side, competitor in self.competitors.items()
943
+ },
944
+ "results": {
945
+ side: side_result.public() for side, side_result in self.results.items()
946
+ },
947
+ "states": copy.deepcopy(self.states),
948
+ "summary": copy.deepcopy(self.summary),
949
+ }
950
+ if include_traces:
951
+ result["traces"] = copy.deepcopy(self.traces)
952
+ return result
953
+
954
+ async def _run(self) -> None:
955
+ tasks: dict[str, asyncio.Task[SideResult]] = {}
956
+ try:
957
+ self.status = "running"
958
+ await self._publish(
959
+ "race.started",
960
+ {
961
+ "race_id": self.id,
962
+ "seed": self.seed,
963
+ "mode": self.mode,
964
+ "max_steps": self.max_steps,
965
+ "competitors": {
966
+ side: competitor.public()
967
+ for side, competitor in self.competitors.items()
968
+ },
969
+ },
970
+ )
971
+ tasks = {
972
+ side: asyncio.create_task(
973
+ self._run_side(side, competitor),
974
+ name=f"tetris-{self.id}-{side}",
975
+ )
976
+ for side, competitor in self.competitors.items()
977
+ }
978
+ # asyncio.timeout cancels the awaited child tasks before control
979
+ # reaches the TimeoutError handler, so publish the intended stop
980
+ # reason before entering the deadline scope.
981
+ self._stop_reason = "race_timeout"
982
+ try:
983
+ async with asyncio.timeout(self.race_timeout_seconds):
984
+ await self._finish_competitors(tasks)
985
+ except TimeoutError:
986
+ self._stop_reason = "race_timeout"
987
+ await self._stop_tasks(tasks)
988
+ self.status = "timed_out"
989
+ else:
990
+ self.status = (
991
+ "incomplete"
992
+ if any(
993
+ result.status in {"error", "guard_reached"}
994
+ or (
995
+ result.status == "stopped"
996
+ and result.reason != "opponent_finished"
997
+ )
998
+ for result in self.results.values()
999
+ )
1000
+ else "finished"
1001
+ )
1002
+ self.summary = self._build_summary()
1003
+ await self._publish("race.finished", self.summary)
1004
+ except asyncio.CancelledError:
1005
+ await self._stop_tasks(tasks)
1006
+ self.status = "cancelled"
1007
+ raise
1008
+ except Exception: # noqa: BLE001 - fail closed and terminate both runners
1009
+ self._stop_reason = "internal_error"
1010
+ await self._stop_tasks(tasks)
1011
+ self.status = "error"
1012
+ if len(self.results) == 2:
1013
+ self.summary = self._build_summary()
1014
+ await self._publish("race.finished", self.summary)
1015
+ finally:
1016
+ # A fast race can finish before the browser opens its event stream.
1017
+ # Reclaim the initial reconnect-grace sleeper as soon as the race is
1018
+ # terminal. If that sleeper initiated cancellation, let it unwind
1019
+ # naturally instead of making the two tasks cancel each other.
1020
+ if self.status != "cancelled":
1021
+ orphan_task = self._orphan_task
1022
+ self._orphan_task = None
1023
+ if orphan_task is not None:
1024
+ orphan_task.cancel()
1025
+ await asyncio.gather(orphan_task, return_exceptions=True)
1026
+ self.completed_at = self.clock()
1027
+ self._done.set()
1028
+ async with self._condition:
1029
+ self._condition.notify_all()
1030
+
1031
+ async def _finish_competitors(
1032
+ self, tasks: Mapping[str, asyncio.Task[SideResult]]
1033
+ ) -> None:
1034
+ if self.mode == "first_failure":
1035
+ pending = set(tasks.values())
1036
+ while pending:
1037
+ done, pending = await asyncio.wait(
1038
+ pending, return_when=asyncio.FIRST_COMPLETED
1039
+ )
1040
+ results = [task.result() for task in done]
1041
+ self._record_results(results)
1042
+ if any(result.status in {"game_over", "error"} for result in results):
1043
+ self._stop_reason = "opponent_finished"
1044
+ await self._stop_tasks(tasks)
1045
+ break
1046
+ return
1047
+ self._record_results(await asyncio.gather(*tasks.values()))
1048
+
1049
+ def _record_results(self, values: Any) -> None:
1050
+ for result in values:
1051
+ if isinstance(result, SideResult):
1052
+ self.results[result.side] = result
1053
+
1054
+ async def _stop_tasks(self, tasks: Mapping[str, asyncio.Task[SideResult]]) -> None:
1055
+ for task in tasks.values():
1056
+ if not task.done():
1057
+ task.cancel()
1058
+ completed = await asyncio.gather(*tasks.values(), return_exceptions=True)
1059
+ self._record_results(completed)
1060
+ for side, competitor in self.competitors.items():
1061
+ if side not in self.results and tasks:
1062
+ self.results[side] = SideResult(
1063
+ side=side,
1064
+ competitor_id=competitor.id,
1065
+ display_name=competitor.display_name,
1066
+ status="stopped",
1067
+ reason=self._stop_reason,
1068
+ score=0,
1069
+ lines=0,
1070
+ pieces=0,
1071
+ duration_ms=0,
1072
+ total_request_ms=0,
1073
+ )
1074
+
1075
+ async def _run_side(self, side: str, competitor: Competitor) -> SideResult:
1076
+ game = Game()
1077
+ started = self.clock()
1078
+ status = "finished"
1079
+ reason = "step_limit"
1080
+ total_request_ms = 0.0
1081
+ limit = self.max_steps if self.mode == "steps" else self.first_failure_guard
1082
+ try:
1083
+ await self._publish(
1084
+ "side.started",
1085
+ {
1086
+ "side": side,
1087
+ "competitor": competitor.public(),
1088
+ "game": game.public(),
1089
+ },
1090
+ )
1091
+ for step in range(limit or 0):
1092
+ piece = self.sequence[step]
1093
+ next_piece = self.sequence[step + 1]
1094
+ placements = enumerate_placements(game, piece)
1095
+ if not placements:
1096
+ game.game_over = True
1097
+ status = "game_over"
1098
+ reason = "no_legal_placement"
1099
+ break
1100
+ request, state = build_decision_request(
1101
+ game,
1102
+ piece,
1103
+ next_piece,
1104
+ placements,
1105
+ competitor.request_model,
1106
+ )
1107
+ before = game.public()
1108
+ legal_choices = frozenset(placement.id for placement in placements)
1109
+ try:
1110
+ self.provider_calls += 1
1111
+ outcome = await self.adapter.decide(
1112
+ competitor,
1113
+ request,
1114
+ legal_choices,
1115
+ )
1116
+ except ArenaUpstreamError as exc:
1117
+ status = "error"
1118
+ reason = "provider_error"
1119
+ trace = {
1120
+ "step": step + 1,
1121
+ "piece": piece,
1122
+ "state": state,
1123
+ "board_before": before["board"],
1124
+ "action": None,
1125
+ "error": str(exc),
1126
+ }
1127
+ self.traces[side].append(trace)
1128
+ await self._publish(
1129
+ "side.error",
1130
+ {"side": side, "step": step + 1, "message": str(exc)},
1131
+ )
1132
+ break
1133
+ if not isinstance(outcome, DecisionOutcome):
1134
+ raise TypeError("decision adapter returned an invalid outcome")
1135
+ if not math.isfinite(outcome.request_ms) or outcome.request_ms < 0:
1136
+ raise TypeError("decision adapter returned an invalid request duration")
1137
+ total_request_ms += outcome.request_ms
1138
+ chosen = next(
1139
+ (
1140
+ placement
1141
+ for placement in placements
1142
+ if placement.id == outcome.choice
1143
+ ),
1144
+ None,
1145
+ )
1146
+ if chosen is None:
1147
+ # Never replace a model's invalid choice with a heuristic choice.
1148
+ status = "error"
1149
+ reason = "invalid_choice"
1150
+ self.traces[side].append(
1151
+ {
1152
+ "step": step + 1,
1153
+ "piece": piece,
1154
+ "state": state,
1155
+ "board_before": before["board"],
1156
+ "action": outcome.choice,
1157
+ "error": "invalid_choice",
1158
+ }
1159
+ )
1160
+ await self._publish(
1161
+ "side.error",
1162
+ {
1163
+ "side": side,
1164
+ "step": step + 1,
1165
+ "message": "The model returned an invalid placement.",
1166
+ },
1167
+ )
1168
+ break
1169
+ game = apply_placement(game, chosen)
1170
+ after = game.public()
1171
+ trace = {
1172
+ "step": step + 1,
1173
+ "piece": piece,
1174
+ "next_piece": next_piece,
1175
+ "state": state,
1176
+ "board_before": before["board"],
1177
+ "action": chosen.id,
1178
+ "board_after": after["board"],
1179
+ "score": game.score,
1180
+ "lines": game.lines,
1181
+ "pieces": game.pieces,
1182
+ "request_ms": round(outcome.request_ms, 3),
1183
+ "provider_model": outcome.provider_model,
1184
+ }
1185
+ if outcome.inference_ms is not None:
1186
+ trace["provider_inference_ms"] = round(outcome.inference_ms, 3)
1187
+ self.traces[side].append(trace)
1188
+ step_event = {
1189
+ "side": side,
1190
+ "step": step + 1,
1191
+ "piece": piece,
1192
+ "next_piece": next_piece,
1193
+ "choice": chosen.id,
1194
+ "request_ms": round(outcome.request_ms, 3),
1195
+ "cleared_rows": list(chosen.cleared_rows),
1196
+ "game": after,
1197
+ }
1198
+ self.states[side] = step_event
1199
+ await self._publish("side.step", step_event)
1200
+ if game.game_over:
1201
+ status = "game_over"
1202
+ reason = "top_out"
1203
+ break
1204
+ else:
1205
+ if self.mode == "first_failure":
1206
+ status = "guard_reached"
1207
+ reason = "operational_guard"
1208
+ except asyncio.CancelledError:
1209
+ status = "stopped"
1210
+ reason = self._stop_reason
1211
+ except Exception: # noqa: BLE001 - isolate one runner without exposing internals
1212
+ status = "error"
1213
+ reason = "internal_error"
1214
+ await self._publish(
1215
+ "side.error",
1216
+ {
1217
+ "side": side,
1218
+ "message": "The server could not complete this model's turn.",
1219
+ },
1220
+ )
1221
+
1222
+ result = SideResult(
1223
+ side=side,
1224
+ competitor_id=competitor.id,
1225
+ display_name=competitor.display_name,
1226
+ status=status,
1227
+ reason=reason,
1228
+ score=game.score,
1229
+ lines=game.lines,
1230
+ pieces=game.pieces,
1231
+ duration_ms=max(0.0, (self.clock() - started) * 1_000),
1232
+ total_request_ms=total_request_ms,
1233
+ )
1234
+ await self._publish(
1235
+ "side.finished",
1236
+ {
1237
+ "side": side,
1238
+ # The in-board card intentionally contains game results only.
1239
+ "result": result.public(include_duration=False),
1240
+ },
1241
+ )
1242
+ return result
1243
+
1244
+ def _build_summary(self) -> dict[str, Any]:
1245
+ left = self.results["left"]
1246
+ right = self.results["right"]
1247
+ comparable_speed = (
1248
+ self.status == "finished"
1249
+ and self.mode == "steps"
1250
+ and left.reason == "step_limit"
1251
+ and right.reason == "step_limit"
1252
+ and left.pieces == right.pieces == self.max_steps
1253
+ )
1254
+ speed = (
1255
+ _advantage(left.total_request_ms, right.total_request_ms, lower_wins=True)
1256
+ if comparable_speed
1257
+ else {"winner": "unavailable", "percent": None, "unbounded": False}
1258
+ )
1259
+ comparable_score = (
1260
+ self.status == "finished"
1261
+ and left.status in {"finished", "game_over"}
1262
+ and right.status in {"finished", "game_over"}
1263
+ )
1264
+ return {
1265
+ "race_id": self.id,
1266
+ "status": self.status,
1267
+ "mode": self.mode,
1268
+ "left": left.public(),
1269
+ "right": right.public(),
1270
+ "speed": {
1271
+ **speed,
1272
+ "basis": "server_observed_request_to_response_e2e_total",
1273
+ "comparable": comparable_speed,
1274
+ },
1275
+ "score": {
1276
+ **(
1277
+ _advantage(float(left.score), float(right.score), lower_wins=False)
1278
+ if comparable_score
1279
+ else {"winner": "unavailable", "percent": None, "unbounded": False}
1280
+ ),
1281
+ "basis": "final_game_score",
1282
+ "comparable": comparable_score,
1283
+ },
1284
+ }
1285
+
1286
+
1287
+ @dataclass
1288
+ class Admission:
1289
+ """Keep quota accounting without retaining a completed race's traces."""
1290
+
1291
+ session: RaceSession | None
1292
+ created_at: float
1293
+ call_budget: int
1294
+ completed_at: float | None = None
1295
+ provider_calls: int = 0
1296
+
1297
+ def refresh(self) -> None:
1298
+ if self.session is not None and self.session.completed_at is not None:
1299
+ self.completed_at = self.session.completed_at
1300
+ self.provider_calls = self.session.provider_calls
1301
+ self.session = None
1302
+
1303
+ def active(self) -> bool:
1304
+ return self.session is not None and self.session.status in {"created", "running"}
1305
+
1306
+ def charged_calls(self) -> int:
1307
+ if self.session is not None:
1308
+ if self.session.completed_at is None:
1309
+ return self.call_budget
1310
+ return self.session.provider_calls
1311
+ return self.provider_calls
1312
+
1313
+
1314
+ class RaceManager:
1315
+ def __init__(
1316
+ self,
1317
+ adapter: Any,
1318
+ *,
1319
+ max_concurrent_races: int = MAX_CONCURRENT_RACES,
1320
+ client_active_races: int = DEFAULT_CLIENT_ACTIVE_RACES,
1321
+ client_races_per_minute: int = DEFAULT_CLIENT_RACES_PER_MINUTE,
1322
+ client_calls_per_window: int | None = None,
1323
+ client_call_window_seconds: float = DEFAULT_CLIENT_CALL_WINDOW_SECONDS,
1324
+ max_steps_limit: int = DEFAULT_MAX_CONFIGURED_STEPS,
1325
+ first_failure_guard: int | None = None,
1326
+ race_timeout_seconds: float = DEFAULT_RACE_TIMEOUT_SECONDS,
1327
+ stream_grace_seconds: float = DEFAULT_STREAM_GRACE_SECONDS,
1328
+ completed_ttl_seconds: float = DEFAULT_COMPLETED_TTL_SECONDS,
1329
+ clock=time.perf_counter,
1330
+ ):
1331
+ if not 1 <= max_steps_limit <= MAX_CONFIGURED_STEPS:
1332
+ raise ValueError("max_steps_limit is out of range")
1333
+ if first_failure_guard is None:
1334
+ first_failure_guard = max_steps_limit
1335
+ if client_calls_per_window is None:
1336
+ client_calls_per_window = max(
1337
+ DEFAULT_CLIENT_CALLS_PER_WINDOW, 2 * max_steps_limit
1338
+ )
1339
+ if max_concurrent_races < 1:
1340
+ raise ValueError("max_concurrent_races must be positive")
1341
+ if client_active_races < 1 or client_active_races > max_concurrent_races:
1342
+ raise ValueError("client_active_races must fit global capacity")
1343
+ if not 1 <= client_races_per_minute <= MAX_CLIENT_RACES_PER_MINUTE:
1344
+ raise ValueError("client_races_per_minute is out of range")
1345
+ if client_calls_per_window < 2 * max_steps_limit:
1346
+ raise ValueError("client_calls_per_window must permit one full race")
1347
+ if client_call_window_seconds <= 0:
1348
+ raise ValueError("client_call_window_seconds must be positive")
1349
+ if not 1 <= first_failure_guard <= max_steps_limit:
1350
+ raise ValueError("first_failure_guard is out of range")
1351
+ if race_timeout_seconds <= 0:
1352
+ raise ValueError("race_timeout_seconds must be positive")
1353
+ if stream_grace_seconds <= 0:
1354
+ raise ValueError("stream_grace_seconds must be positive")
1355
+ if completed_ttl_seconds <= 0:
1356
+ raise ValueError("completed_ttl_seconds must be positive")
1357
+ self.adapter = adapter
1358
+ self.max_concurrent_races = max_concurrent_races
1359
+ self.client_active_races = client_active_races
1360
+ self.client_races_per_minute = client_races_per_minute
1361
+ self.client_calls_per_window = client_calls_per_window
1362
+ self.client_call_window_seconds = client_call_window_seconds
1363
+ self.max_steps_limit = max_steps_limit
1364
+ self.first_failure_guard = first_failure_guard
1365
+ self.race_timeout_seconds = race_timeout_seconds
1366
+ self.stream_grace_seconds = stream_grace_seconds
1367
+ self.completed_ttl_seconds = completed_ttl_seconds
1368
+ self.clock = clock
1369
+ self.sessions: dict[str, RaceSession] = {}
1370
+ self._client_history: dict[str, deque[Admission]] = {}
1371
+ self._lock = asyncio.Lock()
1372
+ self._closing = False
1373
+ self._closed = asyncio.Event()
1374
+ self._reaper_task: asyncio.Task[None] | None = None
1375
+
1376
+ @classmethod
1377
+ def from_environment(
1378
+ cls,
1379
+ adapter: Any,
1380
+ env: Mapping[str, str] | None = None,
1381
+ ) -> RaceManager:
1382
+ env = os.environ if env is None else env
1383
+ max_steps_limit = _environment_int(
1384
+ env, "TETRIS_MAX_STEPS", DEFAULT_MAX_CONFIGURED_STEPS,
1385
+ minimum=1, maximum=MAX_CONFIGURED_STEPS,
1386
+ )
1387
+ return cls(
1388
+ adapter,
1389
+ max_steps_limit=max_steps_limit,
1390
+ first_failure_guard=_environment_int(
1391
+ env, "TETRIS_FIRST_FAILURE_GUARD", max_steps_limit,
1392
+ minimum=1, maximum=max_steps_limit,
1393
+ ),
1394
+ client_active_races=_environment_int(
1395
+ env, "TETRIS_CLIENT_ACTIVE_RACES", DEFAULT_CLIENT_ACTIVE_RACES,
1396
+ minimum=1, maximum=MAX_CONCURRENT_RACES,
1397
+ ),
1398
+ client_races_per_minute=_environment_int(
1399
+ env, "TETRIS_CLIENT_RACES_PER_MINUTE", DEFAULT_CLIENT_RACES_PER_MINUTE,
1400
+ minimum=1, maximum=MAX_CLIENT_RACES_PER_MINUTE,
1401
+ ),
1402
+ client_calls_per_window=_environment_int(
1403
+ env, "TETRIS_CLIENT_CALLS_PER_WINDOW",
1404
+ max(DEFAULT_CLIENT_CALLS_PER_WINDOW, 2 * max_steps_limit),
1405
+ minimum=2 * max_steps_limit, maximum=10_000,
1406
+ ),
1407
+ race_timeout_seconds=_environment_float(
1408
+ env,
1409
+ "TETRIS_RACE_TIMEOUT_SECONDS",
1410
+ DEFAULT_RACE_TIMEOUT_SECONDS,
1411
+ minimum=1,
1412
+ maximum=3_600,
1413
+ ),
1414
+ stream_grace_seconds=_environment_float(
1415
+ env,
1416
+ "TETRIS_STREAM_GRACE_SECONDS",
1417
+ DEFAULT_STREAM_GRACE_SECONDS,
1418
+ minimum=1,
1419
+ maximum=300,
1420
+ ),
1421
+ completed_ttl_seconds=_environment_float(
1422
+ env,
1423
+ "TETRIS_COMPLETED_TTL_SECONDS",
1424
+ DEFAULT_COMPLETED_TTL_SECONDS,
1425
+ minimum=30,
1426
+ maximum=3_600,
1427
+ ),
1428
+ )
1429
+
1430
+ def catalog(self) -> tuple[Competitor, ...]:
1431
+ return tuple(self.adapter.catalog())
1432
+
1433
+ def public_config(self) -> dict[str, Any]:
1434
+ return {
1435
+ "competitors": [competitor.public() for competitor in self.catalog()],
1436
+ "defaults": {
1437
+ "max_steps": min(DEFAULT_MAX_STEPS, self.max_steps_limit),
1438
+ "seed": 42,
1439
+ },
1440
+ "limits": {
1441
+ "max_steps": self.max_steps_limit,
1442
+ "first_failure_guard": self.first_failure_guard,
1443
+ },
1444
+ "modes": ["steps", "first_failure"],
1445
+ }
1446
+
1447
+ async def create(self, value: Any, *, client_id: str = "unknown") -> RaceSession:
1448
+ request = self._validate(value)
1449
+ if not isinstance(client_id, str) or not 1 <= len(client_id) <= 128:
1450
+ raise ArenaValidationError("Invalid client identity.")
1451
+ async with self._lock:
1452
+ if self._closing:
1453
+ raise ArenaValidationError("The arena is shutting down.")
1454
+ self._prune_locked()
1455
+ self._ensure_reaper_locked()
1456
+ self._prune_admissions_locked()
1457
+ active = sum(
1458
+ session.status in {"created", "running"}
1459
+ for session in self.sessions.values()
1460
+ )
1461
+ if active >= self.max_concurrent_races:
1462
+ raise ArenaQuotaError("The arena is busy. Try again shortly.", retry_after=2)
1463
+ history = self._client_history.get(client_id, deque())
1464
+ active_for_client = sum(admission.active() for admission in history)
1465
+ if active_for_client >= self.client_active_races:
1466
+ raise ArenaQuotaError("A race is already active for this client.", retry_after=2)
1467
+ now = self.clock()
1468
+ recent = [
1469
+ admission for admission in history
1470
+ if now - admission.created_at < DEFAULT_CLIENT_WINDOW_SECONDS
1471
+ ]
1472
+ if len(recent) >= self.client_races_per_minute:
1473
+ retry = max(
1474
+ 1,
1475
+ math.ceil(
1476
+ DEFAULT_CLIENT_WINDOW_SECONDS - (now - recent[0].created_at)
1477
+ ),
1478
+ )
1479
+ raise ArenaQuotaError("Race start rate limit reached.", retry_after=retry)
1480
+ cost = 2 * (request["max_steps"] or self.first_failure_guard)
1481
+ used = sum(admission.charged_calls() for admission in history)
1482
+ if used + cost > self.client_calls_per_window:
1483
+ raise ArenaQuotaError("Provider call budget reached.", retry_after=60)
1484
+ if sum(map(len, self._client_history.values())) >= MAX_ADMISSION_HISTORY:
1485
+ raise ArenaQuotaError("The arena is busy. Try again shortly.", retry_after=60)
1486
+ race_id = secrets.token_urlsafe(18)
1487
+ session = RaceSession(
1488
+ race_id=race_id,
1489
+ adapter=self.adapter,
1490
+ left=request["left"],
1491
+ right=request["right"],
1492
+ client_id=client_id,
1493
+ seed=request["seed"],
1494
+ mode=request["mode"],
1495
+ max_steps=request["max_steps"],
1496
+ first_failure_guard=self.first_failure_guard,
1497
+ race_timeout_seconds=self.race_timeout_seconds,
1498
+ stream_grace_seconds=self.stream_grace_seconds,
1499
+ clock=self.clock,
1500
+ )
1501
+ self.sessions[race_id] = session
1502
+ self._client_history.setdefault(client_id, deque()).append(
1503
+ Admission(session, session.created_at, session.call_budget)
1504
+ )
1505
+ session.start()
1506
+ return session
1507
+
1508
+ def get(self, race_id: str) -> RaceSession:
1509
+ self._prune_locked()
1510
+ session = self.sessions.get(race_id)
1511
+ if session is None:
1512
+ raise KeyError(race_id)
1513
+ return session
1514
+
1515
+ async def aclose(self) -> None:
1516
+ async with self._lock:
1517
+ if self._closing:
1518
+ wait_for_close = True
1519
+ reaper = None
1520
+ sessions = ()
1521
+ else:
1522
+ wait_for_close = False
1523
+ self._closing = True
1524
+ reaper = self._reaper_task
1525
+ self._reaper_task = None
1526
+ sessions = tuple(self.sessions.values())
1527
+ if wait_for_close:
1528
+ await self._closed.wait()
1529
+ return
1530
+ try:
1531
+ if reaper is not None:
1532
+ reaper.cancel()
1533
+ await asyncio.gather(reaper, return_exceptions=True)
1534
+ await asyncio.gather(
1535
+ *(session.cancel("shutdown") for session in sessions),
1536
+ return_exceptions=True,
1537
+ )
1538
+ async with self._lock:
1539
+ self.sessions.clear()
1540
+ self._client_history.clear()
1541
+ finally:
1542
+ self._closed.set()
1543
+
1544
+ def _ensure_reaper_locked(self) -> None:
1545
+ if self._reaper_task is None or self._reaper_task.done():
1546
+ self._reaper_task = asyncio.create_task(
1547
+ self._reap_loop(),
1548
+ name="tetris-race-reaper",
1549
+ )
1550
+
1551
+ async def _reap_loop(self) -> None:
1552
+ interval = min(30.0, max(0.1, self.completed_ttl_seconds / 2))
1553
+ try:
1554
+ while True:
1555
+ await asyncio.sleep(interval)
1556
+ async with self._lock:
1557
+ self._prune_locked()
1558
+ self._prune_admissions_locked()
1559
+ except asyncio.CancelledError:
1560
+ return
1561
+
1562
+ def _prune_locked(self) -> None:
1563
+ now = self.clock()
1564
+ for race_id, session in tuple(self.sessions.items()):
1565
+ if (
1566
+ session.completed_at is not None
1567
+ and now - session.completed_at >= self.completed_ttl_seconds
1568
+ ):
1569
+ self.sessions.pop(race_id, None)
1570
+ if len(self.sessions) < MAX_RETAINED_RACES:
1571
+ return
1572
+ completed = sorted(
1573
+ (
1574
+ session
1575
+ for session in self.sessions.values()
1576
+ if session.completed_at is not None
1577
+ ),
1578
+ key=lambda session: session.completed_at or session.created_at,
1579
+ )
1580
+ for session in completed[: len(self.sessions) - MAX_RETAINED_RACES + 1]:
1581
+ self.sessions.pop(session.id, None)
1582
+
1583
+ def _prune_admissions_locked(self) -> None:
1584
+ now = self.clock()
1585
+ for client_id, history in tuple(self._client_history.items()):
1586
+ for admission in history:
1587
+ admission.refresh()
1588
+ retained = deque(
1589
+ admission for admission in history
1590
+ if admission.completed_at is None or now - max(
1591
+ admission.created_at, admission.completed_at
1592
+ ) < max(DEFAULT_CLIENT_WINDOW_SECONDS, self.client_call_window_seconds)
1593
+ )
1594
+ if retained:
1595
+ self._client_history[client_id] = retained
1596
+ else:
1597
+ self._client_history.pop(client_id, None)
1598
+
1599
+ async def cancel(self, race_id: str, *, reason: str = "user_cancelled") -> None:
1600
+ async with self._lock:
1601
+ session = self.sessions.pop(race_id, None)
1602
+ if session is None:
1603
+ raise KeyError(race_id)
1604
+ await session.cancel(reason)
1605
+
1606
+ def _validate(self, value: Any) -> dict[str, Any]:
1607
+ if not isinstance(value, dict):
1608
+ raise ArenaValidationError("Race request must be a JSON object.")
1609
+ allowed = {"left", "right", "seed", "mode", "max_steps"}
1610
+ if set(value) - allowed:
1611
+ raise ArenaValidationError("Race request contains unsupported fields.")
1612
+ by_id = {competitor.id: competitor for competitor in self.catalog()}
1613
+ selected: dict[str, Competitor] = {}
1614
+ for side in ("left", "right"):
1615
+ competitor_id = value.get(side)
1616
+ competitor = by_id.get(competitor_id)
1617
+ if competitor is None or not competitor.ready:
1618
+ raise ArenaValidationError(
1619
+ f"{side} must select a configured competitor."
1620
+ )
1621
+ selected[side] = competitor
1622
+ seed = value.get("seed", 42)
1623
+ if isinstance(seed, bool) or not isinstance(seed, (str, int)):
1624
+ raise ArenaValidationError("seed must be a string or integer.")
1625
+ if isinstance(seed, str) and (not seed or len(seed) > 128):
1626
+ raise ArenaValidationError("seed must contain 1 to 128 characters.")
1627
+ if isinstance(seed, int) and not -(2**63) <= seed < 2**63:
1628
+ raise ArenaValidationError("integer seed must fit in 64 bits.")
1629
+ mode = value.get("mode", "steps")
1630
+ if mode not in {"steps", "first_failure"}:
1631
+ raise ArenaValidationError("mode must be steps or first_failure.")
1632
+ max_steps = value.get("max_steps", min(DEFAULT_MAX_STEPS, self.max_steps_limit))
1633
+ if mode == "steps":
1634
+ if (
1635
+ isinstance(max_steps, bool)
1636
+ or not isinstance(max_steps, int)
1637
+ or not 1 <= max_steps <= self.max_steps_limit
1638
+ ):
1639
+ raise ArenaValidationError(
1640
+ f"max_steps must be between 1 and {self.max_steps_limit}."
1641
+ )
1642
+ else:
1643
+ if "max_steps" in value and value["max_steps"] is not None:
1644
+ raise ArenaValidationError(
1645
+ "max_steps must be null in first_failure mode."
1646
+ )
1647
+ max_steps = None
1648
+ return {**selected, "seed": seed, "mode": mode, "max_steps": max_steps}
1649
+
1650
+
1651
+ def encode_sse(event: dict[str, Any] | None) -> bytes:
1652
+ if event is None:
1653
+ return b": keep-alive\n\n"
1654
+ data = json.dumps(event["data"], ensure_ascii=False, separators=(",", ":"))
1655
+ return f"id: {event['id']}\nevent: {event['type']}\ndata: {data}\n\n".encode()
1656
+
1657
+
1658
+ __all__ = [
1659
+ "LOCAL_MODELS",
1660
+ "ArenaUpstreamError",
1661
+ "ArenaValidationError",
1662
+ "Competitor",
1663
+ "DecisionOutcome",
1664
+ "Game",
1665
+ "HTTPDecisionAdapter",
1666
+ "RaceManager",
1667
+ "RaceSession",
1668
+ "apply_placement",
1669
+ "build_decision_request",
1670
+ "create_piece_sequence",
1671
+ "encode_sse",
1672
+ "enumerate_placements",
1673
+ ]