OpenJev 9B
OpenJev 9B is an open, locally deployable typed decision model for choice,
noul (yes/no/unknown-style decisions), and ordinal score questions over a
shared state. It uses Qwen3.5-9B-Base with Kev's pointer-style decision head and
a local Go API boundary that isolates sibling questions.
This release does not claim parity with hosted Jev. No Jev output was used as a training label, soft target, synthetic teacher, calibration target, selection signal, or fallback response. No fresh same-item Jev evaluation was run for this checkpoint.
Identity
| Field | Value |
|---|---|
| Model | openjev-9b-hard-final-w25 |
| Base | Qwen/Qwen3.5-9B-Base@68c46c4b3498877f3ef123c856ecfde50c39f404 |
| Kev code | jaredpalmer/kev@0fe8fc97c2bcc247fa3efb6e5c32af4e99770e91 |
| Weights SHA256 | 67b87afe12689bf6a95f6669d64250eae70e8744d0700019220942396355f8b8 |
| Head SHA256 | 56910f6a89bfb0cc5642764d7f983f5bad3d025dabcbad38590765455e7622ad |
| Temperature | 1.9543199368684914 |
| Precision | BF16 |
| Maximum state | 16,384 tokens in the validated serving contract |
| Question semantics | strict isolated execution |
The four model-*.safetensors files contain merged full weights. head.pt
contains the typed Kev head. SHA256SUMS binds every checkpoint file.
Run locally
Use the exact Kev revision for the reference backend:
git clone https://github.com/jaredpalmer/kev.git
cd kev
git checkout 0fe8fc97c2bcc247fa3efb6e5c32af4e99770e91
uv sync --extra serve
KEV_TEMPERATURE=1.9543199368684914 \
uv run --extra serve python -m kev.serve \
--run iFurySt/openjev --port 8008
The Kev endpoint implements the System One-shaped decision API. The OpenJev source tree additionally contains a Go boundary with input validation, authentication, overload handling, question isolation, readiness, benchmarks, and an offline deployment layout. It has no Jev key or cloud fallback path.
Training method
The selected checkpoint has three relevant stages:
- A rank-16 LoRA continuation mixed 8,000 audited local records with 8,000 deterministic Kev replay records for one pass.
- A hard-label full-weight BF16 continuation ran one pass over 11,320 records
/ 16,279 questions using six A100-SXM4-80GB GPUs, FSDP2, seed
92323, and LR/head LR1e-6. - R2 froze six WiSE-FT candidates and selected a 75% local parent / 25% hard-final interpolation using DEVELOPMENT and CALIBRATION only. This stage adds no optimizer update or new label.
A matched Qwen3.5-27B teacher research arm was trained during R1, but the released checkpoint uses the hard-label arm, not that teacher arm. Historical K0 data did contain some local-27B soft targets; those response rows are not redistributed and are disclosed as a byte-reproduction gap.
See TRAINING.md and DATA.md in this repository for the full round ledger,
data counts, source revisions, licenses, negative results, and hardware.
Quality evidence
The independently opened capability-v1 confirmation contains 1,200 records / 1,920 questions. On the 1,800-question clean paired population:
| Model | Accuracy | Multiclass Brier | Coverage at <=5% empirical error |
|---|---|---|---|
| OpenJev 9B | 0.93611 | 0.08726 | 0.97611 |
| K0 parent | 0.94556 | 0.07848 | 0.98944 |
| Raw Qwen3.5-9B scoring adapter | 0.81333 | 0.28569 | 0.30889 |
These are synthetic confirmation results, not proof of general intelligence or
Jev parity. A later corrected DEVELOPMENT-only comparison on true released
Kev-9B found OpenJev / Kev accuracy 0.757624 / 0.738898 and Brier
0.328546 / 0.355000; it is not a locked test.
Serving measurements
The same multilingual three-question Go request was measured after two warmups and 20 requests at each concurrency. Every one of 160 measured CPU/GPU requests returned HTTP 200.
| Device | c1 throughput / p50 | c8 throughput / p50 | Peak model memory |
|---|---|---|---|
| 16-core Xeon CPU | 0.5699 req/s / 1,692.7 ms | 0.5037 req/s / 15,678.8 ms | 13.7 GiB RSS |
| A100-SXM4-80GB | 10.7123 req/s / 85.7 ms | 20.0055 req/s / 342.7 ms | 23.6 GiB device |
CPU is appropriate for offline, fallback, or very low-QPS use. A GPU is recommended for interactive concurrent serving. Discrete CPU/GPU decisions matched; probability bits are not promised identical across devices or dynamic batches.
Data and reproducibility
OpenJev publishes source builders, pinned revisions, hashes, roles, and attribution. Original deterministic fixtures can be released under their declared terms. Third-party rows that require per-origin review, contain a benchmark canary, or have share-alike/attribution constraints remain download/build-recipe-only. Jev responses and private teacher-response ledgers are excluded.
That means the model is open and independently verifiable, while exact
from-scratch byte reproduction still has a documented gap around historical K0
rows/targets. See DATA.md; do not infer that all upstream dataset rows are
owned by OpenJev.
Limitations
- No fresh Jev baseline and no parity claim.
- The model inherits biases and factual limitations from Qwen, Kev, and the listed training sources.
- Numeric/date/business-day and long-context joins remain known weak areas in consumed reporting-only diagnostics; those rows were not mined for training.
- Concurrent GPU batching can change rounded probabilities by up to 0.0018 in the measured fixture without changing the discrete decision.
- Active GPU compute is not preempted on client cancellation; queued work can be cancelled cooperatively.
- The API is close to, but not certified identical to, every hosted Jev error, usage, refusal, and rate-limit behavior.
License and attribution
OpenJev code and released weights are provided under Apache-2.0. This model is
derived from Apache-2.0 Qwen and Kev assets. See LICENSE and NOTICE.
“Jev” and “TypeSafe” are used only to describe compatibility and the research
target; this project is not affiliated with or endorsed by TypeSafe.
- Downloads last month
- 13
Model tree for iFurySt/openjev
Base model
Qwen/Qwen3.5-9B-Base