Kviskr punctuation (Norwegian), INT8 ONNX
Superseded by
skolfus/kviskr-tegnsetting-no, which is about 14× smaller, scores higher on Norwegian dictation and spontaneous speech, and is built so it cannot change a word. This model is kept as a research baseline.
An INT8-quantized ONNX export of 1-800-BAD-CODE/xlm-roberta_punctuation_fullstop_truecase (XLM-RoBERTa token classifier, 47 languages), evaluated on Norwegian dictation. Boolean outputs are cast to int64 for ONNX Runtime on macOS.
Files
| File | Contents |
|---|---|
model_embed.onnx |
INT8-quantized model (~279 MB) |
tokenizer.json |
XLM-R tokenizer (swift-transformers compatible) |
tokenizer_config.json, config.json |
Tokenizer and model config |
Evaluation (10 Jun 2026)
Measured on real verbatim ASR output (NB-Whisper large verbatim) over the National Library's evaluation set (200 clips, many speakers, five domains) and a labelled dictation set:
- Sentence-end F1 0.81, comma F1 0.53, truecasing 97.0%
- Case and punctuation error rate on dictation: 29.6% (raw) → 9.8%
Known issue
The model classifies subwords, so it can break abbreviations: bl.a. → b.l...a., m.v. → m...v.. On a held-out text benchmark this affected 9.3% of chunks (279 of 3,000). On ASR text it is rare (1 of 327 clips), since the speech model writes «blant annet», not «bl.a.». The successor model fixes this by predicting per word.
License
Apache 2.0, the same as the base model. Quantization and evaluation by Kviskr.
- Downloads last month
- 56