Danish punctuation restoration and truecasing

This model adds punctuation (, . ? !) and capitalisation to lowercase, unpunctuated Danish text, such as the output of a speech recogniser. Words are never changed, added or removed, so it can be used on transcripts with word timestamps. The architecture, base model and label scheme are identical to RyeAI/ekko-pnc; only the training data differs.

Input: hej mit navn er emil hvordan går det i dag jeg har været på kontoret og bagefter skal jeg hjem

Output: Hej, mit navn er Emil. Hvordan går det i dag? Jeg har været på kontoret, og bagefter skal jeg hjem.

How to use

With transformers (PyTorch)

from transformers import AutoTokenizer, AutoModelForTokenClassification
import torch

repo = "syvai/danish-pnc"
tok = AutoTokenizer.from_pretrained(repo)
model = AutoModelForTokenClassification.from_pretrained(repo).eval()
labels = model.config.id2label

def punctuate(text):
    words = text.split()
    enc = tok(words, is_split_into_words=True, return_tensors="pt", truncation=True, max_length=256)
    with torch.no_grad():
        pred = model(**enc).logits[0].argmax(-1).tolist()
    out, seen, cap_next = [], set(), True
    for ti, wid in enumerate(enc.word_ids()):
        if wid is None or wid in seen:
            continue
        seen.add(wid)
        lab = labels[pred[ti]]                 # e.g. "O", ",|U", ".", "?|U"
        punct = "" if lab[0] == "O" else lab[0]
        w = words[wid]
        if cap_next or "|U" in lab:
            w = w[0].upper() + w[1:]
        out.append(w + punct)
        cap_next = punct in (".", "?", "!")
    return " ".join(out)

print(punctuate("hej mit navn er emil hvordan går det i dag"))

For texts longer than 256 tokens, use the sliding-window helper pnc_infer.py in this repository (128-token windows with a 12-word overlap, every input word kept in order):

python pnc_infer.py syvai/danish-pnc "hej mit navn er emil hvordan går det i dag"

With ONNX Runtime

pnc.onnx (FP32) and pnc.int8.onnx (dynamic int8, about 22 MB) take input_ids, attention_mask and token_type_ids and return logits of shape [batch, seq, 10].

import onnxruntime as ort, numpy as np
from transformers import AutoTokenizer
tok = AutoTokenizer.from_pretrained("syvai/danish-pnc")
sess = ort.InferenceSession("pnc.int8.onnx")
enc = tok("hej mit navn er emil".split(), is_split_into_words=True, return_tensors="np")
logits = sess.run(None, {k: enc[k] for k in ["input_ids", "attention_mask", "token_type_ids"]})[0]
pred = logits[0].argmax(-1)   # map with config.json id2label, first sub-token of each word

Labels

Ten labels, one per word (predicted on the word's first sub-token): punctuation after the word (O none, ,, ., ?, !) combined with |U when the word should start with an uppercase letter. The model produces an initial capital only; it cannot restore mixed case such as iPhone.

Training data

All data was streamed from the Hugging Face Hub; nothing was used in full. Sources were chosen for modern, human-punctuated Danish and mixed with the following sampling weights:

Source Weight Register
Danish Dynaword opensubtitles 0.24 Film and TV dialogue
fineweb-2 dan_Latn 0.24 Web text
Dynaword ft 0.12 Folketinget transcripts
Dynaword wikipedia 0.08 Encyclopaedia
Dynaword ep 0.07 Europarl transcripts
Dynaword hest 0.07 Forum posts
Dynaword nordjyllandnews, tv2r 0.09 News
Dynaword mosel_voxpopuli 0.05 Parliament speech transcripts
Dynaword danske-taler, wiki-comments 0.04 Speeches, talk pages

Text was lowercased and stripped of punctuation to form the input; the removed punctuation and original capitalisation form the labels. Speaker tags, dialogue dashes, quotes and URLs were removed; semicolons map to commas; colons map to a full stop or comma depending on the following word; abbreviations and ordinals (kr., 8. oktober) are not treated as sentence ends. Documents with implausible punctuation density, heavy capitalisation or pre-1948 orthography were dropped. Training windows were 16 to 180 words starting at random offsets so the model learns to handle chunks that do not begin at a sentence start.

Training: ElectraForTokenClassification from jonfd/electra-small-nordic, 150k steps, batch 32, max length 256, AdamW lr 1e-4 with 3k warmup and cosine decay, weight decay 0.01, fp32 on an Apple M4. The data pipeline and training code are in data.py in this repository.

Evaluation

Held-out set of 1,750 windows (24 to 160 words) from documents excluded from training by id hash, spanning all eleven sources. The reference model was scored on exactly the same inputs with the same code (eval_results.json has the raw numbers, eval_log.jsonl the hourly training curve).

Metric RyeAI/ekko-pnc This model
Sentence boundary F1 0.722 0.859
Boundary precision 0.735 0.874
Boundary recall 0.710 0.844
Comma F1 0.670 0.787
Full stop F1 0.696 0.825
Question F1 0.479 0.756
Exclamation F1 0.000 0.271
Casing F1 0.797 0.898
Token accuracy 0.875 0.927

Sentence boundary F1 per source:

Source RyeAI/ekko-pnc This model
opensubtitles 0.724 0.892
fineweb2 0.654 0.776
ft 0.794 0.948
ep 0.834 0.956
wikipedia 0.756 0.894
hest 0.579 0.682
nordjyllandnews 0.782 0.885
tv2r 0.807 0.876
mosel_voxpopuli 0.729 0.733
danske-taler 0.736 0.771
wiki-comments 0.716 0.835

Caveats. This eval set shares corpora and labelling conventions with the training data, which favours this model. RyeAI/ekko-pnc was tuned for continuous meeting speech and evaluated on human-reviewed meeting transcripts, which were not available here. The mosel_voxpopuli row, the only genuine speech-transcript source, is the fairest single comparison. Exclamation marks are rare in the eval set and that score is noisy.

Limitations

  • Punctuation is predicted from words alone; the model cannot hear pauses or intonation.
  • Question marks and exclamation marks are the weakest classes.
  • Capitalisation is initial-uppercase only. Keep existing capitals from the input if you have them.
  • Trained on edited transcripts and written text; spontaneous conversational speech with disfluencies is under-represented.

License

CC BY 4.0 (see LICENSE). The base model jonfd/electra-small-nordic and the training sources (Danish Dynaword, CC0; fineweb-2, ODC-By) carry their own terms.

Downloads last month
18
Safetensors
Model size
21.9M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for syvai/danish-pnc

Finetuned
(4)
this model

Datasets used to train syvai/danish-pnc