typhoon-asr-realtime-nemo-ctc

A Thai CTC model built entirely from standard NeMo classes built on typhoon-ai/typhoon-asr-realtime: EncDecCTCModelBPE whose ConformerEncoder carries the frozen 17 RT layers plus 4 newly trained layers (26M trainable) and a ConvASRDecoder over the RT BPE-2048 tokenizer. Loads with restore_from anywhere NeMo is installed; using only standard NeMo modules keeps the usual NeMo/Riva/TensorRT export paths.

Results

Plain transcription (CER%, same protocol: gigaspeech2 bench in-domain / TVSpeech out-of-domain 30-s talk-show):

model params gigaspeech2 CER ↓ TVSpeech CER ↓
typhoon-asr-realtime-nemo-ctc (greedy CTC) 135M 8.54 13.46
typhoon-asr-realtime (RNN-T) 109M 6.89 9.92
typhoon-whisper-medium 769M 4.81 7.66
typhoon-asr-qwen-1.7b-ctx 2.07B 4.86 6.87
Qwen3-ASR-1.7B (original) 2.07B 6.09 10.63
Qwen3-ASR-0.6B (original) 0.80B 8.20 11.45
typhoon-whisper-large-v3 1.55B 4.69 6.32
nectec/Pathumma-whisper-th-large-v3 1.55B 5.84 10.36
biodatlab/whisper-th-large-v3-combined 1.55B 15.78† 14.91

† dominated by 3 repetition-loop utterances; excluding them: 8.39.

Read this model's CER in role: it trades a few CER points against the larger autoregressive decoders for frame-level CTC posteriors — the thing that enables keyword spotting, word boosting, and timestamps below — while adding only 26M parameters on top of the typhoon-asr-realtime encoder.

Trained on ~10,800 hours of Thai speech (GigaSpeech2-th); see the Typhoon ASR Real-time technical report.

Contextual biasing — vs every route we measured

Same benchmark (wayu-ai/thai-contextasr-bench), same frozen scorer. Cells = CER% ↓ / entity recall ↑. Whisper models take the list as a decoder prompt (zero-shot); typhoon-asr-qwen-1.7b-ctx is a LoRA fine-tuned for biasing; this model spots and splices at decode time (word_boost.py, τ=−2).

typhoon-asr-realtime-nemo-ctc typhoon-asr-qwen-1.7b-ctx typhoon-whisper-medium typhoon-whisper-large-v3 nectec/Pathumma-whisper-th-large-v3 biodatlab/whisper-th-large-v3-combined Qwen3-ASR-1.7B (original)
params 135M 2.07B 769M 1.55B 1.55B 1.55B 2.07B
none (baseline) 9.56 / .158 5.11 / .441 7.05 / .383 4.83 / .459 5.23 / .416 5.68 / .482 5.25 / .396
bias@0 (gold) 5.97 / .852 3.39 / .754 4.92 / .658 N/A 3.40 / .823 6.37 / .842 4.48 / .644
bias@10 7.07 / .849 4.02 / .717 8.09 / .597 N/A 6.01 / .772 8.72 / .783 5.97 / .603
bias@50 7.51 / .839 4.30 / .640 N/A N/A N/A N/A 6.71 / .522
bias@500 11.12 / .757 4.98 / .483 N/A N/A N/A N/A 5.87 / .413
distractor false alarms ↓ .003–.006 ≤.0004 .003 — .0003 .0033 .0125

N/A: typhoon-whisper's prompt pathway collapses on any list (echo loops, CER 200+ — erased by its prompt-free fine-tune); the fine-tuned medium and Pathumma/Thonburian cannot take lists beyond ~10–15 entries because Whisper's 448-token decoder window is shared between the prompt and the transcript. This model and typhoon-asr-qwen-1.7b-ctx have no list-size limit; only they remain standing past bias@10 — and this model's recall stays flat (.85 → .76 from 0 to 500 distractors) where every other route either dies or decays.

Recall and CER vs bias-list size: this model stays flat where whisper routes hit the prompt-window limit at N=10 and typhoon-qwen decays

Rule of thumb: pick this model when entity recall, list-size robustness, compute, or streaming matter — it finds more of your keywords at ~1/15 the size with zero biasing training. Pick typhoon-asr-qwen-1.7b-ctx when absolute CER and tolerance to huge dirty lists matter most. They also compose: typhoon-qwen writes the transcript, this model supplies keyword hits and timestamps from one cheap encoder pass.

Usage

from nemo.collections.asr.models import EncDecCTCModelBPE
m = EncDecCTCModelBPE.restore_from("typhoon-asr-realtime-nemo-ctc.nemo")
print(m.transcribe(["clip.wav"])[0].text)

Keyword spotting on the CTC posteriors (bundled kws_spot.py):

python kws_spot.py --model typhoon-asr-realtime-nemo-ctc.nemo \
    --audio clip.wav --keywords "วโรรส,เฟซบุ๊ก"

Word boosting (bundled word_boost.py — spot then splice into the transcript, threshold --tau):

python word_boost.py --model typhoon-asr-realtime-nemo-ctc.nemo \
    --audio clip.wav --keywords "วโรรส,เฟซบุ๊ก" --tau -1.5

On thai-contextasr-bench (gold list): CER 9.56 → 5.97 with entity recall .158 → .852 at tau −2 (FA .003); tau −1.5 balanced (recall .758, FA .0014). Keep lists curated (≤50 entries) at aggressive tau.

Reference

Typhoon ASR: see the Typhoon ASR Real-time technical report.

License & attribution

CC-BY-4.0 — this checkpoint embeds encoder weights of typhoon-ai/typhoon-asr-realtime (CC-BY-4.0) verbatim; attribution to the Typhoon team is required. Newly trained layers and bundled scripts follow the same license.

Acknowledgements

We thank Kunat Pipatanakul and Potsawee Manakul (wayu-ai) for their valuable feedback.

Downloads last month
58
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for typhoon-ai/typhoon-asr-realtime-nemo-ctc

Space using typhoon-ai/typhoon-asr-realtime-nemo-ctc 1

Collection including typhoon-ai/typhoon-asr-realtime-nemo-ctc

Paper for typhoon-ai/typhoon-asr-realtime-nemo-ctc