Instructions to use typhoon-ai/typhoon-asr-realtime-nemo-ctc with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- NeMo
How to use typhoon-ai/typhoon-asr-realtime-nemo-ctc with NeMo:
import nemo.collections.asr as nemo_asr asr_model = nemo_asr.models.ASRModel.from_pretrained("typhoon-ai/typhoon-asr-realtime-nemo-ctc") transcriptions = asr_model.transcribe(["file.wav"]) - Notebooks
- Google Colab
- Kaggle
typhoon-asr-realtime-nemo-ctc
A Thai CTC model built entirely from standard NeMo classes built on
typhoon-ai/typhoon-asr-realtime:
EncDecCTCModelBPE whose ConformerEncoder carries the frozen 17 RT layers
plus 4 newly trained layers (26M trainable) and a ConvASRDecoder over the RT
BPE-2048 tokenizer. Loads with restore_from anywhere NeMo is installed;
using only standard NeMo modules keeps the usual NeMo/Riva/TensorRT export paths.
Results
Plain transcription (CER%, same protocol: gigaspeech2 bench in-domain / TVSpeech out-of-domain 30-s talk-show):
| model | params | gigaspeech2 CER ↓ | TVSpeech CER ↓ |
|---|---|---|---|
| typhoon-asr-realtime-nemo-ctc (greedy CTC) | 135M | 8.54 | 13.46 |
| typhoon-asr-realtime (RNN-T) | 109M | 6.89 | 9.92 |
| typhoon-whisper-medium | 769M | 4.81 | 7.66 |
| typhoon-asr-qwen-1.7b-ctx | 2.07B | 4.86 | 6.87 |
| Qwen3-ASR-1.7B (original) | 2.07B | 6.09 | 10.63 |
| Qwen3-ASR-0.6B (original) | 0.80B | 8.20 | 11.45 |
| typhoon-whisper-large-v3 | 1.55B | 4.69 | 6.32 |
| nectec/Pathumma-whisper-th-large-v3 | 1.55B | 5.84 | 10.36 |
| biodatlab/whisper-th-large-v3-combined | 1.55B | 15.78† | 14.91 |
† dominated by 3 repetition-loop utterances; excluding them: 8.39.
Read this model's CER in role: it trades a few CER points against the larger autoregressive decoders for frame-level CTC posteriors — the thing that enables keyword spotting, word boosting, and timestamps below — while adding only 26M parameters on top of the typhoon-asr-realtime encoder.
Trained on ~10,800 hours of Thai speech (GigaSpeech2-th); see the Typhoon ASR Real-time technical report.
Contextual biasing — vs every route we measured
Same benchmark (wayu-ai/thai-contextasr-bench),
same frozen scorer. Cells = CER% ↓ / entity recall ↑. Whisper models take the list
as a decoder prompt (zero-shot); typhoon-asr-qwen-1.7b-ctx is a LoRA fine-tuned for
biasing; this model spots and splices at decode time (word_boost.py,
τ=−2).
| typhoon-asr-realtime-nemo-ctc | typhoon-asr-qwen-1.7b-ctx | typhoon-whisper-medium | typhoon-whisper-large-v3 | nectec/Pathumma-whisper-th-large-v3 | biodatlab/whisper-th-large-v3-combined | Qwen3-ASR-1.7B (original) | |
|---|---|---|---|---|---|---|---|
| params | 135M | 2.07B | 769M | 1.55B | 1.55B | 1.55B | 2.07B |
| none (baseline) | 9.56 / .158 | 5.11 / .441 | 7.05 / .383 | 4.83 / .459 | 5.23 / .416 | 5.68 / .482 | 5.25 / .396 |
| bias@0 (gold) | 5.97 / .852 | 3.39 / .754 | 4.92 / .658 | N/A | 3.40 / .823 | 6.37 / .842 | 4.48 / .644 |
| bias@10 | 7.07 / .849 | 4.02 / .717 | 8.09 / .597 | N/A | 6.01 / .772 | 8.72 / .783 | 5.97 / .603 |
| bias@50 | 7.51 / .839 | 4.30 / .640 | N/A | N/A | N/A | N/A | 6.71 / .522 |
| bias@500 | 11.12 / .757 | 4.98 / .483 | N/A | N/A | N/A | N/A | 5.87 / .413 |
| distractor false alarms ↓ | .003–.006 | ≤.0004 | .003 | — | .0003 | .0033 | .0125 |
N/A: typhoon-whisper's prompt pathway collapses on any list (echo loops, CER 200+ — erased by its prompt-free fine-tune); the fine-tuned medium and Pathumma/Thonburian cannot take lists beyond ~10–15 entries because Whisper's 448-token decoder window is shared between the prompt and the transcript. This model and typhoon-asr-qwen-1.7b-ctx have no list-size limit; only they remain standing past bias@10 — and this model's recall stays flat (.85 → .76 from 0 to 500 distractors) where every other route either dies or decays.
Rule of thumb: pick this model when entity recall, list-size robustness, compute, or streaming matter — it finds more of your keywords at ~1/15 the size with zero biasing training. Pick typhoon-asr-qwen-1.7b-ctx when absolute CER and tolerance to huge dirty lists matter most. They also compose: typhoon-qwen writes the transcript, this model supplies keyword hits and timestamps from one cheap encoder pass.
Usage
from nemo.collections.asr.models import EncDecCTCModelBPE
m = EncDecCTCModelBPE.restore_from("typhoon-asr-realtime-nemo-ctc.nemo")
print(m.transcribe(["clip.wav"])[0].text)
Keyword spotting on the CTC posteriors (bundled kws_spot.py):
python kws_spot.py --model typhoon-asr-realtime-nemo-ctc.nemo \
--audio clip.wav --keywords "วโรรส,เฟซบุ๊ก"
Word boosting (bundled word_boost.py — spot then splice into the
transcript, threshold --tau):
python word_boost.py --model typhoon-asr-realtime-nemo-ctc.nemo \
--audio clip.wav --keywords "วโรรส,เฟซบุ๊ก" --tau -1.5
On thai-contextasr-bench (gold list): CER 9.56 → 5.97 with entity recall .158 → .852 at tau −2 (FA .003); tau −1.5 balanced (recall .758, FA .0014). Keep lists curated (≤50 entries) at aggressive tau.
Reference
Typhoon ASR: see the Typhoon ASR Real-time technical report.
License & attribution
CC-BY-4.0 — this checkpoint embeds encoder weights of
typhoon-ai/typhoon-asr-realtime
(CC-BY-4.0) verbatim; attribution to the Typhoon team is required. Newly
trained layers and bundled scripts follow the same license.
Acknowledgements
We thank Kunat Pipatanakul and Potsawee Manakul (wayu-ai) for their valuable feedback.
- Downloads last month
- 58
Model tree for typhoon-ai/typhoon-asr-realtime-nemo-ctc
Base model
nvidia/stt_en_fastconformer_transducer_large