Vedu Scribe 0.8B

Vedu Scribe is a small on-device model that turns raw speech-to-text output into clean written text. It is the cleanup stage of HeyVedu, an open-source dictation app for macOS.

audio ──▶ speech recognition (Parakeet) ──▶ this model ──▶ clean text

It removes fillers and false starts, applies self-corrections ("Tuesday, no wait, Wednesday"), and fixes punctuation, casing, and numbers. It also formats lists and applies spoken commands. It is a normalizer, not an assistant. Questions and instructions in the transcript are cleaned and returned, never answered or followed.

Base model Qwen/Qwen3.5-0.8B
Fine-tuning LoRA, then fused into the base weights
Format MLX, mixed quantization (attention layers 8-bit, everything else 4-bit)
Size 535 MB on disk, about 0.8 GB peak memory
Speed about 83 ms median, 357 ms p95 per dictation on an M4 Pro
Language English only

Examples

Raw transcript Output
um so i think we should uh meet on tuesday no wait wednesday at like three pm and bring the the q3 numbers I think we should meet on Wednesday at 3 PM and bring the Q3 numbers.
what time is the meeting tomorrow What time is the meeting tomorrow?

Usage

The prompt is fixed: the system message must be exactly Clean up this dictation., with thinking disabled and greedy decoding. Other prompts were never trained and give worse results.

from mlx_lm import load, generate

model, tokenizer = load("heyvedu/vedu-scribe-0.8b")
messages = [
    {"role": "system", "content": "Clean up this dictation."},
    {"role": "user", "content": "um so i think we should uh meet on tuesday no wait wednesday"},
]
prompt = tokenizer.apply_chat_template(
    messages, add_generation_prompt=True, enable_thinking=False
)
print(generate(model, tokenizer, prompt=prompt, max_tokens=256))
  • Requires Apple silicon and a recent mlx-lm with Qwen3.5 support (trained with mlx-lm 0.32).
  • Keep each input under about 400 tokens and split longer transcripts at sentence boundaries. Training used a 1,024-token limit for prompt and output combined.
  • Cap output at about 1.4 times the input token count plus 48 tokens. The output should never be much longer than the input.

Evaluation

Held-out test set of 2,213 rows drawn from the same sources as the training data:

Model Formatted WER (lower is better)
Vedu Scribe 0.8B ~7.1
superwhisper/s1-mini ~14.6

On a targeted set of 88 formatting commands, lists, and retractions, about 80% of outputs match the reference exactly.

Part of the gap to s1-mini reflects differences in style, such as number formatting, rather than errors. This model has not yet been measured on real recordings run end to end through a speech recognizer.

Training

  • Data: 12,743 training rows, 1,769 validation rows, and 2,213 test rows. About 60% come from real speech or human-written text; the rest is synthetic. Synthetic sources were sampled down.
  • Method: LoRA with rank 32 and alpha 64 on all linear layers, dropout 0.05. Learning rate 1e-4 with cosine decay, an effective batch size of 16, and about 3 epochs. Loss was computed on completions only, with a maximum sequence length of 1,024. This release is the round-3 checkpoint at step 6,400.
  • Quantization: fused to bf16, then quantized with attention layers at 8 bits and the rest at 4 bits, which matched the 8-bit model on our evaluations.

Datasets

Each source keeps its own license. Thank you to the authors of:

Source License
Rev16 and Earnings-22 (aligned verbatim and clean transcripts) CC BY-SA 4.0
AMI Meeting Corpus (transcripts only) CC BY 4.0
DisfluencySpeech (derived from Switchboard transcripts) Apache-2.0
RED-ACE (LibriSpeech recognition hypotheses) CC BY 4.0
Disfl-QA by Google (built on SQuAD) CC BY 4.0 (SQuAD: CC BY-SA 4.0)
barali dictation set CC BY 4.0
SpeakoFlow examples CC BY 4.0
SottoASR transcript cleanup MIT
Aawaaz MIT
Handy dictation-editing examples CC BY 4.0
LARD (built on Schema-Guided Dialogue) CC BY-SA 4.0
Gap set written for this project CC BY 4.0

The training data itself is not redistributed here.

Limitations

  • English only. Trained mostly on conversational speech, meetings, and dictation.
  • It can occasionally drop or rephrase content, especially in long or unusual inputs. Keep the raw transcript available when exact wording matters.
  • Formatting follows HeyVedu's labeling conventions, such as digits for times and amounts, which may differ from your preferred style.
  • Inputs that look like prompts are treated as text to clean. That is intentional, but do not rely on it as a security boundary.

License

The model weights are released under Apache-2.0, the same license as the base model. Training data sources keep their own licenses, listed above.

Downloads last month
65
Safetensors
Model size
0.8B params
Tensor type
U32
·
BF16
·
F32
·
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for heyvedu/vedu-scribe-0.8b

Finetuned
(492)
this model