Instructions to use heyvedu/vedu-scribe-0.8b with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use heyvedu/vedu-scribe-0.8b with MLX:
# Make sure mlx-lm is installed # pip install --upgrade mlx-lm # Generate text with mlx-lm from mlx_lm import load, generate model, tokenizer = load("heyvedu/vedu-scribe-0.8b") prompt = "Write a story about Einstein" messages = [{"role": "user", "content": prompt}] prompt = tokenizer.apply_chat_template( messages, add_generation_prompt=True ) text = generate(model, tokenizer, prompt=prompt, verbose=True) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Pi
How to use heyvedu/vedu-scribe-0.8b with Pi:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "heyvedu/vedu-scribe-0.8b"
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "mlx-lm": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "heyvedu/vedu-scribe-0.8b" } ] } } }Run Pi
# Start Pi in your project directory: pi
- MLX LM
How to use heyvedu/vedu-scribe-0.8b with MLX LM:
Generate or start a chat session
# Install MLX LM uv tool install mlx-lm # Interactive chat REPL mlx_lm.chat --model "heyvedu/vedu-scribe-0.8b"
Run an OpenAI-compatible server
# Install MLX LM uv tool install mlx-lm # Start the server mlx_lm.server --model "heyvedu/vedu-scribe-0.8b" # Calling the OpenAI-compatible server with curl curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "heyvedu/vedu-scribe-0.8b", "messages": [ {"role": "user", "content": "Hello"} ] }' - Hermes Agent
How to use heyvedu/vedu-scribe-0.8b with Hermes Agent:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "heyvedu/vedu-scribe-0.8b"
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default heyvedu/vedu-scribe-0.8b
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use heyvedu/vedu-scribe-0.8b with OpenClaw:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "heyvedu/vedu-scribe-0.8b"
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "heyvedu/vedu-scribe-0.8b" \ --custom-provider-id mlx-lm \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Vedu Scribe 0.8B
Vedu Scribe is a small on-device model that turns raw speech-to-text output into clean written text. It is the cleanup stage of HeyVedu, an open-source dictation app for macOS.
audio ──▶ speech recognition (Parakeet) ──▶ this model ──▶ clean text
It removes fillers and false starts, applies self-corrections ("Tuesday, no wait, Wednesday"), and fixes punctuation, casing, and numbers. It also formats lists and applies spoken commands. It is a normalizer, not an assistant. Questions and instructions in the transcript are cleaned and returned, never answered or followed.
| Base model | Qwen/Qwen3.5-0.8B |
| Fine-tuning | LoRA, then fused into the base weights |
| Format | MLX, mixed quantization (attention layers 8-bit, everything else 4-bit) |
| Size | 535 MB on disk, about 0.8 GB peak memory |
| Speed | about 83 ms median, 357 ms p95 per dictation on an M4 Pro |
| Language | English only |
Examples
| Raw transcript | Output |
|---|---|
um so i think we should uh meet on tuesday no wait wednesday at like three pm and bring the the q3 numbers |
I think we should meet on Wednesday at 3 PM and bring the Q3 numbers. |
what time is the meeting tomorrow |
What time is the meeting tomorrow? |
Usage
The prompt is fixed: the system message must be exactly Clean up this dictation., with
thinking disabled and greedy decoding. Other prompts were never trained and give worse results.
from mlx_lm import load, generate
model, tokenizer = load("heyvedu/vedu-scribe-0.8b")
messages = [
{"role": "system", "content": "Clean up this dictation."},
{"role": "user", "content": "um so i think we should uh meet on tuesday no wait wednesday"},
]
prompt = tokenizer.apply_chat_template(
messages, add_generation_prompt=True, enable_thinking=False
)
print(generate(model, tokenizer, prompt=prompt, max_tokens=256))
- Requires Apple silicon and a recent
mlx-lmwith Qwen3.5 support (trained withmlx-lm0.32). - Keep each input under about 400 tokens and split longer transcripts at sentence boundaries. Training used a 1,024-token limit for prompt and output combined.
- Cap output at about 1.4 times the input token count plus 48 tokens. The output should never be much longer than the input.
Evaluation
Held-out test set of 2,213 rows drawn from the same sources as the training data:
| Model | Formatted WER (lower is better) |
|---|---|
| Vedu Scribe 0.8B | ~7.1 |
| superwhisper/s1-mini | ~14.6 |
On a targeted set of 88 formatting commands, lists, and retractions, about 80% of outputs match the reference exactly.
Part of the gap to s1-mini reflects differences in style, such as number formatting, rather than errors. This model has not yet been measured on real recordings run end to end through a speech recognizer.
Training
- Data: 12,743 training rows, 1,769 validation rows, and 2,213 test rows. About 60% come from real speech or human-written text; the rest is synthetic. Synthetic sources were sampled down.
- Method: LoRA with rank 32 and alpha 64 on all linear layers, dropout 0.05. Learning rate 1e-4 with cosine decay, an effective batch size of 16, and about 3 epochs. Loss was computed on completions only, with a maximum sequence length of 1,024. This release is the round-3 checkpoint at step 6,400.
- Quantization: fused to bf16, then quantized with attention layers at 8 bits and the rest at 4 bits, which matched the 8-bit model on our evaluations.
Datasets
Each source keeps its own license. Thank you to the authors of:
| Source | License |
|---|---|
| Rev16 and Earnings-22 (aligned verbatim and clean transcripts) | CC BY-SA 4.0 |
| AMI Meeting Corpus (transcripts only) | CC BY 4.0 |
| DisfluencySpeech (derived from Switchboard transcripts) | Apache-2.0 |
| RED-ACE (LibriSpeech recognition hypotheses) | CC BY 4.0 |
| Disfl-QA by Google (built on SQuAD) | CC BY 4.0 (SQuAD: CC BY-SA 4.0) |
| barali dictation set | CC BY 4.0 |
| SpeakoFlow examples | CC BY 4.0 |
| SottoASR transcript cleanup | MIT |
| Aawaaz | MIT |
| Handy dictation-editing examples | CC BY 4.0 |
| LARD (built on Schema-Guided Dialogue) | CC BY-SA 4.0 |
| Gap set written for this project | CC BY 4.0 |
The training data itself is not redistributed here.
Limitations
- English only. Trained mostly on conversational speech, meetings, and dictation.
- It can occasionally drop or rephrase content, especially in long or unusual inputs. Keep the raw transcript available when exact wording matters.
- Formatting follows HeyVedu's labeling conventions, such as digits for times and amounts, which may differ from your preferred style.
- Inputs that look like prompts are treated as text to clean. That is intentional, but do not rely on it as a security boundary.
License
The model weights are released under Apache-2.0, the same license as the base model. Training data sources keep their own licenses, listed above.
- Downloads last month
- 65
4-bit