You need to agree to share your contact information to access this model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

Log in or Sign Up to review the conditions and access this model content.

DhVaani — TTS for 27 Indian languages

DhVaani is a fast, zero-shot voice-cloning text-to-speech model covering 27 Indian languages. Give it a few seconds of reference audio + its transcript and any target text, and it speaks that text in the reference voice — at 24 kHz.

Fine-tuned from ZipVoice (a 123M-parameter flow-matching TTS), with a character-level tokenizer that handles Indian script.

Quick start

# install torch/torchaudio for your machine first (GPU shown; use /whl/cpu for CPU):
pip install torch torchaudio 
pip install -r requirements.txt

transformers AutoModel (recommended)

import os,torch, soundfile as sf
from transformers import AutoModel
os.environ.setdefault("HF_TOKEN", "<your-hf-token>")
DEV = "cuda" if torch.cuda.is_available() else "cpu"
model = AutoModel.from_pretrained("ARTPARK-IISc/DhVaani-0.5",
                                  trust_remote_code=True).to(DEV).eval()

audio = model.synthesize(
    text="നമസ്‌കാരം, സുഖമാണോ?",
    prompt_wav="reference.wav",          # a few seconds of the voice to clone
    prompt_text="reference transcript",
)                                        # float32 numpy @ model.sampling_rate (24 kHz)
sf.write("out.wav", audio, model.sampling_rate)

Or standalone (no transformers)

from dhvaani import DhVaani
tts = DhVaani()                          # auto GPU/CPU
tts.synthesize(text="வணக்கம்!", prompt_wav="ref.wav",
               prompt_text="…", out_path="out.wav")
python dhvaani.py --prompt-wav ref.wav --prompt-text "…" --text "…" --out out.wav

Knobs: num_step (16 = quality, 8 = faster), guidance_scale, speed, seed.

Languages & data

Trained pooled from multiple TTS datasets like IISc SYSPIN, IndicTTS (SPRINGLab)and Rasa at 16 kHz.

well-represented mid low-resource
English, Hindi, Bengali, Marathi, Kannada, Telugu, Maithili, Magahi, Chhattisgarhi, Bhojpuri, Assamese, Tamil Gujarati, Malayalam, Manipuri, Odia, Punjabi, Nepali, Sindhi, Konkani, Santali, Bodo, Urdu Sanskrit, Dogri, Kashmiri, Rajasthani

Scripts: Devanagari, Bengali–Assamese, Tamil, Telugu, Kannada, Malayalam, Gujarati, Gurmukhi, Odia, Perso-Arabic, Ol Chiki, Latin.

What's in this repo

file what
config.json · modeling_dhvaani.py the AutoModel (trust_remote_code) wrapper
dhvaani.py standalone API + CLI (no transformers needed)
model.safetensors weights (491 MB)
model.json · tokens.txt architecture config · 1058-char vocabulary
requirements.txt dependencies
_backend/ the trimmed, inference-only model code DhVaani runs on
samples/ example outputs

Training

  • Fine-tuned from k2-fsa/ZipVoice; the base Chinese+English text embedding (360 tokens) was re-initialised for the 1058-token Indic character vocab, the flow-matching acoustic backbone transferred.

Limitations

  • Code-switching (English words inside Indic text) runs — Latin is in the vocab — but wasn't explicitly trained, so mixed-language prosody can be rough at boundaries.
  • Data imbalance: low-resource languages (Rajasthani, Kashmiri, Dogri, Sanskrit) are ~10× smaller and weaker.
  • Out-of-vocabulary characters (rare diacritics, emoji) are silently dropped.
  • Best with a clean, single-speaker, 3–10 s reference clip.

Credits & license

Apache-2.0, following the base model. Built on ZipVoice (k2-fsa). Please also respect the licenses of the training corpora (IndicTTS, Rasa, IISc SYSPIN).

Downloads last month
92
Safetensors
Model size
0.1B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for ARTPARK-IISc/DhVaani-0.5

Base model

k2-fsa/ZipVoice
Finetuned
(4)
this model

Space using ARTPARK-IISc/DhVaani-0.5 1