4.1 KiB
WeLe Voice Service
Speech in, speech out. This process holds the GPU models and nothing else — it has no idea what the CRM is. Orchestration, auth and business logic stay in the Node service, so voice is a channel into the same agent, not a parallel system with its own brain.
browser ──audio──► node :4000 ──audio──► this :4100 ──► IndicConformer / Whisper
│ │
└────────── same graph, agents, ───────────┘
guardrails as text chat
│
browser ◄──audio───── node ◄──audio──── this ◄── Indic Parler-TTS
Models
| Job | Model | Notes |
|---|---|---|
| Endpointing | Silero VAD | 512-sample frames @16 kHz, 300 ms pre-roll |
| Indic ASR | ai4bharat/indic-conformer-600m-multilingual |
22 Indian languages, CTC decoding |
| English ASR + language ID | openai/whisper-small |
multilingual on purpose — the .en build cannot identify languages |
| TTS | ai4bharat/indic-parler-tts |
21 languages, streaming |
The AI4Bharat repos are gated. Access is auto-approved, but the download
needs an authenticated account: sign in to huggingface.co, accept the terms on
both model pages, then put a read token in .env as HF_TOKEN.
Why two ASR models
IndicConformer decodes as the language you name — it does not detect one, and
English is not among its 22 codes. Whisper covers English and can identify the
spoken language in a single decoder step. So the default mode is auto:
audio → Whisper mel + 1 decoder step → language ID
├─ "en" → Whisper transcribes (mel already computed — no extra cost)
└─ Indic → IndicConformer with the detected code
Below 0.60 confidence the caller's preferred language wins instead of a coin toss. That matters for Tanglish, where a short code-mixed sentence can honestly land either side.
Setup
python -m venv --system-site-packages .venv # reuses the system torch build
.venv/Scripts/python -m pip install -r requirements.txt
cp .env.example .env # add HF_TOKEN
The venv deliberately inherits system site-packages: torch is ~2.5 GB and
already installed with CUDA. Note that parler-tts pins transformers==4.46.1
inside the venv only — the system install is untouched.
Run
npm run voice # from the parent directory
# or
.venv/Scripts/python -m app.server
First start downloads several GB and warms both models. GET /health reports
device, models, sample rate and current VRAM.
Protocol
One WebSocket at /ws/voice, JSON control frames plus binary audio.
| Direction | Message |
|---|---|
| → | binary — 16 kHz mono PCM16 mic frames |
| → | {"type":"config","lang":"auto","prefer":"ta"} |
| → | {"type":"speak","text":"…","id":"…"} |
| → | {"type":"cancel"} — stop speaking now |
| ← | {"type":"speech_start"} — VAD opened a turn (drives barge-in) |
| ← | {"type":"transcript","text":…,"lang":…,"detected":…,"confidence":…} |
| ← | {"type":"audio_start","sample_rate":24000} then binary float32 chunks |
Tuning
| Env | Default | Effect |
|---|---|---|
VAD_SILENCE_MS |
700 | trailing silence that ends a turn — lower feels snappier, truncates people who pause |
VAD_MIN_SPEECH_MS |
250 | ignores coughs and door slams |
VAD_PREFIX_MS |
300 | audio kept from before detection, so word onsets survive |
VOICE_DEFAULT_LANG |
ta |
tiebreak when language ID is unsure |
PRELOAD_ENGLISH |
true |
set false to load Whisper lazily if VRAM is tight |
STT_DECODING |
ctc |
rnnt is more accurate but decodes autoregressively |
VRAM
Roughly 4.7 GB of the 6 GB card with all three models resident. If that proves
too tight, PRELOAD_ENGLISH=false defers Whisper (~0.5 GB) until the first
English utterance.
Scripts
.venv/Scripts/python probe_access.py # which repos the token can reach
.venv/Scripts/python probe_models.py # load, VRAM, time-to-first-audio