Files
Agentic-AI/voice-service/README.md
T

4.1 KiB

WeLe Voice Service

Speech in, speech out. This process holds the GPU models and nothing else — it has no idea what the CRM is. Orchestration, auth and business logic stay in the Node service, so voice is a channel into the same agent, not a parallel system with its own brain.

browser ──audio──► node :4000 ──audio──► this :4100 ──► IndicConformer / Whisper
                      │                                          │
                      └────────── same graph, agents, ───────────┘
                                  guardrails as text chat
                      │
browser ◄──audio───── node ◄──audio──── this ◄── Indic Parler-TTS

Models

Job Model Notes
Endpointing Silero VAD 512-sample frames @16 kHz, 300 ms pre-roll
Indic ASR ai4bharat/indic-conformer-600m-multilingual 22 Indian languages, CTC decoding
English ASR + language ID openai/whisper-small multilingual on purpose — the .en build cannot identify languages
TTS ai4bharat/indic-parler-tts 21 languages, streaming

The AI4Bharat repos are gated. Access is auto-approved, but the download needs an authenticated account: sign in to huggingface.co, accept the terms on both model pages, then put a read token in .env as HF_TOKEN.

Why two ASR models

IndicConformer decodes as the language you name — it does not detect one, and English is not among its 22 codes. Whisper covers English and can identify the spoken language in a single decoder step. So the default mode is auto:

audio → Whisper mel + 1 decoder step → language ID
          ├─ "en"  → Whisper transcribes (mel already computed — no extra cost)
          └─ Indic → IndicConformer with the detected code

Below 0.60 confidence the caller's preferred language wins instead of a coin toss. That matters for Tanglish, where a short code-mixed sentence can honestly land either side.

Setup

python -m venv --system-site-packages .venv      # reuses the system torch build
.venv/Scripts/python -m pip install -r requirements.txt
cp .env.example .env                             # add HF_TOKEN

The venv deliberately inherits system site-packages: torch is ~2.5 GB and already installed with CUDA. Note that parler-tts pins transformers==4.46.1 inside the venv only — the system install is untouched.

Run

npm run voice          # from the parent directory
# or
.venv/Scripts/python -m app.server

First start downloads several GB and warms both models. GET /health reports device, models, sample rate and current VRAM.

Protocol

One WebSocket at /ws/voice, JSON control frames plus binary audio.

Direction Message
→ binary — 16 kHz mono PCM16 mic frames
→ {"type":"config","lang":"auto","prefer":"ta"}
→ {"type":"speak","text":"…","id":"…"}
→ {"type":"cancel"} — stop speaking now
← {"type":"speech_start"} — VAD opened a turn (drives barge-in)
← {"type":"transcript","text":…,"lang":…,"detected":…,"confidence":…}
← {"type":"audio_start","sample_rate":24000} then binary float32 chunks

Tuning

Env Default Effect
VAD_SILENCE_MS 700 trailing silence that ends a turn — lower feels snappier, truncates people who pause
VAD_MIN_SPEECH_MS 250 ignores coughs and door slams
VAD_PREFIX_MS 300 audio kept from before detection, so word onsets survive
VOICE_DEFAULT_LANG ta tiebreak when language ID is unsure
PRELOAD_ENGLISH true set false to load Whisper lazily if VRAM is tight
STT_DECODING ctc rnnt is more accurate but decodes autoregressively

VRAM

Roughly 4.7 GB of the 6 GB card with all three models resident. If that proves too tight, PRELOAD_ENGLISH=false defers Whisper (~0.5 GB) until the first English utterance.

Scripts

.venv/Scripts/python probe_access.py    # which repos the token can reach
.venv/Scripts/python probe_models.py    # load, VRAM, time-to-first-audio