# WeLe Voice Service Speech in, speech out. This process holds the GPU models and nothing else — it has no idea what the CRM is. Orchestration, auth and business logic stay in the Node service, so **voice is a channel into the same agent**, not a parallel system with its own brain. ``` browser ──audio──► node :4000 ──audio──► this :4100 ──► IndicConformer / Whisper │ │ └────────── same graph, agents, ───────────┘ guardrails as text chat │ browser ◄──audio───── node ◄──audio──── this ◄── Indic Parler-TTS ``` ## Models | Job | Model | Notes | |---|---|---| | Endpointing | Silero VAD | 512-sample frames @16 kHz, 300 ms pre-roll | | Indic ASR | `ai4bharat/indic-conformer-600m-multilingual` | 22 Indian languages, CTC decoding | | English ASR + language ID | `openai/whisper-small` | multilingual on purpose — the `.en` build cannot identify languages | | TTS | `ai4bharat/indic-parler-tts` | 21 languages, streaming | **The AI4Bharat repos are gated.** Access is auto-approved, but the download needs an authenticated account: sign in to huggingface.co, accept the terms on both model pages, then put a read token in `.env` as `HF_TOKEN`. ## Why two ASR models IndicConformer decodes *as* the language you name — it does not detect one, and English is not among its 22 codes. Whisper covers English and can identify the spoken language in a single decoder step. So the default mode is `auto`: ``` audio → Whisper mel + 1 decoder step → language ID ├─ "en" → Whisper transcribes (mel already computed — no extra cost) └─ Indic → IndicConformer with the detected code ``` Below **0.60** confidence the caller's preferred language wins instead of a coin toss. That matters for Tanglish, where a short code-mixed sentence can honestly land either side. ## Setup ```bash python -m venv --system-site-packages .venv # reuses the system torch build .venv/Scripts/python -m pip install -r requirements.txt cp .env.example .env # add HF_TOKEN ``` The venv deliberately inherits system site-packages: torch is ~2.5 GB and already installed with CUDA. Note that `parler-tts` pins `transformers==4.46.1` **inside the venv only** — the system install is untouched. ## Run ```bash npm run voice # from the parent directory # or .venv/Scripts/python -m app.server ``` First start downloads several GB and warms both models. `GET /health` reports device, models, sample rate and current VRAM. ## Protocol One WebSocket at `/ws/voice`, JSON control frames plus binary audio. | Direction | Message | |---|---| | → | binary — 16 kHz mono PCM16 mic frames | | → | `{"type":"config","lang":"auto","prefer":"ta"}` | | → | `{"type":"speak","text":"…","id":"…"}` | | → | `{"type":"cancel"}` — stop speaking now | | ← | `{"type":"speech_start"}` — VAD opened a turn (drives barge-in) | | ← | `{"type":"transcript","text":…,"lang":…,"detected":…,"confidence":…}` | | ← | `{"type":"audio_start","sample_rate":24000}` then binary float32 chunks | ## Tuning | Env | Default | Effect | |---|---|---| | `VAD_SILENCE_MS` | 700 | trailing silence that ends a turn — lower feels snappier, truncates people who pause | | `VAD_MIN_SPEECH_MS` | 250 | ignores coughs and door slams | | `VAD_PREFIX_MS` | 300 | audio kept from before detection, so word onsets survive | | `VOICE_DEFAULT_LANG` | `ta` | tiebreak when language ID is unsure | | `PRELOAD_ENGLISH` | `true` | set `false` to load Whisper lazily if VRAM is tight | | `STT_DECODING` | `ctc` | `rnnt` is more accurate but decodes autoregressively | ## VRAM Roughly 4.7 GB of the 6 GB card with all three models resident. If that proves too tight, `PRELOAD_ENGLISH=false` defers Whisper (~0.5 GB) until the first English utterance. ## Scripts ```bash .venv/Scripts/python probe_access.py # which repos the token can reach .venv/Scripts/python probe_models.py # load, VRAM, time-to-first-audio ```