106 lines
4.1 KiB
Markdown
106 lines
4.1 KiB
Markdown
# WeLe Voice Service
|
|
|
|
Speech in, speech out. This process holds the GPU models and nothing else — it
|
|
has no idea what the CRM is. Orchestration, auth and business logic stay in the
|
|
Node service, so **voice is a channel into the same agent**, not a parallel
|
|
system with its own brain.
|
|
|
|
```
|
|
browser ──audio──► node :4000 ──audio──► this :4100 ──► IndicConformer / Whisper
|
|
│ │
|
|
└────────── same graph, agents, ───────────┘
|
|
guardrails as text chat
|
|
│
|
|
browser ◄──audio───── node ◄──audio──── this ◄── Indic Parler-TTS
|
|
```
|
|
|
|
## Models
|
|
|
|
| Job | Model | Notes |
|
|
|---|---|---|
|
|
| Endpointing | Silero VAD | 512-sample frames @16 kHz, 300 ms pre-roll |
|
|
| Indic ASR | `ai4bharat/indic-conformer-600m-multilingual` | 22 Indian languages, CTC decoding |
|
|
| English ASR + language ID | `openai/whisper-small` | multilingual on purpose — the `.en` build cannot identify languages |
|
|
| TTS | `ai4bharat/indic-parler-tts` | 21 languages, streaming |
|
|
|
|
**The AI4Bharat repos are gated.** Access is auto-approved, but the download
|
|
needs an authenticated account: sign in to huggingface.co, accept the terms on
|
|
both model pages, then put a read token in `.env` as `HF_TOKEN`.
|
|
|
|
## Why two ASR models
|
|
|
|
IndicConformer decodes *as* the language you name — it does not detect one, and
|
|
English is not among its 22 codes. Whisper covers English and can identify the
|
|
spoken language in a single decoder step. So the default mode is `auto`:
|
|
|
|
```
|
|
audio → Whisper mel + 1 decoder step → language ID
|
|
├─ "en" → Whisper transcribes (mel already computed — no extra cost)
|
|
└─ Indic → IndicConformer with the detected code
|
|
```
|
|
|
|
Below **0.60** confidence the caller's preferred language wins instead of a coin
|
|
toss. That matters for Tanglish, where a short code-mixed sentence can honestly
|
|
land either side.
|
|
|
|
## Setup
|
|
|
|
```bash
|
|
python -m venv --system-site-packages .venv # reuses the system torch build
|
|
.venv/Scripts/python -m pip install -r requirements.txt
|
|
cp .env.example .env # add HF_TOKEN
|
|
```
|
|
|
|
The venv deliberately inherits system site-packages: torch is ~2.5 GB and
|
|
already installed with CUDA. Note that `parler-tts` pins `transformers==4.46.1`
|
|
**inside the venv only** — the system install is untouched.
|
|
|
|
## Run
|
|
|
|
```bash
|
|
npm run voice # from the parent directory
|
|
# or
|
|
.venv/Scripts/python -m app.server
|
|
```
|
|
|
|
First start downloads several GB and warms both models. `GET /health` reports
|
|
device, models, sample rate and current VRAM.
|
|
|
|
## Protocol
|
|
|
|
One WebSocket at `/ws/voice`, JSON control frames plus binary audio.
|
|
|
|
| Direction | Message |
|
|
|---|---|
|
|
| → | binary — 16 kHz mono PCM16 mic frames |
|
|
| → | `{"type":"config","lang":"auto","prefer":"ta"}` |
|
|
| → | `{"type":"speak","text":"…","id":"…"}` |
|
|
| → | `{"type":"cancel"}` — stop speaking now |
|
|
| ← | `{"type":"speech_start"}` — VAD opened a turn (drives barge-in) |
|
|
| ← | `{"type":"transcript","text":…,"lang":…,"detected":…,"confidence":…}` |
|
|
| ← | `{"type":"audio_start","sample_rate":24000}` then binary float32 chunks |
|
|
|
|
## Tuning
|
|
|
|
| Env | Default | Effect |
|
|
|---|---|---|
|
|
| `VAD_SILENCE_MS` | 700 | trailing silence that ends a turn — lower feels snappier, truncates people who pause |
|
|
| `VAD_MIN_SPEECH_MS` | 250 | ignores coughs and door slams |
|
|
| `VAD_PREFIX_MS` | 300 | audio kept from before detection, so word onsets survive |
|
|
| `VOICE_DEFAULT_LANG` | `ta` | tiebreak when language ID is unsure |
|
|
| `PRELOAD_ENGLISH` | `true` | set `false` to load Whisper lazily if VRAM is tight |
|
|
| `STT_DECODING` | `ctc` | `rnnt` is more accurate but decodes autoregressively |
|
|
|
|
## VRAM
|
|
|
|
Roughly 4.7 GB of the 6 GB card with all three models resident. If that proves
|
|
too tight, `PRELOAD_ENGLISH=false` defers Whisper (~0.5 GB) until the first
|
|
English utterance.
|
|
|
|
## Scripts
|
|
|
|
```bash
|
|
.venv/Scripts/python probe_access.py # which repos the token can reach
|
|
.venv/Scripts/python probe_models.py # load, VRAM, time-to-first-audio
|
|
```
|