WeLe Agentic AI: LangGraph multi-agent CRM assistant with voice

This commit is contained in:
2026-08-28 02:16:03 +05:30
commit 105e58e02a
69 changed files with 11501 additions and 0 deletions
+105
View File
@@ -0,0 +1,105 @@
# WeLe Voice Service
Speech in, speech out. This process holds the GPU models and nothing else — it
has no idea what the CRM is. Orchestration, auth and business logic stay in the
Node service, so **voice is a channel into the same agent**, not a parallel
system with its own brain.
```
browser ──audio──► node :4000 ──audio──► this :4100 ──► IndicConformer / Whisper
│ │
└────────── same graph, agents, ───────────┘
guardrails as text chat
│
browser ◄──audio───── node ◄──audio──── this ◄── Indic Parler-TTS
```
## Models
| Job | Model | Notes |
|---|---|---|
| Endpointing | Silero VAD | 512-sample frames @16 kHz, 300 ms pre-roll |
| Indic ASR | `ai4bharat/indic-conformer-600m-multilingual` | 22 Indian languages, CTC decoding |
| English ASR + language ID | `openai/whisper-small` | multilingual on purpose — the `.en` build cannot identify languages |
| TTS | `ai4bharat/indic-parler-tts` | 21 languages, streaming |
**The AI4Bharat repos are gated.** Access is auto-approved, but the download
needs an authenticated account: sign in to huggingface.co, accept the terms on
both model pages, then put a read token in `.env` as `HF_TOKEN`.
## Why two ASR models
IndicConformer decodes *as* the language you name — it does not detect one, and
English is not among its 22 codes. Whisper covers English and can identify the
spoken language in a single decoder step. So the default mode is `auto`:
```
audio → Whisper mel + 1 decoder step → language ID
├─ "en" → Whisper transcribes (mel already computed — no extra cost)
└─ Indic → IndicConformer with the detected code
```
Below **0.60** confidence the caller's preferred language wins instead of a coin
toss. That matters for Tanglish, where a short code-mixed sentence can honestly
land either side.
## Setup
```bash
python -m venv --system-site-packages .venv # reuses the system torch build
.venv/Scripts/python -m pip install -r requirements.txt
cp .env.example .env # add HF_TOKEN
```
The venv deliberately inherits system site-packages: torch is ~2.5 GB and
already installed with CUDA. Note that `parler-tts` pins `transformers==4.46.1`
**inside the venv only** — the system install is untouched.
## Run
```bash
npm run voice # from the parent directory
# or
.venv/Scripts/python -m app.server
```
First start downloads several GB and warms both models. `GET /health` reports
device, models, sample rate and current VRAM.
## Protocol
One WebSocket at `/ws/voice`, JSON control frames plus binary audio.
| Direction | Message |
|---|---|
| → | binary — 16 kHz mono PCM16 mic frames |
| → | `{"type":"config","lang":"auto","prefer":"ta"}` |
| → | `{"type":"speak","text":"…","id":"…"}` |
| → | `{"type":"cancel"}` — stop speaking now |
| ← | `{"type":"speech_start"}` — VAD opened a turn (drives barge-in) |
| ← | `{"type":"transcript","text":…,"lang":…,"detected":…,"confidence":…}` |
| ← | `{"type":"audio_start","sample_rate":24000}` then binary float32 chunks |
## Tuning
| Env | Default | Effect |
|---|---|---|
| `VAD_SILENCE_MS` | 700 | trailing silence that ends a turn — lower feels snappier, truncates people who pause |
| `VAD_MIN_SPEECH_MS` | 250 | ignores coughs and door slams |
| `VAD_PREFIX_MS` | 300 | audio kept from before detection, so word onsets survive |
| `VOICE_DEFAULT_LANG` | `ta` | tiebreak when language ID is unsure |
| `PRELOAD_ENGLISH` | `true` | set `false` to load Whisper lazily if VRAM is tight |
| `STT_DECODING` | `ctc` | `rnnt` is more accurate but decodes autoregressively |
## VRAM
Roughly 4.7 GB of the 6 GB card with all three models resident. If that proves
too tight, `PRELOAD_ENGLISH=false` defers Whisper (~0.5 GB) until the first
English utterance.
## Scripts
```bash
.venv/Scripts/python probe_access.py # which repos the token can reach
.venv/Scripts/python probe_models.py # load, VRAM, time-to-first-audio
```