Files
Agentic-AI/README.md
T

182 lines
9.0 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# WeLe Agentic AI
A LangGraph multi-agent service that sits beside the WeLe CRM and answers questions about live CRM data — with charts, tables and generated documents rather than paragraphs of numbers.
Runs as its own service (default `:4000`), not inside the CRM. The CRM stays a CRM; the assistant can be restarted, scaled, or given a second channel without touching it.
---
## Architecture
```
CRM Chat UI ─┐
(WhatsApp) ─┤ 1. CHANNEL GATEWAY auth (CRM JWT) · session · rate limit · normalise
(Website) ─┤ ↓
(Mobile) ─┘ 2. ORCHESTRATION intake → supervise (supervisor chain) → finalize
↓
3. GUARDRAILS RBAC · policy · approval · content safety · audit
↓ (enforced at the tool boundary, not the graph)
4. AGENT LAYER 8 read-only domain agents (agent chain), in parallel
↓
5. TOOLS reads → MongoDB | writes → CRM REST API
↓
6. DATA MongoDB (live CRM) · Redis (memory + cache)
↓
7. OUTPUT typed blocks: metrics · chart · table · file · text
```
### Key decisions
**Reads hit Mongo, writes go through the CRM API.** Direct Mongo gives rich aggregation for reporting; routing writes through the CRM's own routes keeps lead scoring, stage transitions, Socket.IO broadcasts and its audit trail intact. Writing to Mongo directly would produce records the CRM never learns about.
**Guardrails live in `defineTool`, not in graph nodes.** A node-level check is bypassed the moment someone adds an edge. A wrapper around the tool itself cannot be: if a model can call it, it went through RBAC → policy → approval → audit.
**Sub-agents are read-only; the supervisor performs every write.** The approval gate uses LangGraph's `interrupt()`, which suspends the graph owning the checkpointer. A sub-agent invoked as a tool is a *nested* graph with no checkpointer — an interrupt raised there bubbles out as an error, gets caught as a failed delegation, and the operator is never actually asked while the assistant claims something "needs approval". Hoisting writes to the supervisor puts every interrupt in the checkpointed graph.
**Claude only, with a model failover chain.** Each tier resolves an ordered list of Claude model ids tried left to right, so an overloaded or erroring model hands the turn to the next instead of failing the request:
```
LLM_CHAIN_SUPERVISOR=claude-opus-5,claude-sonnet-5
LLM_CHAIN_AGENT=claude-sonnet-5,claude-haiku-4-5
```
Claude 5 API rules are encoded once in `orchestration/providers.js`: `temperature`/`top_p`/`budget_tokens` are removed (a 400 if sent), thinking is adaptive and must not be disabled, and depth is tuned with `output_config.effort`.
**Answers are typed blocks, not strings.** The frontend renders metrics rows, charts, tables and file cards natively. Same payload can drive a different channel later.
---
## Setup
```bash
npm install
cp .env.example .env # fill in ANTHROPIC_API_KEY, MONGODB_URI, CRM_JWT_SECRET
npm start # or: npm run dev
```
`CRM_JWT_SECRET` **must** match the CRM's `JWT_SECRET` — that is how a CRM login is accepted here, with the same role and permissions.
Redis is optional: if unreachable the service falls back to an in-process store and logs a warning, so local dev works without it.
### Frontend
The chat page is `frontend/src/pages/AIAssistant.jsx` in the CRM repo, routed at **`/assistant`**.
```bash
# in the CRM frontend
echo "VITE_AGENT_URL=http://localhost:4000" > .env.local
npm run dev
```
---
## API
| Method | Path | Purpose |
|---|---|---|
| `POST` | `/api/agent/chat` | One turn, buffered JSON |
| `POST` | `/api/agent/stream` | One turn, SSE progress + final blocks |
| `POST` | `/api/agent/approve` | Resume a run paused for approval |
| `GET` | `/api/agent/sessions` | Caller's recent conversations |
| `GET` | `/api/agent/sessions/:id` | One conversation's history |
| `DELETE` | `/api/agent/sessions/:id` | Delete a conversation |
| `GET` | `/api/agent/capabilities` | Agents/tools this caller may use |
| `GET` | `/api/agent/audit/:sessionId` | Tool-call audit trail (admin) |
| `GET` | `/artifacts/:id` | Download a generated file |
| `GET` | `/health` | Mongo / Redis / CRM reachability |
All routes take `Authorization: Bearer <CRM token>`.
---
## Agents
| Agent | Owns |
|---|---|
| Lead | Search, profiles, funnel breakdowns, trends, stale-lead hygiene |
| Analytics | Business snapshot, conversion funnel, team performance, source ROI, enrolments |
| Conversation | WhatsApp/FB/IG inbox — history, triage, escalations, volumes |
| Scheduling | Follow-ups and demos, overdue work, the calendar |
| Campaign | Broadcast campaigns and delivery performance |
| Calls | Telephony volumes, answer rates, call records |
| Course | Course/batch catalogue, workshop funnel |
| People | Users, roles, trainers, DSR submissions |
An agent whose whole toolset is behind permissions the caller lacks is hidden from routing entirely, rather than being offered and then failing.
---
## Output
`create_metrics`, `create_chart`, `create_table`, `generate_document` (xlsx / docx / pptx / pdf / csv).
Charts are rendered server-side to SVG (theme-aware in the browser) and rasterised to PNG via `@resvg/resvg-js` for embedding in documents — all pure-JS or prebuilt-binary, no native toolchain needed on Windows.
The palette is validated for colour-vision deficiency in both light and dark modes. Three light-mode slots fall below 3:1 contrast, so every chart ships **direct labels and a data table** — identity is never carried by colour alone.
---
## Data notes discovered against the live database
These are encoded in the tools and the supervisor prompt, and matter for anyone extending this:
- **`Lead.enrolled` is `false` on every one of the ~4,500 lead records** — nothing in the CRM sets it. Enrolment and revenue truth lives in `enrollmentlogs` (`paymentStatus = SUCCESS`). Reporting off the lead flag reports a permanent zero.
- `enrollmentlogs.primaryMobile` stores bare 10-digit mobiles while `leads.phone_number` is 91-prefixed. A naive join matches nothing; `paidPhoneVariants()` in `tools/crm/_shared.js` handles it.
- Phone number is the join key across every collection.
- `current_stage: converted` is used on only a handful of leads — it is not a reliable conversion signal.
- Call logs stop in April 2026; an empty window means an inactive feed, not zero activity.
---
## Voice
Speech-to-speech lives in `voice-service/` (Python, GPU) and enters through
`src/gateway/voice.js` as a WebSocket channel — so a spoken question runs the
same graph, agents and guardrails as a typed one.
```bash
npm run voice # start the GPU service on :4100 first
npm start # the agent service exposes ws://localhost:4000/api/agent/voice
```
Silero VAD does the endpointing, `ai4bharat/indic-conformer-600m-multilingual`
transcribes 22 Indian languages, `openai/whisper-small` handles English *and*
identifies the language, and `ai4bharat/indic-parler-tts` speaks the reply.
Language defaults to auto-detect per utterance, so Tamil and English can be
mixed across turns without touching a setting.
The interesting problem was not audio: a turn takes 17–46 s, and that is dead
silence in voice. So the bridge speaks an acknowledgement within ~1 s, narrates
each agent delegation aloud, and reads the answer sentence-by-sentence as it is
composed. Barge-in is server-side — the VAD opening a turn cancels TTS mid-word.
Details, protocol and tuning: `voice-service/README.md`.
## Claude notes
- Models: `claude-opus-5` (supervisor), `claude-sonnet-5` (agents), `claude-haiku-4-5` (fallback). Use the exact ids — never append date suffixes.
- Roughly $0.10–0.30 per question (supervisor + sub-agents + tool loops ≈ 20–60k tokens).
- Check the account can serve a request with `node scripts/t-providers.mjs`.
## Tests
Diagnostics against live data and the real API (no mocks):
```bash
node scripts/t-providers.mjs # can Claude serve a request right now?
node scripts/t-llm.mjs # Claude 5 models + tool calling
node scripts/t-tools.mjs # all read tools against live Mongo
node scripts/t-chart.mjs # renders every chart form to SVG + PNG
node scripts/t-docs.mjs # builds xlsx/docx/pptx/pdf/csv
node scripts/t-docverify.mjs # structural verification of those files
node scripts/t-e2e.mjs # full agent turns over HTTP
node scripts/t-stream.mjs # SSE streaming
node scripts/t-approve.mjs # approval interrupt + resume (--approve to accept)
```
---
## Adding a channel
Add an entry to `src/gateway/channels.js` declaring how the principal is established, whether it may mutate CRM state, and which block types it can render — then an adapter producing a normalised request. The orchestrator does not change. Public channels should be `readOnly: true`; the policy engine then refuses every non-read tool on them.