Code-switching Hindi Β· English Β· Hinglish Β· Real-time AMFI Market Data Β· Cryptographic Hash Chain Audit
Key Highlights β’ Live UI Gallery β’ System Architecture β’ Quickstart β’ Verification & Tests β’ Compliance
Generic generative AI assistants hallucinate when answering financial questionsβfabricating returns, inventing expense ratios, and confusing fund houses. In regulated wealth management, financial hallucination is an existential compliance violation.
VoiceFinAI is a production-grade, voice-first financial advisor designed specifically for Indian Mutual Funds. It solves the hallucination crisis through a dual-engine architecture:
- Low-Latency Streaming Voice Interface: Sub-second conversational voice in Hinglish, Hindi, and English using Deepgram Nova-3 speech recognition and neural text-to-speech.
- The
AuditStreamCryptographic Verification Firewall: Generative models are constrained to emit parameterized templates ({F1.cagr3},{F1.nav}). Every placeholder, index, and numerical claim is mathematically checked against a verified AMFI fund snapshot before a single syllable is voiced or a single pixel is rendered.
If an LLM hallucinates an ungrounded return or attempts to recommend a fund outside the verified snapshot, AuditStream intercepts it in real-time, logs the violation into a SHA-256 hash-chain ledger, and falls back to deterministic AMFI templates.
User Voice βββΆ [Deepgram Nova-3] βββΆ Intent Router βββΆ AMFI Snapshot βββΆ Groq LPU (Qwen/Claude)
β
[Parametric Template]
β
βΌ
Audio Stream βββ [ElevenLabs/TTS] βββ [Audited Output] βββ [ AuditStream Firewall ]
β’ Zero Bare Numbers
β’ Strict AMC Match
β’ SHA-256 Hash Chain
- Rule 1 β Entity Existence: Ensures referenced fund slots (
{F1},{F2}) strictly exist in the current filtered universe snapshot. - Rule 2 β Figure Accuracy: Generative models never generate bare numbers. All NAVs, 1y/3y/5y CAGRs, minimum SIPs, and expense ratios are injected directly from the deterministic verified snapshot.
- Rule 3 β Cryptographic Audit Ledger: Every turn produces a tamper-proof SHA-256 hash chained to the preceding turn (
prev_entry_hashβentry_hash), creating an immutable record of all spoken and displayed financial advice.
- Deepgram Nova-3: Optimized speech-to-text with conversational Voice Activity Detection (VAD: 800ms silence threshold) tuned for Indian accents and English-Hindi code-switching.
- Groq LPU Acceleration: Generates parameterized natural-language summaries in ~550ms using
qwen/qwen3.8-27band Claude Sonnet fallbacks. - Adaptive Speech Streaming: Word-by-word typewriter interface synchronized with natural audio playback.
- Normalizes colloquial Indian financial terminology (
"2 lakh SIP","paanch saal horizon","surakshit fund","inme se kaunsa lu","safe fund dikhao"). - Automatic language detection dynamically preserves user dialect preference across turns.
- Curated universe of 100+ Direct Growth Mutual Funds.
- Real-time NAV & historical returns synchronization from AMFI via
mfapi.in. - Strict AMC Isolation: When a user explicitly mentions a fund house (e.g. Mirae Asset, Parag Parikh, Quant, HDFC), VoiceFinAI enforces hard filteringβnever substituting competitor funds.
sequenceDiagram
autonumber
actor User as π€ Investor
participant UI as π₯οΈ Single-Page UI
participant Flask as π Flask Server
participant DG as ποΈ Deepgram Nova-3
participant Router as π§ Intent Router
participant AMFI as π AMFI / MFAPI Snapshot
participant LLM as π§ Groq LPU (Qwen/Claude)
participant Firewall as π‘οΈ AuditStream Firewall
participant TTS as π ElevenLabs / Edge TTS
User->>UI: Speaks voice query ("Moderate risk fund dikhao")
UI->>Flask: POST /transcribe (Audio blob)
Flask->>DG: Stream audio payload
DG-->>Flask: Transcript ("Moderate risk fund dikhao")
Flask->>Router: Parse turn & extract intent
Router-->>Flask: Intent: DISCOVER (Risk: Moderate)
Flask->>AMFI: Build verified Snapshot (Direct Funds, live NAVs)
AMFI-->>Flask: Snapshot (F1, F2, F3 with verified CAGRs)
Flask->>LLM: Prompt with parameterized slot schema
LLM-->>Flask: Template: "{F1} has {F1.risk} risk and {F1.cagr3} return..."
Flask->>Firewall: Validate template against Snapshot
alt Template Valid
Firewall-->>Flask: Approved & Grounded Text
else Violation Detected
Firewall-->>Flask: Fallback to Audited Bank Template
end
Flask->>TTS: Synthesize approved text
TTS-->>Flask: Audio Token
Flask-->>UI: Cards JSON + Voice Token + Live Text
UI-->>User: Synchronized Audio Playback + Interactive Cards
| Pipeline Stage | Provider / Component | Mean Latency | Reliability |
|---|---|---|---|
| Speech-to-Text (ASR) | Deepgram Nova-3 | ~240ms | 99.9% uptime |
| Intent Normalization | Rule & Aho-Corasick Automaton | <2ms | Deterministic |
| Snapshot Retrieval | Local Disk / AMFI API Cache | <15ms | Zero network lag |
| Template Generation | Groq LPU (qwen/qwen3.8-27b) |
~550ms | Audited fallback |
| Firewall Validation | core.auditstream.AuditStream |
<4ms | 100% mathematical |
| Neural Speech (TTS) | ElevenLabs / Edge TTS | ~280ms (TTFB) | Dual provider |
| End-to-End Turn | Complete Roundtrip | <1.2s | Real-time conversation |
- Python 3.10, 3.11, or 3.12
- Git
git clone https://github.com/Aan9758/VoiceFinAI.git
cd VoiceFinAI
# Create virtual environment
python -m venv .venv
# Activate virtual environment
# On Windows PowerShell:
.venv\Scripts\Activate.ps1
# On Linux / macOS:
source .venv/bin/activatepip install -r requirements.txtVoiceFinAI comes with an offline mock engine and fallback template bank, allowing full local evaluation without external API keys!
To enable cloud voice recognition and LLM inference:
cp .env.example .envEdit .env with your keys:
DEEPGRAM_API_KEY=your_deepgram_key
GROQ_API_KEY=your_groq_key
GROQ_MODEL=qwen/qwen3.8-27b
ELEVENLABS_API_KEY=your_elevenlabs_key
ELEVENLABS_VOICE_ID=your_voice_idpython app.pyOpen http://127.0.0.1:5000 in your browser.
π‘ Offline UI Test Harness: Run
python tests/offline_e2e_server.pyto inspect the complete voice recording and card rendering loop with zero external API calls.
VoiceFinAI is engineered with rigorous test-driven safety. The test suite includes 84 comprehensive test cases covering AMC hard-filtering, snapshot isolation, session memory, intent routing, and firewall interception.
# Run the entire test suite
python -m pytestOutput:
============================= test session starts ==============================
rootdir: C:\Users\amans\OneDrive\Desktop\VOICEFINAI
configfile: pytest.ini
testpaths: tests, test_intent_router.py
collected 84 items
tests/test_amc_filter.py ................ [ 19%]
tests/test_app_routes.py ............ [ 33%]
tests/test_tts_provider.py .... [ 38%]
test_intent_router.py .................................................... [100%]
============================== 84 passed in 0.43s ==============================
Run the standalone audit-chain verifier to watch the AuditStream firewall detect poisoned LLM templates in real-time:
python verify_audit_chain.pyVoiceFinAI is designed in alignment with SEBI (Securities and Exchange Board of India) guidelines for digital financial tools:
- Direct Plans Exclusively: All suggested mutual funds are Direct Plans (
Direct Plan - Growth), minimizing expense ratios and eliminating distributor commission biases. - Categorization Heuristics: Adheres strictly to SEBI Categorization circulars (Large Cap, Mid Cap, Small Cap, Flexi Cap, Liquid, Banking & PSU Debt).
- No Direct Execution: VoiceFinAI is an informational analysis engine; it does not take custody of funds, access demat accounts, or execute trades.
- Mandatory Disclaimers: Discloses risk levels and historical return caveats on every screen and spoken turn: "Mutual fund investments are subject to market risks. Past performance is not an indicator of future returns."
VoiceFinAI/
βββ core/
β βββ amfi.py # Live AMFI / MFAPI sync & NAV calculations
β βββ auditstream.py # Mathematical zero-hallucination verification firewall
β βββ dialogue.py # Multi-turn conversational session manager
β βββ generate.py # Prompt templating & Groq/Claude integration
β βββ pipeline.py # End-to-end turn orchestrator
β βββ retrieval2.py # Portfolio snapshot builder with epoch rotation
β βββ snapshot.py # Data structures for immutable financial snapshots
β βββ tts.py # ElevenLabs & Edge TTS neural voice manager
β βββ universe.py # 100+ SEBI curated fund catalog & search
βββ data/
β βββ universe.json # Curated mutual fund master catalog
βββ docs/
β βββ assets/ # High-resolution screenshots and vector banners
βββ scripts/
β βββ capture_screenshots.py # Automated Playwright documentation screenshots
βββ static/ # Audio previews, avatar assets & greetings
βββ templates/
β βββ index.html # Single-conversation glassmorphic web UI
βββ tests/
β βββ offline_e2e_server.py # Offline provider-less test harness
β βββ test_amc_filter.py # AMC isolation & safety unit tests
β βββ test_app_routes.py # Flask HTTP & audio route tests
β βββ test_tts_provider.py # Voice provider fallback tests
βββ .env.example # Environment configuration template
βββ .gitignore # Clean zero-leak gitignore
βββ app.py # Main Flask entrypoint
βββ intent_router.py # Bilingual Hinglish/English intent classifier
βββ LICENSE # MIT License
βββ pytest.ini # Pytest configuration
βββ README.md # Executive documentation
βββ requirements.txt # Production dependencies
βββ verify_audit_chain.py # Cryptographic ledger verification script
Aman Saraswat
Indian Institute of Technology Guwahati (IIT Guwahati)
- GitHub: @Aan9758
- Project: VoiceFinAI
Built with Deepgram Nova-3, Groq LPU, ElevenLabs, and Flask.



