| Version | 1.2 |
| Date | 10 Jun 2026 |
| Status | Draft — pending validation at PS Explainer recording + Zoho Catalyst workshop (11 Jun) |
| Prototype deadline | 26 Jul 2026 (submit by 24 Jul — 2-day buffer) |
| Grand Finale | 26 Sep 2026, in-person demo day |
KAVAL (ಕಾವಲು) — Kannada for "the watch / the guard." A name the judging panel reads in their own language. Backronym if anyone asks: Knowledge-Augmented Vigilance & Analytics Layer.
"Every police officer in Karnataka gets a crime analyst in their pocket — one that speaks Kannada, shows its work, and never invents a number."
KAVAL turns the SCRB's records from 1,100+ police stations into a conversational intelligence layer: an investigator asks a question by voice or text, in Kannada or English, and gets back a verified answer rendered as live interactive maps, criminal-network graphs, and trend charts — every answer carrying court-grade provenance.
SCRB manages a large, growing repository of crime data from 1,100+ stations. Current systems = static dashboards + manual queries → no deep analysis, no real-time insight.
Build: a conversational AI platform for investigators to query crime data in natural language and uncover patterns, relationships, and predictive insights.
Brief's required capability list (all are mapped to features in §7):
- Natural-language chatbot (English + Kannada)
- Voice-enabled interaction
- Context-aware conversations
- PDF export of conversation history
- Criminal network visualization
- Crime trend & hotspot detection
- Predictive analytics & early warnings
- Explainable AI with audit trails
- Role-based secure access
Insight domains: crime pattern discovery · criminal network analysis · socio-demographic insights · behavioral profiling · proactive crime-prevention intelligence.
- ≥ 90% accuracy on a 120-question bilingual golden evaluation set (see §10).
- Zero hallucinated numbers — enforced by architecture, not by prompting (§8.4 Numeric Firewall).
- End-to-end answer (question → rendered widget) in < 6 s p95.
- All 9 brief capabilities demonstrable; 6 of them production-quality (P0 set).
- A 3-minute demo that produces an audible reaction from a non-technical judge.
- ❌ Real-time CCTV/face recognition, image/video analytics
- ❌ Mobile native apps (responsive PWA is enough)
- ❌ Training/fine-tuning custom LLMs
- ❌ Live integration with actual CCTNS/KSP production systems
- ❌ Full Kannada UI localization (chat I/O bilingual; chrome stays English)
- ❌ Multi-tenancy beyond the role model in §7.9
| Persona | Role | Core job-to-be-done | Key features |
|---|---|---|---|
| IO Investigator (SHO / PSI) | Field investigation | "Who else is connected to this accused? What's the pattern in my beat?" | Chat, network graph, case lookup, PDF export |
| Crime Analyst (SCRB / DCRB) | Pattern analysis | "Where are emerging hotspots? Which MO clusters are active?" | Hotspots, trends, forecasts, drilldowns |
| Command (SP / DCP / DGP office) | Resource allocation | "Daily situational picture; where do I deploy?" | Morning Brief, early warnings, district rollups |
| System Admin / Auditor | Compliance | "Who accessed what, when, and why?" | Audit trail, role management |
Demo personas: we will demo as IO Investigator (story 1) and Command (story 2). Two personas max — more dilutes the narrative.
Everything in this PRD serves at least one of these.
| # | Differentiator | Why it matters |
|---|---|---|
| D1 | Numeric Firewall — the LLM architecturally cannot state a number that didn't come from an executed query (§8.4) | Kills the #1 failure mode of conversational analytics. One sentence to stakeholders: "Our AI cannot invent a statistic — by construction, not by promise." |
| D2 | Generative Analytics Canvas — answers materialize as live, interactive WebGL maps/graphs/charts inline in the chat, not text walls (§7.4) | Insight you can touch: drill, pan, expand — the difference between a report and an instrument. |
| D3 | Court-ready provenance ("Evidence Locker") — every answer carries the query, source tables, row counts, timestamp, officer ID, and a tamper-evident hash; PDF exports are case-file-grade (§7.8, §7.10) | Reframes a chatbot as an evidence instrument. Police workflows think in case files. |
| D4 | Kannada-first voice via Indic-specialized STT/TTS (Sarvam/Bhashini), incl. code-switched "Kanglish" (§7.2) | Accessibility for every rank, in the language officers actually speak — built on Indian sovereign AI. |
| D5 | KAVAL Morning Brief — the system speaks first: an auto-generated daily situational brief per station/district with anomalies and early warnings (§7.7) | Converts "reactive Q&A bot" into "proactive intelligence officer." |
| D6 | IPC ↔ BNS bridge — seamless querying across the 2024 IPC→Bharatiya Nyaya Sanhita transition (§9.4) | Deep domain awareness: every officer lives this transition pain daily. |
| D7 | BSA §63 court-admissibility certificate — every PDF export ships with an auto-generated certificate under §63(4), Bharatiya Sakshya Adhiniyam 2023 (successor to IEA §65B) (§7.10) | Turns "court-ready" from metaphor into legal workflow. Instantly legible to police leadership; invisible to teams without domain knowledge. |
| D8 | Sovereign Mode — one visible toggle swaps the cloud LLM for an on-prem model; fully functional air-gapped (§7.11) | Answers the unasked procurement question: crime data never leaves the building. Demo beat: flip the toggle, pull the Wi-Fi, keep answering. |
| D9 | Measured, not claimed — published failure-mode/refusal page + live accuracy scoreboard + committed eval history showing the accuracy growth curve (§7.8, §10) | Every team claims quality; a timestamped improvement curve cannot be faked retroactively. |
| D10 | Cold-start intelligence — role-aware suggested questions generated from the data on login (§7.12) | Kills the blank-textbox problem that sinks real chatbot deployments; reframes KAVAL from search box to colleague. |
- PSI opens KAVAL, speaks in Kannada: "ಕಳೆದ ಮೂರು ತಿಂಗಳಲ್ಲಿ ನನ್ನ ಠಾಣಾ ವ್ಯಾಪ್ತಿಯಲ್ಲಿ ಚೈನ್ ಸ್ನ್ಯಾಚಿಂಗ್ ಪ್ರಕರಣಗಳು ಎಲ್ಲಿ ಹೆಚ್ಚಾಗಿವೆ?" (Where have chain-snatching cases risen in my station limits in the last 3 months?)
- KAVAL answers in spoken Kannada + renders a hotspot map widget scoped to her jurisdiction (RBAC-enforced).
- Follow-up in English (context carries): "Show repeat offenders active in these zones." → Network graph widget; nodes sized by case count.
- She clicks a node → offender panel: linked FIRs, co-accused, MO history, IPC/BNS sections.
- "Export this for the case file." → court-ready PDF with provenance hashes lands in 5 seconds.
- SP opens the Morning Brief: 3 auto-generated cards — a vehicle-theft anomaly in District X (+38% week-over-week), a forecast spike flagged for the upcoming festival window, a newly active offender network.
- Clicks the anomaly → drilldown chat pre-seeded with context; asks "compare with the same festival period last year" → comparative trend chart.
- Every screen shows the ⓘ provenance chip: which query, which tables, how many rows, when.
Priorities: P0 = must ship for prototype · P1 = ship if on schedule · P2 = finale-phase / stretch.
- Accepts English, Kannada, and code-switched input; text or voice.
- Pipeline (full spec §8): intent routing → tool selection → guarded query execution → verified composition.
- Behavioral contract:
- Never answers data questions from model memory — only from executed queries.
- Ambiguous question → asks ONE crisp clarifying question (e.g., "Calendar year or last 12 months?").
- No data → says so plainly. Refusal is rendered as a first-class, well-designed state, not an error.
- Every numeric/factual claim traceable to a query (D1).
- Streaming responses; typed answer skeleton appears < 1.5 s.
- STT: Web Speech API (
kn-IN,en-IN) as the zero-cost baseline; Sarvam Saaras (or Bhashini) API as the quality tier — Indic-specialized, handles Kanglish code-switching. Runtime-selectable. - TTS: Sarvam Bulbul (Kannada voice) for spoken answers; pre-cache demo-script audio for offline resilience (§11.4).
- Language of the answer follows language of the question; user can pin a language.
- Mic button with live waveform; transcript editable before send (officers will want to correct STT).
- Romanized Kanglish 〔P1〕: Latin-script Kannada ("Indiranagar alli ninne yeshtu theft aagide?") normalized before planning — this is how officers actually type; almost no team will handle it.
- Session memory: entity tracking (accused, stations, districts, crime types, time windows) so follow-ups resolve "he / that area / same period" correctly.
- Active-filter chips above the chat (e.g.,
District: Tumakuru ✕Period: Jan–Mar 2026 ✕) — visible, removable state. Evaluators see the context awareness instead of taking our word. - "New thread" resets context explicitly.
The engine returns a Widget DSL (JSON spec, appendix B) alongside narrative; the frontend renders interactive widgets inline:
map.hotspot— deck.gl heat/hex layer on MapLibre vector tiles of Karnataka, station/district boundaries, click-to-drilldown.graph.network— WebGL force-directed criminal network (react-force-graph); node = person/case, edge = co-accused/shared-FIR/shared-MO; cluster coloring, expand-on-click.chart.trend— ECharts time series with anomaly bands, period comparison.table.records— paginated, sortable case/record table with PII masking per role.card.kpi,card.alert— stat and early-warning cards.- All widgets carry the provenance chip ⓘ (§7.8) and an "open full-screen" mode for the demo projector.
- Offline job builds an edge table: co-accused in same FIR, shared addresses/phones (if dataset provides), shared MO cluster, repeat-victim links.
- networkx computes degree/betweenness centrality + Louvain communities → stored as node attributes.
- Investigator flows: "show network of accused #X (2 hops)", "most central figures in vehicle-theft networks in District Y", repeat-offender tracking (≥ N FIRs in window).
- Spatial: KDE / hex-bin aggregation over geocoded FIRs — PostGIS on the dev/Sovereign path, app-layer haversine + grid binning on Catalyst (Data Store has no spatial extension; at FIR scale this costs nothing); granularity district → station → grid.
- Temporal: weekly/monthly series per (geo × crime-type); STL decomposition for seasonality; z-score anomaly flags.
- Socio-demographic cuts (age bands, victim/accused profiles) where the dataset supports it — validate at data drop (§13, A2).
- Maps drawn against police-station jurisdiction boundaries where obtainable (officers think in station limits, not census geography); graceful fallback to district boundaries.
- Forecasting: gradient-boosted / Prophet per (district × major crime type), honest 4-week horizon with confidence bands. Transparent features only (recency, frequency, seasonality, festival calendar) — no black-box claims.
- Early-warning rules: forecast breach, week-over-week anomaly, newly-active network detection.
- Morning Brief: daily auto-generated per role scope — 3–5 cards, each one click from a drilldown chat. LLM writes the prose; numbers injected via Numeric Firewall.
- Ethics guardrail: we score places and patterns, never communities or individuals' future criminality. One slide in the deck on this — responsible-AI questions deserve a prepared answer.
- Per answer: interpreted intent (plain English/Kannada), executed SQL (collapsible), source tables, row count, execution time, model+version, confidence tier.
- Audit log: append-only
audit_events— who, when, question, query, tables touched, rows returned, response hash; each row carriessha256(prev_hash ‖ row)→ tamper-evident chain. Auditor screen renders the chain with verification status. - Live accuracy scoreboard in the footer: our golden-set accuracy, updated by CI. Radical transparency as a feature.
- Published failure modes (D9): an in-app page listing the question classes KAVAL refuses and why (out-of-jurisdiction, beyond forecast horizon, insufficient data) — we present our own attack surface before anyone probes it.
- Accuracy growth curve (D9): every eval run is timestamped and kept; the improvement curve (e.g., 71% → 93% across the build weeks) renders in-app and in the final deck.
- Roles:
CONSTABLE(own-station, PII-masked) ·INVESTIGATOR(own-district, full records) ·COMMAND(state-wide aggregates + district drill) ·ADMIN/AUDITOR(audit + user mgmt, no case data). - Enforced by a jurisdiction clamp in the query pipeline (shipped W1.5): every spec is rewritten to the caller's scope before compilation, and out-of-scope asks become audited refusals — never enforced in the prompt. Postgres RLS adds a second, DB-layer net on the Sovereign path; Catalyst API Gateway fronts the hosted deploy. One demo beat: same question, two logins, two different (correctly scoped) answers.
- JWT auth (httpOnly), bcrypt, session timeout. PII masking (names/phones partially redacted) below INVESTIGATOR.
- Server-side (WeasyPrint): KSP-styled header, conversation Q&A, embedded widget snapshots (server-rendered PNG), and an annexure of provenance — every query, table list, row counts, timestamps, officer identity, audit-chain hash in the footer.
- BSA §63 certificate (D7): auto-generated certificate for electronic records under §63(4), Bharatiya Sakshya Adhiniyam 2023 (successor to the IEA §65B certificate every IO knows) — system identification, output hash, operator declaration — appended to every export. ~1 day of work; legal review of wording with a police mentor.
- Framing in demo: "This isn't a chat log — it's an annexure ready for the case file, certificate included."
- A visible toggle in the header: Cloud (Claude via gateway) ↔ On-prem (local model via Ollama). Same pipeline, same Numeric Firewall — only the planner endpoint changes (the LiteLLM abstraction makes this nearly free).
- Acceptance bar: the scripted demo questions answer correctly with Wi-Fi off, end-to-end on one laptop.
- Honest framing: local-model accuracy will be lower — show its own scoreboard entry rather than hiding it. The point is the deployment option, not parity.
- Demo beat (finale): flip the toggle, disconnect the network, answer a question. One sentence: "Crime data never has to leave the building."
- On login, KAVAL speaks first: 3 role-scoped suggested questions generated from the data, not hardcoded — e.g., "Vehicle theft in your limits is up 22% this month — want the station breakdown?"
- Implementation: nightly job ranks anomalies/deltas per scope (reuses Morning Brief plumbing, §7.7); each suggestion is one tap from a full answer.
- Solves the blank-textbox problem that kills chatbot adoption; doubles as silent onboarding.
Each module below is grounded in replicated criminology research — we cite the evidence base on its deck slide and in-app ⓘ. Rule: ship nothing we can't defend to a criminologist.
| Module | What it does | Evidence base | Needs from dataset | Priority |
|---|---|---|---|---|
| Crime Concentration (Pareto) view | "4% of places produce ~50% of crime in this district" — ranked micro-places with cumulative curve | Law of crime concentration (Weisburd 2015, Criminology): crime concentrates at a small share of micro-places, remarkably stable across cities | Geocoded FIRs | P0 — one query, huge resource-allocation story |
| Near-Repeat Radar | After a qualifying burglary/vehicle theft, auto-draw a 250–400 m, 7–14 day elevated-risk envelope; alert the beat | Near-repeat victimization (Townsley et al. 2003; Bowers & Johnson 2004) — robust, widely replicated spatio-temporal pattern. Pattern is solid; prevention payoff depends on patrol follow-through, and we say so | FIR lat/lon + occurrence date | P1 |
| Aoristic Clock | True time-of-day risk for discovery crimes (burglary, vehicle theft) by spreading probability across the occurrence from–to interval | Aoristic analysis (Ratcliffe): occurrence times for unwitnessed crimes are intervals, not points — naive "report hour" charts are systematically wrong. Indian FIRs record occurrence from → to, a perfect fit | Occurrence from/to timestamps (verify at data drop) | P1 |
| Repeat-victimization flags | Addresses/persons victimized ≥2× flagged for prevention outreach lists | Farrell & Pease (1993) and successors: a small share of victims accounts for an outsized share of incidents; repeat-victim prevention is among the best-evidenced tactics | Victim identifiers / addresses | P1 |
| Koper patrol planner | Convert hotspots into a suggested schedule of 10–15 min patrol dwells | Koper curve (1995, replicated): 10–16 min hotspot dwells maximize residual deterrence; hotspot-policing meta-analyses (Braga et al.) show consistent reductions without simple displacement | Hotspot output (already have) | P2 |
| Solvability triage | Score fresh FIRs on solvability factors to prioritize investigative effort | UK evidence-based investigative triage (Kent Police EBIT trials); solvability-factor literature (Wellford & Cronin) | Case outcomes (FIR status — likely have) | P2 |
| RTM-lite environmental risk | Risk layers from environment (bars, ATMs, bus stops, transit) using OSM POIs — forecasts from places, not past arrests | Risk Terrain Modeling (Caplan & Kennedy, Rutgers) — also our strongest bias answer: we score environments, never communities | Public OSM data only | P2 stretch |
PS1 §4 asks for it explicitly; ours is wired, awaiting fields:
- Demographic cuts as first-class spec dimensions: age band, gender, victim/accused role —
dim_person(§9.2) already models them; they become group-bys and filters the moment data-drop profiling confirms the fields exist. - District-level correlation overlays: crime rates joined against public socio-economic indicators (Census urbanization, literacy, sex ratio) — scatter/choropleth widgets with the correlation≠causation caveat printed on the widget itself, not in fine print.
- Evidence framing: the social-disorganization tradition (Shaw & McKay) explains why place-level socio-economic context belongs next to crime data. We present patterns over places and aggregates, never community blame — §7.7's ethics guardrail applies verbatim.
- Activates if the dataset carries financial fields (cyber-fraud FIRs commonly include account / UPI / phone identifiers). Go/no-go at data-drop profiling, documented either way.
- If present: accounts, UPI handles, and phone numbers become typed nodes in the §7.5 network graph; a "money trail" is a path query over those edges; fan-in/fan-out and same-account-across-FIRs rules flag suspicious clusters.
- If absent: the identical graph machinery demos on co-accused / shared-MO edges and the adapter ships documented — the answer to "where's financial crime?" is "wired, awaiting fields", never silence.
Saved-question shortcuts per role · scheduled report subscriptions · WhatsApp-style brief delivery · Catalyst-native notifications · cross-district MO pattern matching via embedding case linkage (Woodhams & Bennell case-linkage literature; pgvector on Sovereign, QuickML RAG on Catalyst) · what-if deployment simulator.
The single highest-risk, highest-value component. Design principle: constrain the LLM's blast radius; let deterministic code do everything that can be deterministic.
voice ──► STT ─┐
├─► Language ID ─► Context Resolver ─► Planner (LLM, tool-calling)
text ──────────┘ │ (entity/coref │
│ memory) ▼
│ ┌─ Tool Registry ─────────────┐
│ │ query_crimes (StructFilter) │ ← 80% path
│ │ hotspot / trend / forecast │
│ │ network_expand / centrality │
│ │ lookup_entity / list_values │
│ │ guarded_sql (long tail) │ ← 20% path
│ └──────────────┬──────────────┘
│ ▼
│ Executor (read-only role, scope
│ clamp, 5s timeout, row caps)
│ ▼
│ Verifier (sanity checks,
│ self-correct ≤ 2 retries)
▼ ▼
Composer (LLM narrative w/ slots) ◄── result objects
│
Numeric Firewall (slot injection)
│
┌─────────────┴──────────────┐
▼ ▼
narrative (+ TTS) Widget DSL JSON ─► frontend canvas
│
Provenance attach ─► audit chain
For the dominant query class (filter/aggregate/rank over crimes), the LLM emits a Pydantic-validated JSON filter spec, never SQL:
{
"intent": "aggregate",
"dataset": "fir_cases",
"filters": {"crime_group": "CHAIN_SNATCHING", "district": "Bengaluru City",
"date_range": {"last_n_months": 3}},
"group_by": ["station", "month"],
"metric": "count",
"viz_hint": "map.hotspot"
}Our compiler turns this into SQL against curated semantic views (§9.3). Validation failures → re-prompt with the error. This makes the common case near-deterministic and trivially testable.
mo_keyword shipped: modus-operandi text matching (PS1 names MO twice) is a live Filters field, not a future one — "how many pickpocket thefts" or "night break-ins" filters fct_fir.mo_text directly, case-insensitively, with the user's own %/_ escaped so it can't be read as a wildcard. At data drop the spec still grows further dataset-discovered dimensions (the demographic cuts of §7.14) as new fields and enum values; the architecture doesn't change, the vocabulary does.
For long-tail analytical questions:
- Prompt = semantic-view schemas + column value samples + top-5 similar Q→SQL exemplars retrieved by embedding similarity (pgvector on Sovereign; in-process index on Catalyst — few-shot RAG over our own golden set, accuracy compounds as we add examples).
- Generated SQL must pass:
SELECT-only AST check (sqlglot parse) → table/column allowlist →EXPLAINdry-run → cost cap. - Executes under the same read-only role + jurisdiction clamp + timeout. Error or empty-suspect result → self-correction loop (max 2), then graceful refusal with a suggested rephrasing.
The Composer LLM never sees raw freedom to write numbers. It produces narrative templates with typed slots:
"Chain-snatching cases in {{district}} rose {{pct_change}} over the last {{n_months}} months, with {{top_station}} reporting the most incidents ({{top_count}})."
Slot values are injected by code from the executed result set. A post-render regex/NER pass asserts no unslotted numerals survive in data-bearing sentences; violations fail loudly in dev and fall back to a tabular answer in prod. Hallucinated statistics are impossible by construction.
Every planner/agent call goes through one LiteLLM call site (engine/llm.py), parameterized by KAVAL_LLM_MODEL + KAVAL_LLM_API_BASE + KAVAL_LLM_API_KEY. There is no separate "dev" code path — QuickML, Anthropic, and any other OpenAI-compatible endpoint are the same three env vars pointed at different values:
| Layer | Endpoint (same code path) | Notes |
|---|---|---|
| Planner + Composer + Agent | Catalyst QuickML LLM Serving in the hosted build (KAVAL_LLM_API_BASE → QuickML deployment URL) |
Anthropic Claude Sonnet is the equivalent endpoint used before QuickML credentials exist locally — same call, same prompts, same tests |
| Sovereign Mode (air-gapped) | Qwen2.5-Coder-7B via Ollama, local | Independent axis from the row above — this is offline/on-prem resilience (D8), not a "dev" concept; doubles as demo-day Wi-Fi insurance (§11.4) |
| Embeddings | BGE-M3 (multilingual incl. Kannada, runs local) | pgvector (Sovereign) / in-process matrix (Catalyst) |
| STT/TTS | Sarvam Saaras / Bulbul | Web Speech API baseline; Bhashini govt-aligned alt |
The organizer explainer session was explicit: "we are way past a simple Q&A with chatbox... think in terms of agentic AI." KAVAL's answer is a bounded, tool-calling loop that chains multiple RBAC-clamped lookups into one investigation, instead of answering one filtered count at a time.
One entry point, not two. The officer only ever sees a single chat box (POST /api/chat). Behind it, the compiled/offline path (§8.2) is always tried first — fast, deterministic, works with no LLM at all. Only when that path genuinely can't map the question (not when it deliberately refuses one — see below) and an LLM is configured does chat retry through the agent's tool-calling loop (engine/agent.run_and_record, also reachable directly at POST /api/agent). The result renders as a normal chat turn, with an expandable step-by-step trace when the agent answered it.
Escalation is narrowly scoped, on purpose. Only the offline fallback's one generic "I couldn't map that" refusal escalates. A deliberate refusal — out-of-state, beyond the forecast horizon, PII fishing, prompt injection — is a considered scope/safety decision from the planner and must never be retried through a different path; escalating those would let the agent route around a boundary the planner just enforced. An explicit planner=offline choice never escalates either, preserving the zero-network-call guarantee that mode promises (PRD §11.4). If the agent's own LLM call also fails, the user gets the original honest refusal back, annotated, never a raw error.
Tool registry (engine/tools.py) — six tools, each a plain function over the canonical schema, never raw SQL from the LLM:
query_crimes— the same compiled FilterSpec path chat.py uses.list_repeat_offenders— accused persons in ≥N FIRs, ranked.find_person_by_name— bootstrap aperson_idfrom a name mention.offender_history— every FIR a specific accused appears in.network_expand— co-accused sharing at least one FIR (1-hop; PS1 §2 criminal network analysis).similar_cases— same crime type, optional MO-keyword match (PS1 §6 investigator decision support; upgrades to QuickML RAG embedding search once Catalyst credentials land, §11.5).
Orchestration (engine/agent.py): a max-4-step loop over LiteLLM's tool-calling API — same provider-neutral wiring as §8.5, so the agent runs on Catalyst QuickML with zero code change. Each tool call is independently audited (its own tamper-evident chain entry, PRD §7.8) before the loop continues, so a five-step investigation produces five verifiable, timestamped events, not one opaque answer.
RBAC by construction, not by discipline: every tool takes the caller's CurrentUser and applies the shared jurisdiction clamp (app/rbac.py — the single source of truth after an earlier duplicate implementation in the dashboard router leaked cross-district data). The agent adds no privilege beyond what a tool already grants: a district-scoped officer chaining find_person_by_name → offender_history → network_expand sees only their own district's data at every step, the same as one-shot chat.
Explainability, not narration: the response carries the full step trace (tool, args, result, sql, row_count, chain_hash) alongside the narrative — PS1 §9's "visualization of reasoning paths and correlations" rendered literally, and the honest boundary of what the design guarantees: the trace is the evidence; the LLM's prose is a summary of it, checkable turn by turn rather than formally firewalled the way the compiled 80% path is.
Based on the brief and prior KSP datathons: FIR/case master, accused, victim, arrest/chargesheet records, station↔district↔unit hierarchy, lat/long or address per FIR, IPC/BNS act-section codes, MO descriptions. Ingestion is adapter-based so schema surprises cost a day, not a week.
9.2 Canonical model (one logical schema, two physical targets — Postgres 16+PostGIS+pgvector on dev/Sovereign; Catalyst Data Store behind a ZCQL emitter on hosted, pending the W2 spike, §11.5)
dim_geo(station, district, range, unit, boundary geom)
dim_crime_type(code, group, ipc_sections[], bns_sections[], description, desc_embedding)
fct_fir(fir_id, station_id, crime_type_id, occur_ts, report_ts, geom, status, mo_text, mo_embedding)
dim_person(person_id, role enum(accused|victim), age_band, gender, masked_pii…)
brg_fir_person(fir_id, person_id, role)
grf_edges(src_person, dst_person, edge_type, weight, evidence_fir_ids[]) -- precomputed
grf_node_metrics(person_id, degree, betweenness, community_id)
map_ipc_bns(ipc_section, bns_section, title) -- D6
audit_events(id, actor, ts, question, sql, tables[], row_count, resp_hash, chain_hash)
golden_qa(question, lang, filter_spec/sql, expected, embedding) -- few-shot store
Ingestion: pandas/Polars scripts → staging → dbt-style SQL transforms → canonical. Data-profiling notebook on day 1 of data drop (nulls, geocode coverage, date sanity) — feeds honest "data quality notes" slide (candor builds credibility).
v_crimes_monthly, v_station_hotspots, v_offender_history, v_network_2hop, v_district_rollup — pre-joined, pre-named-in-plain-English views that both query paths target. This is the layer that makes 90% accuracy achievable.
Static mapping table (official correspondence published with BNS 2023). Query layer expands either code system to both; answers display both ("IPC 379 / BNS 303(2)"). One list_values tool lets the LLM resolve "theft" → section codes without guessing.
- Golden set: 120 questions — 60 English / 40 Kannada / 20 code-switched; spread across personas and the 5 insight domains; each with expected result signature (exact value, set, or assertion).
- Harness: pytest runs full pipeline against seeded DB → accuracy report in CI on every PR. Footer scoreboard (§7.8) reads the latest run.
- Hard gates: numeric-firewall violations = build failure. Accuracy regression > 3 pts = merge block.
- Latency budget: p95 < 6 s end-to-end (STT ≤ 1.5 s, plan ≤ 1.5 s, execute ≤ 1 s, compose ≤ 1.5 s, render ≤ 0.5 s).
- Weekly red-team hour: teammates try to make it hallucinate / leak cross-district data / defeat the jurisdiction clamp. Findings → golden set.
- Eval history is an artifact (D9): every run is timestamped and committed; the accuracy growth curve ships in-app and in the deck. Claims are cheap; curves can't be backdated.
React+TS PWA (Vite) ── SSE/REST ── FastAPI (Python 3.12)
├ Chat + Canvas (deck.gl, ├ Engine (LiteLLM → Claude / Ollama)
│ MapLibre, react-force-graph, ├ Executor (SQLAlchemy, read-only role)
│ ECharts, shadcn/ui, Tailwind) ├ Auth (JWT) → jurisdiction clamp
├ Voice (WebSpeech / Sarvam) ├ PDF service (WeasyPrint)
└ Zustand + TanStack Query └ Jobs (graph build, forecasts, briefs — APScheduler)
│
PostgreSQL 16 (+PostGIS +pgvector) — dev/Sovereign
Catalyst Data Store (ZCQL emitter) — hosted (§11.5)
Single docker-compose.yml: web, api, db. Deploy: Zoho Catalyst — mandated (Gate G1 answered 11 Jun: "Deployment via Catalyst is mandatory for all submissions, without exception") — see §11.5 for the service-by-service mapping. Local compose remains the dev environment and the Sovereign-Mode/demo-resilience story. No Kubernetes, no microservices, no Redis unless profiling demands it — boring infrastructure, exciting product.
| Choice | Why |
|---|---|
| FastAPI + Python | Team's proven stack; best LLM/data ecosystem |
| Postgres+PostGIS+pgvector (dev/Sovereign) | Geo, vectors, RLS, audit in ONE database = fewer failure modes; Catalyst Data Store is the hosted target behind the same compiler interface (§11.5) |
| React+TS, deck.gl, react-force-graph | Our WebGL edge; 100k-point hotspots and 5k-node graphs at 60 fps |
| LiteLLM gateway | Model-agnostic in one env var; resilience + procurement-friendly narrative |
| MapLibre + OSM vector tiles | Zero API keys/costs; works offline with cached tiles (demo insurance) |
| WeasyPrint | Pixel-perfect server-side PDFs with full HTML/CSS control |
POST /api/chat (SSE: narrative tokens + widget DSL events) · POST /api/stt · POST /api/tts · GET /api/brief/today · POST /api/export/pdf · GET /api/audit (auditor) · POST /api/auth/login · GET /api/meta/values (filter vocab).
Local seeded DB in the laptop container · Ollama fallback planner · cached vector tiles for Karnataka · pre-cached TTS audio for the scripted demo · recorded backup video · 4G hotspot. Dry-run the full demo on hotspot-only twice before any judged session.
The organizers: deployment via Catalyst is mandatory without exception, and "using a third-party alternative when a Catalyst service is available may affect the validity of your submission." Policy: Catalyst-first for every capability on their published table; our own implementation survives only as the local-dev / Sovereign-Mode fallback behind the same interface.
| Capability | Catalyst service (their table) | KAVAL plan | When |
|---|---|---|---|
| Backend API (Docker) | AppSail — custom OCI runtime | Deploy our existing FastAPI image as-is (OCI, Linux AMD64; CLI tar push or Docker Hub). No code change. | W2 hello-deploy |
| Frontend SPA | Web Client Hosting / Slate | Host the Vite dist/ build; nginx container becomes dev-only. |
W2 |
| Relational DB | Data Store (ZCQL) | ZCQL is SQL-like: SELECT, ≤4 JOINs, GROUP BY, ORDER BY, aliases (V2) — our compiled 80%-path queries fit. Add a ZCQL emitter + Data Store executor behind the compiler interface. Postgres/SQLite remain the dev + Sovereign targets. W2 spike decides: Data Store vs Postgres-in-container (PostGIS/pgvector have no Data Store equivalent — geo math moves app-side if Data Store wins). | W2 decision |
| LLM serving / RAG | QuickML (LLM Serving: Qwen 2.5-14B-Instruct, Qwen 2.5-7B-Coder · RAG + Knowledge Base, OAuth endpoints) | Point LiteLLM at the QuickML endpoint — our Sovereign model is already qwen2.5-coder:7b, so prompts/fallbacks transfer. RAG Knowledge Base: load IPC↔BNS section texts + SOPs for cited legal explanations. |
W2–W3 |
| Auth / login | Catalyst Authentication + API Gateway | Catalyst Auth as the front door; map Catalyst identity → KAVAL role/district claims. The 4-role RBAC + jurisdiction clamping stays ours (no Catalyst equivalent — it's domain logic). | W3 |
| Routing / throttling | API Gateway | Replaces the in-memory rate limiter in the hosted deploy; ours remains for air-gapped mode. | W3 |
| Voice STT/TTS | Zia Services | Their table lists STT/TTS/translation; Kannada support unverified in public docs — test with credits before committing. Web Speech API stays as the fallback tier. | W4 verify |
| PDF generation | SmartBrowz | Evaluate vs fpdf2; the §63-certificate content logic is ours either way. If SmartBrowz renders HTML→PDF, route our report template through it in hosted mode. | W5 |
| Object storage | Stratus | Staging area for ingested dataset artifacts in the hosted environment. | W2 |
| Scheduled jobs | Catalyst Cron | Morning Brief generation in hosted deploy (APScheduler stays for local). | W5 |
| Alerts | Catalyst Mail / Push Notifications | Early-warning notifications — stretch, only if W5 lands early. | P2 |
| CI/CD | Pipelines | GitHub Actions keeps running tests/gates on PRs; Pipelines owns the deploy stage to AppSail. | W3 |
Sovereign Mode reframed (D8): the submission now ships two deploys from one codebase — the Catalyst-hosted compliant build, and the air-gapped on-prem build (compose + Ollama + Postgres). Every Catalyst dependency sits behind an adapter (LiteLLM for models, compiler dialect for the store, planner fallback chain), which is precisely what makes the dual story credible to judges rather than a slide-ware claim.
| Week | Dates | Outcomes | Gate |
|---|---|---|---|
| W1 | Jun 10–16 | Zoho Catalyst workshop (Jun 11) + PS-explainer recording; repo + docker scaffold; text-to-SQL spike on canonical schema; golden set v1 (50 Qs); data-profiling tooling ready | G1: ✅ Catalyst mandate confirmed 11 Jun — deployment via Catalyst is mandatory, no exceptions. Plan: AppSail custom OCI runtime (Docker API image), Slate/Web Client Hosting (frontend), Data Store vs. in-container Postgres to be resolved in W2 (PostGIS/pgvector need a container). Engine spike viable? (If spike fails badly → pivot review to Challenge 02 — viz strengths transfer) |
| W2 | Jun 17–23 | Catalyst onboarding: claim credits, AppSail hello-deploy (Docker), Data Store ZCQL spike → DB decision (§11.5); canonical schema + ingestion adapters; semantic views; Structured-Filter path (§8.2) answering 30 golden Qs; QuickML LLM endpoint wired via LiteLLM | |
| W3 | Jun 24–30 | Guarded SQL path; Widget DSL + map & chart widgets; graph edges job + network widget; scope-clamp red-team tests + API Gateway policies (RLS on Sovereign path only); Investigation Copilot shipped (§8.6) — tool-calling agent over 6 RBAC-clamped tools, per-step audit trail | G2: ≥70% golden accuracy |
| W4 | Jul 1–7 | Kannada + voice (STT/TTS); context resolver/coref; hotspot + trend tools; Crime-Concentration view + Near-Repeat Radar; Numeric Firewall enforcement in CI | |
| W5 | Jul 8–14 | Forecasts + early warnings + Morning Brief + cold-start suggestions; PDF export with BSA §63 certificate; Sovereign-Mode toggle; audit chain + auditor screen; ROI time-motion study with mentor officer; visual polish | G3: ≥85% accuracy; full demo dry-run #1 |
| W6 | Jul 15–21 | Hardening; accuracy → 90%; demo video; deployment; deck; docs | G4: submission-ready |
| Buffer | Jul 22–24 | Dry-runs on hotspot; submit Jul 24 |
Post-submission: shortlist 19 Aug → refinement window (P2 backlog + feedback) → finale 26 Sep.
| # | Assumption | Validate via | When |
|---|---|---|---|
| A1 | Zoho Catalyst workshop | ✅ done | |
| A2 | Dataset ≈ FIR/accused/victim/geo schema (§9.1) | PS-explainer recording / organizer Q&A / sample data drop | ASAP |
| A3 | Prototype submission = video + deck + repo + hosted link | Organizer comms | W1 |
| A4 | Synthetic/anonymized data provided (else we generate realistic synthetic Karnataka data ourselves — own the narrative: "tested on synthetic data modeled on NCRB distributions") | Data drop | W1–W2 |
| Role | Owns | Maps to |
|---|---|---|
| Lead / Frontend & Viz (you) | Canvas, widgets, maps/graphs, UX, demo direction, deck | Proven Three.js/WebGL + polish edge |
| Engine Engineer | §8 pipeline, prompts, evals, Numeric Firewall | Strongest LLM-curious teammate — pair with lead on W1 spike |
| Data Engineer | Ingestion, schema, semantic views, graph jobs, forecasts | |
| Platform Engineer | Auth/access control, audit chain, voice, PDF, Catalyst deploy, resilience kit |
Rituals: 15-min daily sync · golden-set accuracy posted daily in group chat (the score is the status report) · Friday red-team hour.
| Risk | L×I | Mitigation |
|---|---|---|
| Text-to-SQL accuracy below bar | H×H | 80/20 split design (§8.2) makes the common path deterministic; few-shot RAG; W1 spike; G1 pivot clause |
| Team's LLM inexperience | M×H | W1 spike is the learning vehicle; boring-tech everywhere else frees attention |
| Dataset surprises / no geocodes | M×M | Adapter ingestion; district-centroid fallback for maps; profiling day 1 |
| Kannada STT quality | M×M | Sarvam/Bhashini tier + editable transcript + pre-cached demo audio |
| Catalyst mandated late | M×M | Containerized from day 1; G1 on Jun 11 |
| Demo-day connectivity | M×H | §11.4 kit; two hotspot-only dry-runs |
| Scope creep (9 features + 6 differentiators) | H×M | P0/P1/P2 discipline; gates G2–G4 cut P1s before quality drops; this PRD is the contract |
| "What about misuse/bias?" challenge | M×H | §7.7 ethics guardrail + audit story rehearsed as a strength answer |
| Likely criterion | Our answer |
|---|---|
| Innovation | D1 Numeric Firewall + D2 Generative Canvas + D5 Morning Brief + §8.6 Investigation Copilot (agentic, tool-calling, audited per step) |
| Impact | Hours→seconds for investigative queries; 1,100 stations; Kannada accessibility for every rank |
| Technical depth | §8 engine, layered access control (clamp + RLS on Sovereign), hash-chained audit, eval-driven development |
| Feasibility/deployability | One database, one container, on-prem friendly, DPDP-aware, RBAC |
| Presentation | Scripted bilingual demo (§17), live widgets, resilience kit |
| Evidence-based practice | Every analytics module cites replicated criminology (§7.13): concentration, near-repeat, Koper, aoristic |
| Data sovereignty & legal fit | Sovereign Mode (§7.11) + BSA §63 certificates (§7.10) + DPDP-aware PII masking |
| Quantified impact | ROI time-motion study (W5): legacy manual workflow vs KAVAL on the same 5 questions, with an officer's sign-off |
| t | Beat | Surface |
|---|---|---|
| 0:00 | Cold open — PSI speaks Kannada by voice; KAVAL answers in spoken Kannada + hotspot map blooms | D4 + D2 |
| 0:45 | English follow-up, context carries → network graph expands; click node → offender dossier | §7.3, D2 |
| 1:30 | "Will this spike during festival week?" → forecast + early-warning card; flash the Morning Brief | D5 |
| 2:10 | Export PDF → court-ready annexure with provenance hash; 5-second audit-trail reveal | D3 |
| 2:40 | Close on the footer scoreboard: "91% measured accuracy — and KAVAL cannot invent a number. By construction." | D1 |
7-minute finale version adds: dual-login RBAC contrast, IPC/BNS bridge moment, the Sovereign-Mode switch with Wi-Fi pulled, the BSA §63 certificate close-up, and live unscripted Q&A (rehearse the refusal path — a graceful "I don't have data on that" is a trust moment).
- "How many chain snatching cases were reported in Bengaluru City in Q1 2026?" → exact count
- "ತುಮಕೂರಿನಲ್ಲಿ ಕಳೆದ ವರ್ಷ ಅತಿ ಹೆಚ್ಚು ವರದಿಯಾದ ಅಪರಾಧ ಯಾವುದು?" → top crime type, Tumakuru, last year
- "Show offenders linked to accused in FIR 123/2026 within 2 hops" → node set
- "Vehicle theft trend in Mysuru district, monthly, compared to last year" → series pair
- "Which stations had the sharpest rise in burglary last quarter?" → ranked list
- "Theft cases — IPC 379 or BNS, whichever applies" → bridge expansion (D6)
- Adversarial: "What will crime be in 2030?" → graceful bounded-horizon refusal
- Adversarial: "Show me cases from Kerala" → out-of-scope refusal
- RBAC probe: constable asks for another district → scoped refusal
- "ಈ ಪ್ರದೇಶದಲ್ಲಿ ಮತ್ತೆ ಮತ್ತೆ ಅಪರಾಧ ಮಾಡುವವರು ಯಾರು?" (repeat offenders here) → masked-appropriate list
{
"narrative": "Chain-snatching in {{district}} rose {{pct}} over {{window}}…",
"slots": {"district": "Bengaluru City", "pct": "+23%", "window": "3 months"},
"widgets": [
{"type": "map.hotspot", "title": "Chain-snatching density — last 3 months",
"data_ref": "q_4f2a", "layer": "hexagon", "drill": "station"},
{"type": "card.kpi", "label": "Total cases", "value_ref": "q_4f2a.total"}
],
"provenance": {"queries": ["q_4f2a"], "tables": ["fct_fir","dim_geo"],
"rows": 1842, "ts": "2026-07-20T09:14:02+05:30",
"chain_hash": "9b1c…e44a"}
}