Skip to content

Latest commit

 

History

History
535 lines (423 loc) · 47.2 KB

File metadata and controls

535 lines (423 loc) · 47.2 KB

PRD — Project KAVAL (ಕಾವಲು)

Intelligent Conversational AI for the KSP Crime Database

Karnataka State Police Datathon 2026 — Challenge 01

Version 1.2
Date 10 Jun 2026
Status Draft — pending validation at PS Explainer recording + Zoho Catalyst workshop (11 Jun)
Prototype deadline 26 Jul 2026 (submit by 24 Jul — 2-day buffer)
Grand Finale 26 Sep 2026, in-person demo day

KAVAL (ಕಾವಲು) — Kannada for "the watch / the guard." A name the judging panel reads in their own language. Backronym if anyone asks: Knowledge-Augmented Vigilance & Analytics Layer.


1. One-line pitch

"Every police officer in Karnataka gets a crime analyst in their pocket — one that speaks Kannada, shows its work, and never invents a number."

KAVAL turns the SCRB's records from 1,100+ police stations into a conversational intelligence layer: an investigator asks a question by voice or text, in Kannada or English, and gets back a verified answer rendered as live interactive maps, criminal-network graphs, and trend charts — every answer carrying court-grade provenance.


2. Problem statement (verbatim interpretation)

SCRB manages a large, growing repository of crime data from 1,100+ stations. Current systems = static dashboards + manual queries → no deep analysis, no real-time insight.

Build: a conversational AI platform for investigators to query crime data in natural language and uncover patterns, relationships, and predictive insights.

Brief's required capability list (all are mapped to features in §7):

  1. Natural-language chatbot (English + Kannada)
  2. Voice-enabled interaction
  3. Context-aware conversations
  4. PDF export of conversation history
  5. Criminal network visualization
  6. Crime trend & hotspot detection
  7. Predictive analytics & early warnings
  8. Explainable AI with audit trails
  9. Role-based secure access

Insight domains: crime pattern discovery · criminal network analysis · socio-demographic insights · behavioral profiling · proactive crime-prevention intelligence.


3. Goals & non-goals

Goals (prototype, by 24 Jul)

  1. 90% accuracy on a 120-question bilingual golden evaluation set (see §10).
  2. Zero hallucinated numbers — enforced by architecture, not by prompting (§8.4 Numeric Firewall).
  3. End-to-end answer (question → rendered widget) in < 6 s p95.
  4. All 9 brief capabilities demonstrable; 6 of them production-quality (P0 set).
  5. A 3-minute demo that produces an audible reaction from a non-technical judge.

Non-goals (explicitly out of scope — discipline wins hackathons)

  • ❌ Real-time CCTV/face recognition, image/video analytics
  • ❌ Mobile native apps (responsive PWA is enough)
  • ❌ Training/fine-tuning custom LLMs
  • ❌ Live integration with actual CCTNS/KSP production systems
  • ❌ Full Kannada UI localization (chat I/O bilingual; chrome stays English)
  • ❌ Multi-tenancy beyond the role model in §7.9

4. Users & personas

Persona Role Core job-to-be-done Key features
IO Investigator (SHO / PSI) Field investigation "Who else is connected to this accused? What's the pattern in my beat?" Chat, network graph, case lookup, PDF export
Crime Analyst (SCRB / DCRB) Pattern analysis "Where are emerging hotspots? Which MO clusters are active?" Hotspots, trends, forecasts, drilldowns
Command (SP / DCP / DGP office) Resource allocation "Daily situational picture; where do I deploy?" Morning Brief, early warnings, district rollups
System Admin / Auditor Compliance "Who accessed what, when, and why?" Audit trail, role management

Demo personas: we will demo as IO Investigator (story 1) and Command (story 2). Two personas max — more dilutes the narrative.


5. Product thesis — six differentiators

Everything in this PRD serves at least one of these.

# Differentiator Why it matters
D1 Numeric Firewall — the LLM architecturally cannot state a number that didn't come from an executed query (§8.4) Kills the #1 failure mode of conversational analytics. One sentence to stakeholders: "Our AI cannot invent a statistic — by construction, not by promise."
D2 Generative Analytics Canvas — answers materialize as live, interactive WebGL maps/graphs/charts inline in the chat, not text walls (§7.4) Insight you can touch: drill, pan, expand — the difference between a report and an instrument.
D3 Court-ready provenance ("Evidence Locker") — every answer carries the query, source tables, row counts, timestamp, officer ID, and a tamper-evident hash; PDF exports are case-file-grade (§7.8, §7.10) Reframes a chatbot as an evidence instrument. Police workflows think in case files.
D4 Kannada-first voice via Indic-specialized STT/TTS (Sarvam/Bhashini), incl. code-switched "Kanglish" (§7.2) Accessibility for every rank, in the language officers actually speak — built on Indian sovereign AI.
D5 KAVAL Morning Brief — the system speaks first: an auto-generated daily situational brief per station/district with anomalies and early warnings (§7.7) Converts "reactive Q&A bot" into "proactive intelligence officer."
D6 IPC ↔ BNS bridge — seamless querying across the 2024 IPC→Bharatiya Nyaya Sanhita transition (§9.4) Deep domain awareness: every officer lives this transition pain daily.
D7 BSA §63 court-admissibility certificate — every PDF export ships with an auto-generated certificate under §63(4), Bharatiya Sakshya Adhiniyam 2023 (successor to IEA §65B) (§7.10) Turns "court-ready" from metaphor into legal workflow. Instantly legible to police leadership; invisible to teams without domain knowledge.
D8 Sovereign Mode — one visible toggle swaps the cloud LLM for an on-prem model; fully functional air-gapped (§7.11) Answers the unasked procurement question: crime data never leaves the building. Demo beat: flip the toggle, pull the Wi-Fi, keep answering.
D9 Measured, not claimed — published failure-mode/refusal page + live accuracy scoreboard + committed eval history showing the accuracy growth curve (§7.8, §10) Every team claims quality; a timestamped improvement curve cannot be faked retroactively.
D10 Cold-start intelligence — role-aware suggested questions generated from the data on login (§7.12) Kills the blank-textbox problem that sinks real chatbot deployments; reframes KAVAL from search box to colleague.

6. Critical user journeys

Journey A — The investigator (demo story 1)

  1. PSI opens KAVAL, speaks in Kannada: "ಕಳೆದ ಮೂರು ತಿಂಗಳಲ್ಲಿ ನನ್ನ ಠಾಣಾ ವ್ಯಾಪ್ತಿಯಲ್ಲಿ ಚೈನ್ ಸ್ನ್ಯಾಚಿಂಗ್ ಪ್ರಕರಣಗಳು ಎಲ್ಲಿ ಹೆಚ್ಚಾಗಿವೆ?" (Where have chain-snatching cases risen in my station limits in the last 3 months?)
  2. KAVAL answers in spoken Kannada + renders a hotspot map widget scoped to her jurisdiction (RBAC-enforced).
  3. Follow-up in English (context carries): "Show repeat offenders active in these zones." → Network graph widget; nodes sized by case count.
  4. She clicks a node → offender panel: linked FIRs, co-accused, MO history, IPC/BNS sections.
  5. "Export this for the case file." → court-ready PDF with provenance hashes lands in 5 seconds.

Journey B — Command (demo story 2)

  1. SP opens the Morning Brief: 3 auto-generated cards — a vehicle-theft anomaly in District X (+38% week-over-week), a forecast spike flagged for the upcoming festival window, a newly active offender network.
  2. Clicks the anomaly → drilldown chat pre-seeded with context; asks "compare with the same festival period last year" → comparative trend chart.
  3. Every screen shows the ⓘ provenance chip: which query, which tables, how many rows, when.

7. Functional requirements

Priorities: P0 = must ship for prototype · P1 = ship if on schedule · P2 = finale-phase / stretch.

7.1 Conversational engine — NL → verified insight 〔P0 — the crown jewel〕

  • Accepts English, Kannada, and code-switched input; text or voice.
  • Pipeline (full spec §8): intent routing → tool selection → guarded query execution → verified composition.
  • Behavioral contract:
    • Never answers data questions from model memory — only from executed queries.
    • Ambiguous question → asks ONE crisp clarifying question (e.g., "Calendar year or last 12 months?").
    • No data → says so plainly. Refusal is rendered as a first-class, well-designed state, not an error.
    • Every numeric/factual claim traceable to a query (D1).
  • Streaming responses; typed answer skeleton appears < 1.5 s.

7.2 Bilingual + voice interaction 〔P0 voice-in/text; P1 voice-out〕

  • STT: Web Speech API (kn-IN, en-IN) as the zero-cost baseline; Sarvam Saaras (or Bhashini) API as the quality tier — Indic-specialized, handles Kanglish code-switching. Runtime-selectable.
  • TTS: Sarvam Bulbul (Kannada voice) for spoken answers; pre-cache demo-script audio for offline resilience (§11.4).
  • Language of the answer follows language of the question; user can pin a language.
  • Mic button with live waveform; transcript editable before send (officers will want to correct STT).
  • Romanized Kanglish 〔P1〕: Latin-script Kannada ("Indiranagar alli ninne yeshtu theft aagide?") normalized before planning — this is how officers actually type; almost no team will handle it.

7.3 Context-aware conversation 〔P0〕

  • Session memory: entity tracking (accused, stations, districts, crime types, time windows) so follow-ups resolve "he / that area / same period" correctly.
  • Active-filter chips above the chat (e.g., District: Tumakuru ✕ Period: Jan–Mar 2026 ✕) — visible, removable state. Evaluators see the context awareness instead of taking our word.
  • "New thread" resets context explicitly.

7.4 Generative Analytics Canvas 〔P0 — D2〕

The engine returns a Widget DSL (JSON spec, appendix B) alongside narrative; the frontend renders interactive widgets inline:

  • map.hotspot — deck.gl heat/hex layer on MapLibre vector tiles of Karnataka, station/district boundaries, click-to-drilldown.
  • graph.network — WebGL force-directed criminal network (react-force-graph); node = person/case, edge = co-accused/shared-FIR/shared-MO; cluster coloring, expand-on-click.
  • chart.trend — ECharts time series with anomaly bands, period comparison.
  • table.records — paginated, sortable case/record table with PII masking per role.
  • card.kpi, card.alert — stat and early-warning cards.
  • All widgets carry the provenance chip ⓘ (§7.8) and an "open full-screen" mode for the demo projector.

7.5 Criminal network analysis 〔P0 graph viz; P1 metrics〕

  • Offline job builds an edge table: co-accused in same FIR, shared addresses/phones (if dataset provides), shared MO cluster, repeat-victim links.
  • networkx computes degree/betweenness centrality + Louvain communities → stored as node attributes.
  • Investigator flows: "show network of accused #X (2 hops)", "most central figures in vehicle-theft networks in District Y", repeat-offender tracking (≥ N FIRs in window).

7.6 Crime trends & hotspot detection 〔P0〕

  • Spatial: KDE / hex-bin aggregation over geocoded FIRs — PostGIS on the dev/Sovereign path, app-layer haversine + grid binning on Catalyst (Data Store has no spatial extension; at FIR scale this costs nothing); granularity district → station → grid.
  • Temporal: weekly/monthly series per (geo × crime-type); STL decomposition for seasonality; z-score anomaly flags.
  • Socio-demographic cuts (age bands, victim/accused profiles) where the dataset supports it — validate at data drop (§13, A2).
  • Maps drawn against police-station jurisdiction boundaries where obtainable (officers think in station limits, not census geography); graceful fallback to district boundaries.

7.7 Predictive early warning + Morning Brief 〔P1 — D5〕

  • Forecasting: gradient-boosted / Prophet per (district × major crime type), honest 4-week horizon with confidence bands. Transparent features only (recency, frequency, seasonality, festival calendar) — no black-box claims.
  • Early-warning rules: forecast breach, week-over-week anomaly, newly-active network detection.
  • Morning Brief: daily auto-generated per role scope — 3–5 cards, each one click from a drilldown chat. LLM writes the prose; numbers injected via Numeric Firewall.
  • Ethics guardrail: we score places and patterns, never communities or individuals' future criminality. One slide in the deck on this — responsible-AI questions deserve a prepared answer.

7.8 Explainability, provenance & audit 〔P0 — D1/D3〕

  • Per answer: interpreted intent (plain English/Kannada), executed SQL (collapsible), source tables, row count, execution time, model+version, confidence tier.
  • Audit log: append-only audit_events — who, when, question, query, tables touched, rows returned, response hash; each row carries sha256(prev_hash ‖ row) → tamper-evident chain. Auditor screen renders the chain with verification status.
  • Live accuracy scoreboard in the footer: our golden-set accuracy, updated by CI. Radical transparency as a feature.
  • Published failure modes (D9): an in-app page listing the question classes KAVAL refuses and why (out-of-jurisdiction, beyond forecast horizon, insufficient data) — we present our own attack surface before anyone probes it.
  • Accuracy growth curve (D9): every eval run is timestamped and kept; the improvement curve (e.g., 71% → 93% across the build weeks) renders in-app and in the final deck.

7.9 Role-based secure access 〔P0 skeleton; P1 full〕

  • Roles: CONSTABLE (own-station, PII-masked) · INVESTIGATOR (own-district, full records) · COMMAND (state-wide aggregates + district drill) · ADMIN/AUDITOR (audit + user mgmt, no case data).
  • Enforced by a jurisdiction clamp in the query pipeline (shipped W1.5): every spec is rewritten to the caller's scope before compilation, and out-of-scope asks become audited refusals — never enforced in the prompt. Postgres RLS adds a second, DB-layer net on the Sovereign path; Catalyst API Gateway fronts the hosted deploy. One demo beat: same question, two logins, two different (correctly scoped) answers.
  • JWT auth (httpOnly), bcrypt, session timeout. PII masking (names/phones partially redacted) below INVESTIGATOR.

7.10 Court-ready PDF export 〔P0 — D3〕

  • Server-side (WeasyPrint): KSP-styled header, conversation Q&A, embedded widget snapshots (server-rendered PNG), and an annexure of provenance — every query, table list, row counts, timestamps, officer identity, audit-chain hash in the footer.
  • BSA §63 certificate (D7): auto-generated certificate for electronic records under §63(4), Bharatiya Sakshya Adhiniyam 2023 (successor to the IEA §65B certificate every IO knows) — system identification, output hash, operator declaration — appended to every export. ~1 day of work; legal review of wording with a police mentor.
  • Framing in demo: "This isn't a chat log — it's an annexure ready for the case file, certificate included."

7.11 Sovereign Mode 〔P1 — D8〕

  • A visible toggle in the header: Cloud (Claude via gateway) ↔ On-prem (local model via Ollama). Same pipeline, same Numeric Firewall — only the planner endpoint changes (the LiteLLM abstraction makes this nearly free).
  • Acceptance bar: the scripted demo questions answer correctly with Wi-Fi off, end-to-end on one laptop.
  • Honest framing: local-model accuracy will be lower — show its own scoreboard entry rather than hiding it. The point is the deployment option, not parity.
  • Demo beat (finale): flip the toggle, disconnect the network, answer a question. One sentence: "Crime data never has to leave the building."

7.12 Cold-start intelligence 〔P1 — D10〕

  • On login, KAVAL speaks first: 3 role-scoped suggested questions generated from the data, not hardcoded — e.g., "Vehicle theft in your limits is up 22% this month — want the station breakdown?"
  • Implementation: nightly job ranks anomalies/deltas per scope (reuses Morning Brief plumbing, §7.7); each suggestion is one tap from a full answer.
  • Solves the blank-textbox problem that kills chatbot adoption; doubles as silent onboarding.

7.13 Evidence-backed analytics modules

Each module below is grounded in replicated criminology research — we cite the evidence base on its deck slide and in-app ⓘ. Rule: ship nothing we can't defend to a criminologist.

Module What it does Evidence base Needs from dataset Priority
Crime Concentration (Pareto) view "4% of places produce ~50% of crime in this district" — ranked micro-places with cumulative curve Law of crime concentration (Weisburd 2015, Criminology): crime concentrates at a small share of micro-places, remarkably stable across cities Geocoded FIRs P0 — one query, huge resource-allocation story
Near-Repeat Radar After a qualifying burglary/vehicle theft, auto-draw a 250–400 m, 7–14 day elevated-risk envelope; alert the beat Near-repeat victimization (Townsley et al. 2003; Bowers & Johnson 2004) — robust, widely replicated spatio-temporal pattern. Pattern is solid; prevention payoff depends on patrol follow-through, and we say so FIR lat/lon + occurrence date P1
Aoristic Clock True time-of-day risk for discovery crimes (burglary, vehicle theft) by spreading probability across the occurrence from–to interval Aoristic analysis (Ratcliffe): occurrence times for unwitnessed crimes are intervals, not points — naive "report hour" charts are systematically wrong. Indian FIRs record occurrence from → to, a perfect fit Occurrence from/to timestamps (verify at data drop) P1
Repeat-victimization flags Addresses/persons victimized ≥2× flagged for prevention outreach lists Farrell & Pease (1993) and successors: a small share of victims accounts for an outsized share of incidents; repeat-victim prevention is among the best-evidenced tactics Victim identifiers / addresses P1
Koper patrol planner Convert hotspots into a suggested schedule of 10–15 min patrol dwells Koper curve (1995, replicated): 10–16 min hotspot dwells maximize residual deterrence; hotspot-policing meta-analyses (Braga et al.) show consistent reductions without simple displacement Hotspot output (already have) P2
Solvability triage Score fresh FIRs on solvability factors to prioritize investigative effort UK evidence-based investigative triage (Kent Police EBIT trials); solvability-factor literature (Wellford & Cronin) Case outcomes (FIR status — likely have) P2
RTM-lite environmental risk Risk layers from environment (bars, ATMs, bus stops, transit) using OSM POIs — forecasts from places, not past arrests Risk Terrain Modeling (Caplan & Kennedy, Rutgers) — also our strongest bias answer: we score environments, never communities Public OSM data only P2 stretch

7.14 Sociological crime insights 〔P1 — activates at data drop〕

PS1 §4 asks for it explicitly; ours is wired, awaiting fields:

  • Demographic cuts as first-class spec dimensions: age band, gender, victim/accused role — dim_person (§9.2) already models them; they become group-bys and filters the moment data-drop profiling confirms the fields exist.
  • District-level correlation overlays: crime rates joined against public socio-economic indicators (Census urbanization, literacy, sex ratio) — scatter/choropleth widgets with the correlation≠causation caveat printed on the widget itself, not in fine print.
  • Evidence framing: the social-disorganization tradition (Shaw & McKay) explains why place-level socio-economic context belongs next to crime data. We present patterns over places and aggregates, never community blame — §7.7's ethics guardrail applies verbatim.

7.15 Financial crime & transaction link analysis 〔conditional — PS1 §7〕

  • Activates if the dataset carries financial fields (cyber-fraud FIRs commonly include account / UPI / phone identifiers). Go/no-go at data-drop profiling, documented either way.
  • If present: accounts, UPI handles, and phone numbers become typed nodes in the §7.5 network graph; a "money trail" is a path query over those edges; fan-in/fan-out and same-account-across-FIRs rules flag suspicious clusters.
  • If absent: the identical graph machinery demos on co-accused / shared-MO edges and the adapter ships documented — the answer to "where's financial crime?" is "wired, awaiting fields", never silence.

7.16 P2 / finale-phase backlog

Saved-question shortcuts per role · scheduled report subscriptions · WhatsApp-style brief delivery · Catalyst-native notifications · cross-district MO pattern matching via embedding case linkage (Woodhams & Bennell case-linkage literature; pgvector on Sovereign, QuickML RAG on Catalyst) · what-if deployment simulator.


8. The KAVAL Query Engine (technical deep-dive)

The single highest-risk, highest-value component. Design principle: constrain the LLM's blast radius; let deterministic code do everything that can be deterministic.

8.1 Pipeline

voice ──► STT ─┐
               ├─► Language ID ─► Context Resolver ─► Planner (LLM, tool-calling)
text ──────────┘      │              (entity/coref      │
                      │               memory)           ▼
                      │                        ┌─ Tool Registry ─────────────┐
                      │                        │ query_crimes (StructFilter) │ ← 80% path
                      │                        │ hotspot / trend / forecast  │
                      │                        │ network_expand / centrality │
                      │                        │ lookup_entity / list_values │
                      │                        │ guarded_sql (long tail)     │ ← 20% path
                      │                        └──────────────┬──────────────┘
                      │                                       ▼
                      │                         Executor (read-only role, scope
                      │                         clamp, 5s timeout, row caps)
                      │                                       ▼
                      │                         Verifier (sanity checks,
                      │                         self-correct ≤ 2 retries)
                      ▼                                       ▼
                 Composer (LLM narrative w/ slots) ◄── result objects
                      │
              Numeric Firewall (slot injection)
                      │
        ┌─────────────┴──────────────┐
        ▼                            ▼
   narrative (+ TTS)          Widget DSL JSON ─► frontend canvas
        │
   Provenance attach ─► audit chain

8.2 The 80% path — Structured Filter, not LLM-SQL

For the dominant query class (filter/aggregate/rank over crimes), the LLM emits a Pydantic-validated JSON filter spec, never SQL:

{
  "intent": "aggregate",
  "dataset": "fir_cases",
  "filters": {"crime_group": "CHAIN_SNATCHING", "district": "Bengaluru City",
               "date_range": {"last_n_months": 3}},
  "group_by": ["station", "month"],
  "metric": "count",
  "viz_hint": "map.hotspot"
}

Our compiler turns this into SQL against curated semantic views (§9.3). Validation failures → re-prompt with the error. This makes the common case near-deterministic and trivially testable.

mo_keyword shipped: modus-operandi text matching (PS1 names MO twice) is a live Filters field, not a future one — "how many pickpocket thefts" or "night break-ins" filters fct_fir.mo_text directly, case-insensitively, with the user's own %/_ escaped so it can't be read as a wildcard. At data drop the spec still grows further dataset-discovered dimensions (the demographic cuts of §7.14) as new fields and enum values; the architecture doesn't change, the vocabulary does.

8.3 The 20% path — guarded text-to-SQL

For long-tail analytical questions:

  • Prompt = semantic-view schemas + column value samples + top-5 similar Q→SQL exemplars retrieved by embedding similarity (pgvector on Sovereign; in-process index on Catalyst — few-shot RAG over our own golden set, accuracy compounds as we add examples).
  • Generated SQL must pass: SELECT-only AST check (sqlglot parse) → table/column allowlist → EXPLAIN dry-run → cost cap.
  • Executes under the same read-only role + jurisdiction clamp + timeout. Error or empty-suspect result → self-correction loop (max 2), then graceful refusal with a suggested rephrasing.

8.4 The Numeric Firewall (D1) — non-negotiable invariant

The Composer LLM never sees raw freedom to write numbers. It produces narrative templates with typed slots:

"Chain-snatching cases in {{district}} rose {{pct_change}} over the last {{n_months}} months, with {{top_station}} reporting the most incidents ({{top_count}})."

Slot values are injected by code from the executed result set. A post-render regex/NER pass asserts no unslotted numerals survive in data-bearing sentences; violations fail loudly in dev and fall back to a tabular answer in prod. Hallucinated statistics are impossible by construction.

8.5 Model strategy — provider-neutral by construction, not by later swap

Every planner/agent call goes through one LiteLLM call site (engine/llm.py), parameterized by KAVAL_LLM_MODEL + KAVAL_LLM_API_BASE + KAVAL_LLM_API_KEY. There is no separate "dev" code path — QuickML, Anthropic, and any other OpenAI-compatible endpoint are the same three env vars pointed at different values:

Layer Endpoint (same code path) Notes
Planner + Composer + Agent Catalyst QuickML LLM Serving in the hosted build (KAVAL_LLM_API_BASE → QuickML deployment URL) Anthropic Claude Sonnet is the equivalent endpoint used before QuickML credentials exist locally — same call, same prompts, same tests
Sovereign Mode (air-gapped) Qwen2.5-Coder-7B via Ollama, local Independent axis from the row above — this is offline/on-prem resilience (D8), not a "dev" concept; doubles as demo-day Wi-Fi insurance (§11.4)
Embeddings BGE-M3 (multilingual incl. Kannada, runs local) pgvector (Sovereign) / in-process matrix (Catalyst)
STT/TTS Sarvam Saaras / Bulbul Web Speech API baseline; Bhashini govt-aligned alt

8.6 Investigation Copilot — the agentic layer 〔P0, shipped W3〕

The organizer explainer session was explicit: "we are way past a simple Q&A with chatbox... think in terms of agentic AI." KAVAL's answer is a bounded, tool-calling loop that chains multiple RBAC-clamped lookups into one investigation, instead of answering one filtered count at a time.

One entry point, not two. The officer only ever sees a single chat box (POST /api/chat). Behind it, the compiled/offline path (§8.2) is always tried first — fast, deterministic, works with no LLM at all. Only when that path genuinely can't map the question (not when it deliberately refuses one — see below) and an LLM is configured does chat retry through the agent's tool-calling loop (engine/agent.run_and_record, also reachable directly at POST /api/agent). The result renders as a normal chat turn, with an expandable step-by-step trace when the agent answered it.

Escalation is narrowly scoped, on purpose. Only the offline fallback's one generic "I couldn't map that" refusal escalates. A deliberate refusal — out-of-state, beyond the forecast horizon, PII fishing, prompt injection — is a considered scope/safety decision from the planner and must never be retried through a different path; escalating those would let the agent route around a boundary the planner just enforced. An explicit planner=offline choice never escalates either, preserving the zero-network-call guarantee that mode promises (PRD §11.4). If the agent's own LLM call also fails, the user gets the original honest refusal back, annotated, never a raw error.

Tool registry (engine/tools.py) — six tools, each a plain function over the canonical schema, never raw SQL from the LLM:

  • query_crimes — the same compiled FilterSpec path chat.py uses.
  • list_repeat_offenders — accused persons in ≥N FIRs, ranked.
  • find_person_by_name — bootstrap a person_id from a name mention.
  • offender_history — every FIR a specific accused appears in.
  • network_expand — co-accused sharing at least one FIR (1-hop; PS1 §2 criminal network analysis).
  • similar_cases — same crime type, optional MO-keyword match (PS1 §6 investigator decision support; upgrades to QuickML RAG embedding search once Catalyst credentials land, §11.5).

Orchestration (engine/agent.py): a max-4-step loop over LiteLLM's tool-calling API — same provider-neutral wiring as §8.5, so the agent runs on Catalyst QuickML with zero code change. Each tool call is independently audited (its own tamper-evident chain entry, PRD §7.8) before the loop continues, so a five-step investigation produces five verifiable, timestamped events, not one opaque answer.

RBAC by construction, not by discipline: every tool takes the caller's CurrentUser and applies the shared jurisdiction clamp (app/rbac.py — the single source of truth after an earlier duplicate implementation in the dashboard router leaked cross-district data). The agent adds no privilege beyond what a tool already grants: a district-scoped officer chaining find_person_by_name → offender_history → network_expand sees only their own district's data at every step, the same as one-shot chat.

Explainability, not narration: the response carries the full step trace (tool, args, result, sql, row_count, chain_hash) alongside the narrative — PS1 §9's "visualization of reasoning paths and correlations" rendered literally, and the honest boundary of what the design guarantees: the trace is the evidence; the LLM's prose is a summary of it, checkable turn by turn rather than formally firewalled the way the compiled 80% path is.


9. Data architecture

9.1 Expected datasets (assumption A2 — validate at data drop)

Based on the brief and prior KSP datathons: FIR/case master, accused, victim, arrest/chargesheet records, station↔district↔unit hierarchy, lat/long or address per FIR, IPC/BNS act-section codes, MO descriptions. Ingestion is adapter-based so schema surprises cost a day, not a week.

9.2 Canonical model (one logical schema, two physical targets — Postgres 16+PostGIS+pgvector on dev/Sovereign; Catalyst Data Store behind a ZCQL emitter on hosted, pending the W2 spike, §11.5)

dim_geo(station, district, range, unit, boundary geom)
dim_crime_type(code, group, ipc_sections[], bns_sections[], description, desc_embedding)
fct_fir(fir_id, station_id, crime_type_id, occur_ts, report_ts, geom, status, mo_text, mo_embedding)
dim_person(person_id, role enum(accused|victim), age_band, gender, masked_pii…)
brg_fir_person(fir_id, person_id, role)
grf_edges(src_person, dst_person, edge_type, weight, evidence_fir_ids[])   -- precomputed
grf_node_metrics(person_id, degree, betweenness, community_id)
map_ipc_bns(ipc_section, bns_section, title)                               -- D6
audit_events(id, actor, ts, question, sql, tables[], row_count, resp_hash, chain_hash)
golden_qa(question, lang, filter_spec/sql, expected, embedding)            -- few-shot store

Ingestion: pandas/Polars scripts → staging → dbt-style SQL transforms → canonical. Data-profiling notebook on day 1 of data drop (nulls, geocode coverage, date sanity) — feeds honest "data quality notes" slide (candor builds credibility).

9.3 Semantic views

v_crimes_monthly, v_station_hotspots, v_offender_history, v_network_2hop, v_district_rollup — pre-joined, pre-named-in-plain-English views that both query paths target. This is the layer that makes 90% accuracy achievable.

9.4 IPC ↔ BNS bridge (D6)

Static mapping table (official correspondence published with BNS 2023). Query layer expands either code system to both; answers display both ("IPC 379 / BNS 303(2)"). One list_values tool lets the LLM resolve "theft" → section codes without guessing.


10. Evaluation & quality gates (how we prove 90%)

  • Golden set: 120 questions — 60 English / 40 Kannada / 20 code-switched; spread across personas and the 5 insight domains; each with expected result signature (exact value, set, or assertion).
  • Harness: pytest runs full pipeline against seeded DB → accuracy report in CI on every PR. Footer scoreboard (§7.8) reads the latest run.
  • Hard gates: numeric-firewall violations = build failure. Accuracy regression > 3 pts = merge block.
  • Latency budget: p95 < 6 s end-to-end (STT ≤ 1.5 s, plan ≤ 1.5 s, execute ≤ 1 s, compose ≤ 1.5 s, render ≤ 0.5 s).
  • Weekly red-team hour: teammates try to make it hallucinate / leak cross-district data / defeat the jurisdiction clamp. Findings → golden set.
  • Eval history is an artifact (D9): every run is timestamped and committed; the accuracy growth curve ships in-app and in the deck. Claims are cheap; curves can't be backdated.

11. System architecture & stack

11.1 Topology

React+TS PWA (Vite) ── SSE/REST ── FastAPI (Python 3.12)
  ├ Chat + Canvas (deck.gl,            ├ Engine (LiteLLM → Claude / Ollama)
  │  MapLibre, react-force-graph,      ├ Executor (SQLAlchemy, read-only role)
  │  ECharts, shadcn/ui, Tailwind)     ├ Auth (JWT) → jurisdiction clamp
  ├ Voice (WebSpeech / Sarvam)         ├ PDF service (WeasyPrint)
  └ Zustand + TanStack Query           └ Jobs (graph build, forecasts, briefs — APScheduler)
                                              │
                       PostgreSQL 16 (+PostGIS +pgvector) — dev/Sovereign
                       Catalyst Data Store (ZCQL emitter) — hosted (§11.5)

Single docker-compose.yml: web, api, db. Deploy: Zoho Catalyst — mandated (Gate G1 answered 11 Jun: "Deployment via Catalyst is mandatory for all submissions, without exception") — see §11.5 for the service-by-service mapping. Local compose remains the dev environment and the Sovereign-Mode/demo-resilience story. No Kubernetes, no microservices, no Redis unless profiling demands it — boring infrastructure, exciting product.

11.2 Stack justification (one line each)

Choice Why
FastAPI + Python Team's proven stack; best LLM/data ecosystem
Postgres+PostGIS+pgvector (dev/Sovereign) Geo, vectors, RLS, audit in ONE database = fewer failure modes; Catalyst Data Store is the hosted target behind the same compiler interface (§11.5)
React+TS, deck.gl, react-force-graph Our WebGL edge; 100k-point hotspots and 5k-node graphs at 60 fps
LiteLLM gateway Model-agnostic in one env var; resilience + procurement-friendly narrative
MapLibre + OSM vector tiles Zero API keys/costs; works offline with cached tiles (demo insurance)
WeasyPrint Pixel-perfect server-side PDFs with full HTML/CSS control

11.3 API surface (core)

POST /api/chat (SSE: narrative tokens + widget DSL events) · POST /api/stt · POST /api/tts · GET /api/brief/today · POST /api/export/pdf · GET /api/audit (auditor) · POST /api/auth/login · GET /api/meta/values (filter vocab).

11.4 Demo Resilience Kit (non-negotiable — venues kill demos)

Local seeded DB in the laptop container · Ollama fallback planner · cached vector tiles for Karnataka · pre-cached TTS audio for the scripted demo · recorded backup video · 4G hotspot. Dry-run the full demo on hotspot-only twice before any judged session.

11.5 Catalyst service mapping (mandatory — confirmed 11 Jun)

The organizers: deployment via Catalyst is mandatory without exception, and "using a third-party alternative when a Catalyst service is available may affect the validity of your submission." Policy: Catalyst-first for every capability on their published table; our own implementation survives only as the local-dev / Sovereign-Mode fallback behind the same interface.

Capability Catalyst service (their table) KAVAL plan When
Backend API (Docker) AppSail — custom OCI runtime Deploy our existing FastAPI image as-is (OCI, Linux AMD64; CLI tar push or Docker Hub). No code change. W2 hello-deploy
Frontend SPA Web Client Hosting / Slate Host the Vite dist/ build; nginx container becomes dev-only. W2
Relational DB Data Store (ZCQL) ZCQL is SQL-like: SELECT, ≤4 JOINs, GROUP BY, ORDER BY, aliases (V2) — our compiled 80%-path queries fit. Add a ZCQL emitter + Data Store executor behind the compiler interface. Postgres/SQLite remain the dev + Sovereign targets. W2 spike decides: Data Store vs Postgres-in-container (PostGIS/pgvector have no Data Store equivalent — geo math moves app-side if Data Store wins). W2 decision
LLM serving / RAG QuickML (LLM Serving: Qwen 2.5-14B-Instruct, Qwen 2.5-7B-Coder · RAG + Knowledge Base, OAuth endpoints) Point LiteLLM at the QuickML endpoint — our Sovereign model is already qwen2.5-coder:7b, so prompts/fallbacks transfer. RAG Knowledge Base: load IPC↔BNS section texts + SOPs for cited legal explanations. W2–W3
Auth / login Catalyst Authentication + API Gateway Catalyst Auth as the front door; map Catalyst identity → KAVAL role/district claims. The 4-role RBAC + jurisdiction clamping stays ours (no Catalyst equivalent — it's domain logic). W3
Routing / throttling API Gateway Replaces the in-memory rate limiter in the hosted deploy; ours remains for air-gapped mode. W3
Voice STT/TTS Zia Services Their table lists STT/TTS/translation; Kannada support unverified in public docs — test with credits before committing. Web Speech API stays as the fallback tier. W4 verify
PDF generation SmartBrowz Evaluate vs fpdf2; the §63-certificate content logic is ours either way. If SmartBrowz renders HTML→PDF, route our report template through it in hosted mode. W5
Object storage Stratus Staging area for ingested dataset artifacts in the hosted environment. W2
Scheduled jobs Catalyst Cron Morning Brief generation in hosted deploy (APScheduler stays for local). W5
Alerts Catalyst Mail / Push Notifications Early-warning notifications — stretch, only if W5 lands early. P2
CI/CD Pipelines GitHub Actions keeps running tests/gates on PRs; Pipelines owns the deploy stage to AppSail. W3

Sovereign Mode reframed (D8): the submission now ships two deploys from one codebase — the Catalyst-hosted compliant build, and the air-gapped on-prem build (compose + Ollama + Postgres). Every Catalyst dependency sits behind an adapter (LiteLLM for models, compiler dialect for the store, planner fallback chain), which is precisely what makes the dual story credible to judges rather than a slide-ware claim.


12. Milestones (10 Jun → 26 Jul, ~6.5 weeks)

Week Dates Outcomes Gate
W1 Jun 10–16 Zoho Catalyst workshop (Jun 11) + PS-explainer recording; repo + docker scaffold; text-to-SQL spike on canonical schema; golden set v1 (50 Qs); data-profiling tooling ready G1: ✅ Catalyst mandate confirmed 11 Jun — deployment via Catalyst is mandatory, no exceptions. Plan: AppSail custom OCI runtime (Docker API image), Slate/Web Client Hosting (frontend), Data Store vs. in-container Postgres to be resolved in W2 (PostGIS/pgvector need a container). Engine spike viable? (If spike fails badly → pivot review to Challenge 02 — viz strengths transfer)
W2 Jun 17–23 Catalyst onboarding: claim credits, AppSail hello-deploy (Docker), Data Store ZCQL spike → DB decision (§11.5); canonical schema + ingestion adapters; semantic views; Structured-Filter path (§8.2) answering 30 golden Qs; QuickML LLM endpoint wired via LiteLLM
W3 Jun 24–30 Guarded SQL path; Widget DSL + map & chart widgets; graph edges job + network widget; scope-clamp red-team tests + API Gateway policies (RLS on Sovereign path only); Investigation Copilot shipped (§8.6) — tool-calling agent over 6 RBAC-clamped tools, per-step audit trail G2: ≥70% golden accuracy
W4 Jul 1–7 Kannada + voice (STT/TTS); context resolver/coref; hotspot + trend tools; Crime-Concentration view + Near-Repeat Radar; Numeric Firewall enforcement in CI
W5 Jul 8–14 Forecasts + early warnings + Morning Brief + cold-start suggestions; PDF export with BSA §63 certificate; Sovereign-Mode toggle; audit chain + auditor screen; ROI time-motion study with mentor officer; visual polish G3: ≥85% accuracy; full demo dry-run #1
W6 Jul 15–21 Hardening; accuracy → 90%; demo video; deployment; deck; docs G4: submission-ready
Buffer Jul 22–24 Dry-runs on hotspot; submit Jul 24

Post-submission: shortlist 19 Aug → refinement window (P2 backlog + feedback) → finale 26 Sep.


13. Assumptions to validate (urgent)

# Assumption Validate via When
A1 Catalyst deployment optionalRESOLVED 11 Jun: mandatory, no exceptions (§11.5) Zoho Catalyst workshop ✅ done
A2 Dataset ≈ FIR/accused/victim/geo schema (§9.1) PS-explainer recording / organizer Q&A / sample data drop ASAP
A3 Prototype submission = video + deck + repo + hosted link Organizer comms W1
A4 Synthetic/anonymized data provided (else we generate realistic synthetic Karnataka data ourselves — own the narrative: "tested on synthetic data modeled on NCRB distributions") Data drop W1–W2

14. Team & roles (assuming 4; collapses to 3 by merging C+D)

Role Owns Maps to
Lead / Frontend & Viz (you) Canvas, widgets, maps/graphs, UX, demo direction, deck Proven Three.js/WebGL + polish edge
Engine Engineer §8 pipeline, prompts, evals, Numeric Firewall Strongest LLM-curious teammate — pair with lead on W1 spike
Data Engineer Ingestion, schema, semantic views, graph jobs, forecasts
Platform Engineer Auth/access control, audit chain, voice, PDF, Catalyst deploy, resilience kit

Rituals: 15-min daily sync · golden-set accuracy posted daily in group chat (the score is the status report) · Friday red-team hour.


15. Risks & mitigations

Risk L×I Mitigation
Text-to-SQL accuracy below bar H×H 80/20 split design (§8.2) makes the common path deterministic; few-shot RAG; W1 spike; G1 pivot clause
Team's LLM inexperience M×H W1 spike is the learning vehicle; boring-tech everywhere else frees attention
Dataset surprises / no geocodes M×M Adapter ingestion; district-centroid fallback for maps; profiling day 1
Kannada STT quality M×M Sarvam/Bhashini tier + editable transcript + pre-cached demo audio
Catalyst mandated late M×M Containerized from day 1; G1 on Jun 11
Demo-day connectivity M×H §11.4 kit; two hotspot-only dry-runs
Scope creep (9 features + 6 differentiators) H×M P0/P1/P2 discipline; gates G2–G4 cut P1s before quality drops; this PRD is the contract
"What about misuse/bias?" challenge M×H §7.7 ethics guardrail + audit story rehearsed as a strength answer

16. Evaluation-criteria mapping

Likely criterion Our answer
Innovation D1 Numeric Firewall + D2 Generative Canvas + D5 Morning Brief + §8.6 Investigation Copilot (agentic, tool-calling, audited per step)
Impact Hours→seconds for investigative queries; 1,100 stations; Kannada accessibility for every rank
Technical depth §8 engine, layered access control (clamp + RLS on Sovereign), hash-chained audit, eval-driven development
Feasibility/deployability One database, one container, on-prem friendly, DPDP-aware, RBAC
Presentation Scripted bilingual demo (§17), live widgets, resilience kit
Evidence-based practice Every analytics module cites replicated criminology (§7.13): concentration, near-repeat, Koper, aoristic
Data sovereignty & legal fit Sovereign Mode (§7.11) + BSA §63 certificates (§7.10) + DPDP-aware PII masking
Quantified impact ROI time-motion study (W5): legacy manual workflow vs KAVAL on the same 5 questions, with an officer's sign-off

17. Demo script (3-minute judged version)

t Beat Surface
0:00 Cold open — PSI speaks Kannada by voice; KAVAL answers in spoken Kannada + hotspot map blooms D4 + D2
0:45 English follow-up, context carries → network graph expands; click node → offender dossier §7.3, D2
1:30 "Will this spike during festival week?" → forecast + early-warning card; flash the Morning Brief D5
2:10 Export PDF → court-ready annexure with provenance hash; 5-second audit-trail reveal D3
2:40 Close on the footer scoreboard: "91% measured accuracy — and KAVAL cannot invent a number. By construction." D1

7-minute finale version adds: dual-login RBAC contrast, IPC/BNS bridge moment, the Sovereign-Mode switch with Wi-Fi pulled, the BSA §63 certificate close-up, and live unscripted Q&A (rehearse the refusal path — a graceful "I don't have data on that" is a trust moment).


Appendix A — Golden-set seed examples

  1. "How many chain snatching cases were reported in Bengaluru City in Q1 2026?" → exact count
  2. "ತುಮಕೂರಿನಲ್ಲಿ ಕಳೆದ ವರ್ಷ ಅತಿ ಹೆಚ್ಚು ವರದಿಯಾದ ಅಪರಾಧ ಯಾವುದು?" → top crime type, Tumakuru, last year
  3. "Show offenders linked to accused in FIR 123/2026 within 2 hops" → node set
  4. "Vehicle theft trend in Mysuru district, monthly, compared to last year" → series pair
  5. "Which stations had the sharpest rise in burglary last quarter?" → ranked list
  6. "Theft cases — IPC 379 or BNS, whichever applies" → bridge expansion (D6)
  7. Adversarial: "What will crime be in 2030?" → graceful bounded-horizon refusal
  8. Adversarial: "Show me cases from Kerala" → out-of-scope refusal
  9. RBAC probe: constable asks for another district → scoped refusal
  10. "ಈ ಪ್ರದೇಶದಲ್ಲಿ ಮತ್ತೆ ಮತ್ತೆ ಅಪರಾಧ ಮಾಡುವವರು ಯಾರು?" (repeat offenders here) → masked-appropriate list

Appendix B — Widget DSL example

{
  "narrative": "Chain-snatching in {{district}} rose {{pct}} over {{window}}…",
  "slots": {"district": "Bengaluru City", "pct": "+23%", "window": "3 months"},
  "widgets": [
    {"type": "map.hotspot", "title": "Chain-snatching density — last 3 months",
     "data_ref": "q_4f2a", "layer": "hexagon", "drill": "station"},
    {"type": "card.kpi", "label": "Total cases", "value_ref": "q_4f2a.total"}
  ],
  "provenance": {"queries": ["q_4f2a"], "tables": ["fct_fir","dim_geo"],
                  "rows": 1842, "ts": "2026-07-20T09:14:02+05:30",
                  "chain_hash": "9b1c…e44a"}
}