A citation-grounded chatbot over Census 2011 India state reports — every factual claim traces back to an exact PDF page, checksum, and verbatim quote, or the system refuses rather than guesses.
Ask it about literacy rates, population, sex ratios, and district rankings across Karnataka, Odisha, and Madhya Pradesh — or upload your own report. It looks things up, summarizes, compares, ranks, checks internal consistency, and builds charts/tables — and it says "I don't know" instead of making something up.
Watch or download the demo (MP4, about 3½ minutes) · Transcript · Captions
Recorded from the running application on September 7, 2026, with synthetic English narration.
Covers architecture, healthy services, a cited lookup, a follow-up, chart generation, and an
out-of-scope refusal. Model waiting time is edited out and selected frames are held for explanation.
The follow-up comparison failed citation validation in this run; the video preserves and explains
that result. The lookup and chart succeeded. make check passed all 305 tests.
That specific comparison-refusal's root cause (resolve_query could rewrite a follow-up to name
every target while losing the comparative framing itself, so nothing was ever computed to cite) is
now fixed and verified live — see FAILURE_ANALYSIS.md. The video is kept unedited as an honest
record of that session rather than re-recorded.
The video is included in the repository. If GitHub shows a file page instead of a player, use View raw or Download raw file to watch it.
A Next.js + TypeScript client built around one idea: every answer should be checkable in one click.
More: district ranking with the winner highlighted · finished research brief · dark mode · mobile · empty state
- Claude-style sessions: a searchable, date-grouped chat history, rename and delete, and
shareable URLs (
/c/<session_id>). Switch chats while one is still answering. - Answers that show their work: numbered citation chips on every claim, arithmetic shown as computed in code, a clear badge when the agent declines instead of guessing, a per-answer execution trace, and one-click follow-ups ("How does that compare with Odisha?") that exercise the validated conversation memory.
- Hardened by default: the browser reaches FastAPI only through an allowlisted same-origin proxy; the container is read-only, non-root, and has no credentials.
- Dark mode, mobile layout, keyboard shortcuts (⇧⌘O new chat), and live service health.
- Demo video
- The interface
- Measured trust
- What makes this different
- Architecture
- How a request actually flows
- The provenance chain
- What it can do
- Quick start
- Data setup and first initialization
- Run and review
- Safety and behavior
- Verification and evaluation
- API and traces
- Troubleshooting
- Production modes
- Repository layout
- Further reading
A benchmark of 26 questions with answers verified by hand against the source tables, each asked
twice through the live API and scored independently of the agent's own validation code
(backend/app/trust.py, cases in
evals/trust_benchmark.json). Ground truth is facts, not loose
numbers: for example (Odisha, sex ratio, 2011, urban, all persons, 932). Every number must also be
found on its region's row, in the right column, of the quote it cites.
| Check (52 live answers) | Result |
|---|---|
| Wrong answers shown | 0 |
| Numbers not on their region's row and column in their cited quote | 0 of 236 claims checked |
| Correct refusals (GDP, unemployment, cricket, a 2021 census) | 8 / 8 |
| Correct answers | 37 / 44 |
| Unnecessary refusals (the safe failure) | 4 |
| Operational errors | 3 |
| Median latency per question | 20 s |
The cases are built to catch misattribution, not just recall:
- a 2001 figure next to the 2011 one;
- urban versus rural versus total columns;
- a comparison where Madhya Pradesh's child sex ratio equals its urban sex ratio;
- the Scheduled Tribes and Castes tables beside the state figures;
- follow-ups that name neither region nor metric;
- a complete district ranking the agent is known to struggle with.
The misses are honest limitations, not wrong answers. In that run the agent declined questions about 2001, failed closed on district literacy rankings, and three cases gave different outcomes on their two runs. District rankings now read from structured tables: re-running the four ranking cases twice passed 8 of 8, with a median of 5.8 seconds. All are listed in FAILURE_ANALYSIS.md. A replay of real answers with deliberate corruptions (swapped regions, a neighbouring column, a wrong year) shows the scorer catches each one; the previous version passed all four.
This is one run on 2026-09-26 with gemini-2.5-flash. Reproduce it with make trust-benchmark
(billable) and see it in the app at /trust.
A lot of "chat with your documents" demos silently give up citation accuracy the moment the answer requires a computed comparison, a chart, or a district ranking. This one doesn't:
| Typical RAG chatbot | Census Insight Agent | |
|---|---|---|
| Citations | Page number, sometimes | Exact verbatim quote + physical PDF page + SHA-256 checksum, verified by application code — never trusted from the model |
| Comparisons ("how does X compare to Y") | Model does the subtraction in prose | Arithmetic is app-owned; the model never asserts a computed number that wasn't independently calculated |
| Charts/tables | Pre-baked template, or the model fabricates numbers | Model proposes which data; application code hydrates and verifies every cell against a trusted table row before any code executes |
| Code execution | Often absent, or runs unsandboxed | Generated Python passes an AST policy gate, then runs in a non-root, network-disabled, read-only container with resource limits |
| "I don't know" | Rare — models like to please | A structural outcome: insufficient evidence, a failed provenance check, or an unverifiable table cell all produce an explicit refusal |
| Memory | Conversation replayed as prose | Structured, validated claim history — a follow-up re-verifies against live evidence, it never trusts what a past turn said |
Four services, four trust boundaries. The frontend never touches Qdrant, credentials, or the executor; the executor never touches the network, Qdrant, or credentials.
flowchart LR
subgraph P["🖥️ Presentation plane"]
U((User)) --> UI["Next.js :3000\nallowlisted proxy"]
end
subgraph A["🧠 Agent / orchestration plane"]
UI --> API["FastAPI :8000"]
API --> LG["LangGraph"]
LG --> V["Vertex Gemini via ADC"]
LG --> S["Runtime skills\n(skills/*.md)"]
LG --> M["SQLite checkpointer"]
end
subgraph E["📚 Source / evidence plane"]
LG --> R["Hybrid retrieval"]
R --> VE["Vertex embeddings"]
R --> BM["Local BM25"]
R --> Q["Qdrant RRF :6333"]
PDF[("Authoritative PDFs")] --> Q
end
subgraph X["🔒 Isolated execution plane"]
LG --> FQ["Filesystem queue"]
FQ --> EX["Network-disabled executor\n(non-root, read-only, no creds)"]
EX --> AR["Validated PNG / CSV / manifest"]
AR --> API
end
classDef isolated fill:#fef2f2,stroke:#dc2626,color:#7f1d1d
class X,EX isolated
This is what happens between you hitting Enter and an answer appearing — classification, per-target retrieval, evidence assessment, and a citation-validation gate that can trigger one bounded repair before it ever refuses.
sequenceDiagram
autonumber
actor You
participant UI as Next.js
participant API as FastAPI
participant Graph as LangGraph agent
participant Qdrant
participant Gemini as Vertex Gemini
participant Exec as Isolated executor
You->>UI: "Compare literacy rates, Karnataka vs Odisha"
UI->>API: POST /chat/stream
API->>Graph: run(session_id, message)
Graph-->>UI: SSE progress event as each node starts
Graph->>Gemini: classify task + resolve standalone query
Gemini-->>Graph: task_type=comparison, targets=[Karnataka, Odisha]
par per-target retrieval
Graph->>Qdrant: dense + BM25 search — Karnataka
Graph->>Qdrant: dense + BM25 search — Odisha
end
Qdrant-->>Graph: RRF-fused candidates, reserved per target
Graph->>Gemini: assess evidence relevance
Gemini-->>Graph: relevance + entity/metric/unit match flags
alt evidence insufficient
Graph-->>API: citation-safe refusal + limitation
else artifact requested (chart/table)
Graph->>Gemini: propose minimal dataset (labels + values only)
Graph->>Graph: hydrate + verify every cell against trusted evidence
Graph->>Exec: generated Python + input.json (no credentials, no network)
Exec-->>Graph: chart.png / table.csv + source-manifest.json
Graph-->>API: validated artifact + citations
else lookup / comparison / summary / ranking
Graph->>Gemini: synthesize structured claims
Graph->>Graph: validate citations (one repair if invalid, else refuse)
Graph-->>API: cited answer
end
API-->>UI: result event (response + trace_id)
UI-->>You: cited answer · PDF page viewer · trace
Every number the agent ever states can be walked backward to a specific, checksummed page. Nothing skips a link in this chain — a broken link anywhere produces a refusal, not a guess.
flowchart LR
PDF["📄 Original PDF\nchecksum + page count"] --> MAP["Page mapping\nMarkdown-preferred, PDF-anchored"]
MAP --> CHUNK["Page-bounded chunk\nnever spans pages"]
CHUNK --> VEC["Dense + BM25 vectors"]
VEC --> QD["Qdrant point\nchunk + checksum + page"]
QD --> RET["Retrieved candidate"]
RET --> ASSESS["LLM relevance assessment"]
ASSESS --> CLAIM["Structured claim\napp-built, not model-written"]
CLAIM --> CITE["✅ Citation\nexact quote + physical PDF page"]
CITE --> OUT["Answer or artifact"]
style PDF fill:#fee2e2,stroke:#991b1b,color:#7f1d1d
style CITE fill:#dcfce7,stroke:#166534,color:#14532d
style OUT fill:#dbeafe,stroke:#1e40af,color:#1e3a8a
If any link breaks — an excluded visual page, a non-verbatim snippet, an unverifiable table cell — the chain stops and the system says so instead of silently continuing.
| You ask | Task type | What actually happens |
|---|---|---|
| "What was Karnataka's literacy rate in 2011?" | lookup |
Retrieves, cites the exact table cell |
| "How does that compare with Odisha?" | comparison |
Retrieves both sides independently, computes a validated arithmetic difference |
| "Summarize the key findings for Odisha" | summary |
Page-distributed representative retrieval, not just the top-k hits |
| "Which district had the highest gender ratio?" | artifact_table (ranking) |
Full-table scan; the winner is computed by application code, never asserted by the model |
| "Create a bar chart comparing Karnataka and Odisha literacy rates" | artifact_chart |
Model proposes labels/values only; app hydrates and verifies every cell before the isolated executor renders a PNG |
| "Build a table of the urban vs rural breakdown" | artifact_table |
Same hydration/verification path, output is table.csv + table.md |
| "Check whether the population totals in this table are consistent" | inconsistency_analysis |
Deterministic arithmetic tool calls, checked against the same citation-validation gate |
| "Which source pages support those values?" | source_support |
Rehydrates prior validated claims against current Qdrant state — never trusts conversation prose |
| "What was France's unemployment rate in 2011?" | out_of_scope |
Refuses cleanly — no hallucinated cross-corpus answer |
| Upload a PDF, then ask about it | any of the above | Indexed through the same page-bounded, checksum-bound pipeline; the agent sees the live document catalog |
| Deep research: "Gender balance and literacy across Karnataka and Odisha" | research brief | A planner writes 3–5 questions; each runs as a separate validated agent turn; the brief contains only validated claims |
The submission ZIP includes the three original PDFs and supplied Markdown. A Git checkout requires
copying those files into data/source/ as described below.
Credentials are never included. Vertex inference and first-time embeddings are billable.
After completing the ADC and .env configuration below, run:
sh scripts/setup.sh --allow-paid-callsThis builds images, checks source-page coverage, initializes an empty index, prepares Linux queue
permissions, and waits for healthy services. It preserves an existing valid collection. Open
http://localhost:3000. Subsequent starts need only docker compose up -d --wait.
The one-shot queue-init container exits successfully; the four application services stay running.
Windows users should use WSL2 with Docker Desktop integration enabled.
For the complete application, install Docker Engine/Desktop with Compose v2 and the Google Cloud CLI.
For host-side development, also install Python 3.12 and uv. The
backend requires a Vertex-enabled GCP project and Application Default Credentials (ADC); API-key
authentication is unsupported.
gcloud auth application-default login
gcloud auth application-default set-quota-project YOUR_PROJECT_ID
cp .env.example .envIn .env, set GOOGLE_CLOUD_PROJECT, the Gemini model names, and GOOGLE_CREDENTIALS_HOST_PATH to
the absolute host ADC JSON path. Keep GOOGLE_APPLICATION_CREDENTIALS=/var/secrets/google/adc.json;
it is the intentional container-local mount. Never copy credentials into the repository.
Why Vertex/ADC instead of a free-credit API key? Short version: project-scoped GCP identity without distributing a secret, at the cost of a slower reviewer setup. This tradeoff is made explicitly, not overlooked — see DESIGN.md § Tradeoffs.
The assignment corpus is ignored because its PDFs and Markdown total about 115 MB. Copy the supplied files into:
data/source/pdf/
PC11_PCA_Data_Highlights_Karnataka.pdf
PC11_PCA_Data_Highlights_Odisha.pdf
PCA Data Highlights MP.pdf
data/source/markdown/
PC11_PCA_Data_Highlights_Karnataka.md
PC11_PCA_Data_Highlights_Odisha.md
PCA Data Highlights MP.md
Pairing and reviewed page-coverage decisions are checked in under data/manifests/. Run the free dry
run first; it makes no Gemini calls and no Qdrant writes:
docker compose run --rm --no-deps --build backend uv run --frozen python -m backend.app.ingestion.cli ingest --source-dir /app/data/source --dry-run --report-output /app/data/processed/dry-run-report.json
python3 -m json.tool data/processed/dry-run-report.jsonAfter reviewing zero failures and the estimated inputs, initialize an empty collection intentionally. The initializer stops unless the paid-call flag is present, skips a valid 2,058-point collection, and never deletes a non-matching collection:
docker compose up -d qdrant
docker compose run --rm backend uv run --frozen python scripts/initialize_corpus.py --allow-paid-callsInitial ingestion invokes billable Vertex embeddings. A fresh docker compose up --build starts the
services but cannot answer corpus questions until the assignment files are supplied and
initialization is approved.
Ingestion also writes each document's structured table store. For a collection indexed before table stores existed, build them from the indexed chunks with no model calls:
docker compose run --rm backend uv run --frozen python -m backend.app.tables.cliWithout a store, district rankings fall back to scanning chunks.
docker compose up --buildThe first build and FastEmbed load can take several minutes. Wait for docker compose ps to show
four healthy services.
| Service | URL |
|---|---|
| Next.js chat UI | http://localhost:3000 |
Streamlit UI (optional, docker compose --profile streamlit up) |
http://localhost:8501 |
| FastAPI / OpenAPI docs | http://localhost:8000/docs |
| Qdrant dashboard | http://localhost:6333/dashboard |
Try the example prompts above in one conversation — ask the lookup, then the follow-up comparison, then a chart, to see memory, citations, and artifacts all in the same thread.
Stop without removing indexed data:
docker compose downNever add --volumes unless deleting the local Qdrant collection is intentional.
Original PDFs define identity, checksum, page count, and citation page. Supplied Markdown is
preferred; PyMuPDF4LLM is fallback-only. Chunks never cross pages. Twelve unsafe Karnataka chart/map
pages remain quarantined as excluded_unverified_visual; raw OCR never reaches embeddings, Qdrant,
retrieval, or generation.
Every query gets Vertex dense and local BM25 sparse vectors; Qdrant performs RRF. Retrieval is candidate selection only. LangGraph assesses evidence, builds typed claims, selects exact spans from trusted current-run chunks, validates provenance, and answers or refuses. Comparisons reserve evidence budget for both regions.
AsyncSqliteSaver stores successful conversation state in workspace/checkpoints.sqlite. Run-local
errors do not become durable claims. Runtime instructions are discovered from skills/*.md. Artifact
proposals are deterministically hydrated with trusted evidence before generated Python reaches the
unprivileged, network-disabled executor.
The executor has no cloud credentials, Qdrant, Docker socket, network, or session workspace. Docker isolation is not a hardened hostile multi-tenant sandbox; production use should add a stronger sandbox such as microVM isolation.
A GitHub Actions workflow runs the full offline gate on every push,
with no billable calls: format, lint, type check, all 418 Python tests (including the Postgres and
Redis integration tests), an offline UI smoke test, a Compose trust-boundary audit that covers
both overlays, a git-history secret scan, and a build (never a run) of each service's Docker
image. A second job runs the executor under gVisor. Badge at the top of this file reflects the
current main branch.
Run the same gate locally:
uv sync --frozen --all-extras
make check
make secret-scanWith the services and local corpus running, the canonical full offline gate adds an exhaustive read-only collection scan and dense/sparse compatibility queries. It performs no ingestion, Qdrant writes, Gemini generation, or Vertex embeddings:
make verify-offlineThe Next.js client has its own gate: type check, ESLint, Vitest unit tests (including the proxy allowlist), and a production build. CI runs it as a separate job:
make web-install
make web-checkFor UI development against a running backend, make web-dev serves the client with hot reload
on http://localhost:3000.
Useful individual commands include make format-check, make lint, make typecheck, make test,
make ui-smoke, make qdrant-readonly, docker compose config --quiet, and
docker compose exec frontend python -m frontend.security_check.
workspace/checkpoints.sqlite grows without bound — every LangGraph superstep of every turn
writes a full state snapshot, and nothing prunes old sessions automatically. make prune-checkpoints-dry-run lists sessions inactive more than 30 days (--older-than-days to
change that); make prune-checkpoints actually removes their checkpoints, app_sessions row, and
workspace/sessions/<id>/ trace/artifact directory. Neither command touches Qdrant or requires
Vertex/GCP configuration.
The real-corpus retrieval evaluator uses a paid Vertex query embedding and is manual:
docker compose exec backend uv run --frozen python evals/run_retrieval.py --cases evals/real_corpus_cases.json --output /app/data/processed/retrieval-evaluation-report.jsonThe eleven-case live harness is also manual, requires explicit consent, never retries POST /chat,
and stores a sanitized report:
docker compose exec backend uv run --frozen python scripts/live_evaluation.py --allow-paid-calls --output /app/data/processed/live-evaluation-report.jsonVertex connectivity alone can be checked with the billable make verify-vertex command.
curl -sS -X POST http://localhost:8000/sessions
curl -sS -X POST http://localhost:8000/chat -H 'content-type: application/json' -d '{"session_id":"SESSION_ID","message":"What was Karnataka’s literacy rate in 2011?"}'
curl -sS http://localhost:8000/runs/TRACE_ID/traceTrace IDs are allocated at request start and remain retrievable for typed failures. The UI's
trace view shows allowlisted operational fields only. Files under workspace/ are ignored local state.
| Endpoint | Purpose |
|---|---|
POST /chat/stream |
Server-sent events: progress per graph node, then one result or error |
GET /sessions · GET /sessions/{id}/messages |
Chat history and a session's display transcript |
PATCH /sessions/{id} · DELETE /sessions/{id} |
Rename; delete transcript, checkpoints, traces, and artifacts |
POST /documents/upload · GET /documents/uploads/{job_id} |
Upload a PDF (202 + background job) and poll it |
DELETE /documents/{id} |
Remove an uploaded document and its vectors (bundled corpus is protected) |
POST /research/stream |
Deep research: plan, per-section progress/section_done, then the saved report |
POST /runs/{run_id}/feedback |
Rate an answer up or down; stored beside its trace and sent to Langfuse when enabled |
GET /evaluation/scorecard |
Latest trust benchmark run (local, else the copy in evals/) |
GET /documents/{id}/pages/{n}?highlight=… |
Render a physical PDF page, highlighting the quote where the PDF has a text layer |
| Symptom | Fix |
|---|---|
port is already allocated |
Stop the other process/container using 6333. Do not remove the Qdrant volume. |
pytest: executable file not found |
Run uv run --frozen pytest -q; dev tools are not directly on the production image PATH. |
Unknown session |
Call POST /sessions, then pass its returned ID in valid JSON. |
| UI stays pending | Rebuild frontend/backend and inspect their logs; POST /chat is intentionally not retried. |
| Backend unhealthy | Verify ADC mount, project/location/model values, Qdrant, and executor heartbeat. |
Known limitations: excluded visual-page content, progress streaming by graph step rather than by token (answers are only released after citation validation), scanned or vector-text uploads cannot be indexed without OCR review, quote highlighting is unavailable for the bundled PDFs (they have no text layer), loopback-only services without user authentication, and shared-kernel executor isolation unless the gVisor mode is on. See DESIGN.md and FAILURE_ANALYSIS.md for the full, honest accounting — including inputs where this system degrades, why, and what a real fix looks like. Production modes covers the opt-in scaled, sandboxed, and observed deployments.
The default docker compose up is deliberately a single host: one API process, SQLite state, and
in-process upload indexing. Each production concern is an opt-in mode that keeps every trust
guarantee above, and each was verified live against the real stack. Details and failure
behavior are in DESIGN.md.
| Mode | Turn it on | What it changes | How it was verified |
|---|---|---|---|
| Stateless API replicas | make up-scale (needs POSTGRES_PASSWORD) |
Checkpoints, sessions, and transcripts move to Postgres (AsyncPostgresSaver). The per-session lock becomes a Postgres advisory lock, so a turn on one replica waits for an in-flight turn on another, and a crashed replica cannot wedge a session. |
A question answered on replica :8000 was followed up on :8001, then both containers were replaced with state intact. Integration tests run two replicas against real Postgres; they fail if the lock is removed. |
| Durable upload worker | Same overlay | Uploads are indexed by a Celery worker (Redis broker, append-only persistence) outside the API. Late acknowledgement redelivers a job whose worker died, and transient errors retry with backoff. | A 60-page upload was killed mid-index with SIGKILL; Redis redelivered it and it finished with exactly 60 points, no duplicates. |
| gVisor executor | make up-gvisor (Linux with runsc) |
Generated code runs on gVisor's user-space kernel instead of the host kernel, with every existing control still in place. | Isolation checks, all offline cases, and a full queue handoff pass under runsc. The CI job gvisor-executor repeats this on every push. |
| Langfuse observability | LANGFUSE_BASE_URL, LANGFUSE_PUBLIC_KEY, LANGFUSE_SECRET_KEY |
One trace per turn, grouped by session. Every Gemini call is a generation with token usage and latency. Answers get thumbs up/down scores, and trust scorecards become experiment runs with refusal precision and recall. Prompts and answers stay local unless LANGFUSE_CAPTURE_CONTENT=true. |
Checked against a self-hosted Langfuse 4.46 server, plus unit tests against the exact spans the SDK exports. |
make integration-test runs the Postgres and real-Redis tests against throwaway containers, and
CI runs them on every push. Still open: an authenticated gateway with TLS, a microVM per executor
job (Firecracker or Kata Containers), and object storage so replicas can span hosts.
backend/app/ FastAPI, ingestion, retrieval, LangGraph, lineage
backend/tests/ Unit and integration tests
web/ Next.js + TypeScript client, allowlisted API proxy, Vitest tests
frontend/ Original Streamlit client (optional Compose profile)
executor/ Isolated worker and code policy
skills/ Runtime Markdown skills
data/manifests/ Pairing and reviewed coverage decisions
data/source/ Locally supplied corpus (ignored)
evals/ Evaluation cases and retrieval evaluator
scripts/ Checks, replays, initialization, live harness
workspace/ Ignored checkpoints, traces, queue, artifacts
docs/ Decisions and reviewer preparation
| Document | What's in it |
|---|---|
| DESIGN.md | The why behind every trust boundary — provenance, retrieval, artifact execution, ranking, and the tradeoffs taken under time pressure |
| FAILURE_ANALYSIS.md | Real inputs where the system broke or degraded, root cause, and the fix — including gaps left open on purpose |
| docs/DECISIONS.md | Binding platform decisions in one page |
| docs/REQUIREMENTS.md | Assignment requirement → implementation map |
| docs/INTERVIEW_NOTES.md | 60-second architecture pitch and likely evaluator questions |
| docs/VIDEO_SCRIPT.md | Five-minute walkthrough script |
| docs/SUBMISSION_CHECKLIST.md | Pre-submission readiness checklist |







