Skip to content

Repository files navigation

📊 Census Insight Agent

A citation-grounded chatbot over Census 2011 India state reports — every factual claim traces back to an exact PDF page, checksum, and verbatim quote, or the system refuses rather than guesses.

offline-quality Python Next.js TypeScript FastAPI LangGraph Qdrant Docker Compose Tests

Ask it about literacy rates, population, sex ratios, and district rankings across Karnataka, Odisha, and Madhya Pradesh — or upload your own report. It looks things up, summarizes, compares, ranks, checks internal consistency, and builds charts/tables — and it says "I don't know" instead of making something up.

Census Insight Agent: a cited answer with sources, in the Next.js interface

Demo video

Watch the narrated project demo

Watch or download the demo (MP4, about 3½ minutes) · Transcript · Captions

Recorded from the running application on September 7, 2026, with synthetic English narration. Covers architecture, healthy services, a cited lookup, a follow-up, chart generation, and an out-of-scope refusal. Model waiting time is edited out and selected frames are held for explanation. The follow-up comparison failed citation validation in this run; the video preserves and explains that result. The lookup and chart succeeded. make check passed all 305 tests.

That specific comparison-refusal's root cause (resolve_query could rewrite a follow-up to name every target while losing the comparative framing itself, so nothing was ever computed to cite) is now fixed and verified live — see FAILURE_ANALYSIS.md. The video is kept unedited as an honest record of that session rather than re-recorded.

The video is included in the repository. If GitHub shows a file page instead of a player, use View raw or Download raw file to watch it.

The interface

A Next.js + TypeScript client built around one idea: every answer should be checkable in one click.

Deep research brief with verified sections Trust scorecard
Deep research. The agent plans 3–5 questions, runs each through the full citation checks in parallel, and assembles a brief you can export as Markdown or PDF. Declined sections stay declined. Trust scorecard. A live benchmark with hand-verified answers and trap questions, audited independently of the agent's own validation: correct answers, correct refusals, and numbers that don't appear in their own citation.
Interactive chart where each bar links to its source page Citation opens the original PDF page
Charts you can audit. Each bar is a verified source cell: hover for its page, click to open the original PDF. Rankings highlight the winner computed in code. Citation → original PDF page. Every citation chip opens the physical page of the source PDF next to the quoted evidence.
Live progress while the agent works Upload your own PDF
Live progress. Server-sent events show each LangGraph step as it starts; a research brief shows every section working in parallel. Bring your own PDF. Uploads go through the same citation-safe pipeline as the bundled reports, as a background job with live status.

More: district ranking with the winner highlighted · finished research brief · dark mode · mobile · empty state

  • Claude-style sessions: a searchable, date-grouped chat history, rename and delete, and shareable URLs (/c/<session_id>). Switch chats while one is still answering.
  • Answers that show their work: numbered citation chips on every claim, arithmetic shown as computed in code, a clear badge when the agent declines instead of guessing, a per-answer execution trace, and one-click follow-ups ("How does that compare with Odisha?") that exercise the validated conversation memory.
  • Hardened by default: the browser reaches FastAPI only through an allowlisted same-origin proxy; the container is read-only, non-root, and has no credentials.
  • Dark mode, mobile layout, keyboard shortcuts (⇧⌘O new chat), and live service health.

Table of contents


Measured trust

A benchmark of 26 questions with answers verified by hand against the source tables, each asked twice through the live API and scored independently of the agent's own validation code (backend/app/trust.py, cases in evals/trust_benchmark.json). Ground truth is facts, not loose numbers: for example (Odisha, sex ratio, 2011, urban, all persons, 932). Every number must also be found on its region's row, in the right column, of the quote it cites.

Check (52 live answers) Result
Wrong answers shown 0
Numbers not on their region's row and column in their cited quote 0 of 236 claims checked
Correct refusals (GDP, unemployment, cricket, a 2021 census) 8 / 8
Correct answers 37 / 44
Unnecessary refusals (the safe failure) 4
Operational errors 3
Median latency per question 20 s

The cases are built to catch misattribution, not just recall:

  • a 2001 figure next to the 2011 one;
  • urban versus rural versus total columns;
  • a comparison where Madhya Pradesh's child sex ratio equals its urban sex ratio;
  • the Scheduled Tribes and Castes tables beside the state figures;
  • follow-ups that name neither region nor metric;
  • a complete district ranking the agent is known to struggle with.

The misses are honest limitations, not wrong answers. In that run the agent declined questions about 2001, failed closed on district literacy rankings, and three cases gave different outcomes on their two runs. District rankings now read from structured tables: re-running the four ranking cases twice passed 8 of 8, with a median of 5.8 seconds. All are listed in FAILURE_ANALYSIS.md. A replay of real answers with deliberate corruptions (swapped regions, a neighbouring column, a wrong year) shows the scorer catches each one; the previous version passed all four.

This is one run on 2026-09-26 with gemini-2.5-flash. Reproduce it with make trust-benchmark (billable) and see it in the app at /trust.

What makes this different

A lot of "chat with your documents" demos silently give up citation accuracy the moment the answer requires a computed comparison, a chart, or a district ranking. This one doesn't:

Typical RAG chatbot Census Insight Agent
Citations Page number, sometimes Exact verbatim quote + physical PDF page + SHA-256 checksum, verified by application code — never trusted from the model
Comparisons ("how does X compare to Y") Model does the subtraction in prose Arithmetic is app-owned; the model never asserts a computed number that wasn't independently calculated
Charts/tables Pre-baked template, or the model fabricates numbers Model proposes which data; application code hydrates and verifies every cell against a trusted table row before any code executes
Code execution Often absent, or runs unsandboxed Generated Python passes an AST policy gate, then runs in a non-root, network-disabled, read-only container with resource limits
"I don't know" Rare — models like to please A structural outcome: insufficient evidence, a failed provenance check, or an unverifiable table cell all produce an explicit refusal
Memory Conversation replayed as prose Structured, validated claim history — a follow-up re-verifies against live evidence, it never trusts what a past turn said

Architecture

Four services, four trust boundaries. The frontend never touches Qdrant, credentials, or the executor; the executor never touches the network, Qdrant, or credentials.

flowchart LR
  subgraph P["🖥️ Presentation plane"]
    U((User)) --> UI["Next.js :3000\nallowlisted proxy"]
  end
  subgraph A["🧠 Agent / orchestration plane"]
    UI --> API["FastAPI :8000"]
    API --> LG["LangGraph"]
    LG --> V["Vertex Gemini via ADC"]
    LG --> S["Runtime skills\n(skills/*.md)"]
    LG --> M["SQLite checkpointer"]
  end
  subgraph E["📚 Source / evidence plane"]
    LG --> R["Hybrid retrieval"]
    R --> VE["Vertex embeddings"]
    R --> BM["Local BM25"]
    R --> Q["Qdrant RRF :6333"]
    PDF[("Authoritative PDFs")] --> Q
  end
  subgraph X["🔒 Isolated execution plane"]
    LG --> FQ["Filesystem queue"]
    FQ --> EX["Network-disabled executor\n(non-root, read-only, no creds)"]
    EX --> AR["Validated PNG / CSV / manifest"]
    AR --> API
  end

  classDef isolated fill:#fef2f2,stroke:#dc2626,color:#7f1d1d
  class X,EX isolated
Loading

How a request actually flows

This is what happens between you hitting Enter and an answer appearing — classification, per-target retrieval, evidence assessment, and a citation-validation gate that can trigger one bounded repair before it ever refuses.

sequenceDiagram
    autonumber
    actor You
    participant UI as Next.js
    participant API as FastAPI
    participant Graph as LangGraph agent
    participant Qdrant
    participant Gemini as Vertex Gemini
    participant Exec as Isolated executor

    You->>UI: "Compare literacy rates, Karnataka vs Odisha"
    UI->>API: POST /chat/stream
    API->>Graph: run(session_id, message)
    Graph-->>UI: SSE progress event as each node starts
    Graph->>Gemini: classify task + resolve standalone query
    Gemini-->>Graph: task_type=comparison, targets=[Karnataka, Odisha]
    par per-target retrieval
        Graph->>Qdrant: dense + BM25 search — Karnataka
        Graph->>Qdrant: dense + BM25 search — Odisha
    end
    Qdrant-->>Graph: RRF-fused candidates, reserved per target
    Graph->>Gemini: assess evidence relevance
    Gemini-->>Graph: relevance + entity/metric/unit match flags
    alt evidence insufficient
        Graph-->>API: citation-safe refusal + limitation
    else artifact requested (chart/table)
        Graph->>Gemini: propose minimal dataset (labels + values only)
        Graph->>Graph: hydrate + verify every cell against trusted evidence
        Graph->>Exec: generated Python + input.json (no credentials, no network)
        Exec-->>Graph: chart.png / table.csv + source-manifest.json
        Graph-->>API: validated artifact + citations
    else lookup / comparison / summary / ranking
        Graph->>Gemini: synthesize structured claims
        Graph->>Graph: validate citations (one repair if invalid, else refuse)
        Graph-->>API: cited answer
    end
    API-->>UI: result event (response + trace_id)
    UI-->>You: cited answer · PDF page viewer · trace
Loading

The provenance chain (the core guarantee)

Every number the agent ever states can be walked backward to a specific, checksummed page. Nothing skips a link in this chain — a broken link anywhere produces a refusal, not a guess.

flowchart LR
    PDF["📄 Original PDF\nchecksum + page count"] --> MAP["Page mapping\nMarkdown-preferred, PDF-anchored"]
    MAP --> CHUNK["Page-bounded chunk\nnever spans pages"]
    CHUNK --> VEC["Dense + BM25 vectors"]
    VEC --> QD["Qdrant point\nchunk + checksum + page"]
    QD --> RET["Retrieved candidate"]
    RET --> ASSESS["LLM relevance assessment"]
    ASSESS --> CLAIM["Structured claim\napp-built, not model-written"]
    CLAIM --> CITE["✅ Citation\nexact quote + physical PDF page"]
    CITE --> OUT["Answer or artifact"]

    style PDF fill:#fee2e2,stroke:#991b1b,color:#7f1d1d
    style CITE fill:#dcfce7,stroke:#166534,color:#14532d
    style OUT fill:#dbeafe,stroke:#1e40af,color:#1e3a8a
Loading

If any link breaks — an excluded visual page, a non-verbatim snippet, an unverifiable table cell — the chain stops and the system says so instead of silently continuing.

What it can do

You ask Task type What actually happens
"What was Karnataka's literacy rate in 2011?" lookup Retrieves, cites the exact table cell
"How does that compare with Odisha?" comparison Retrieves both sides independently, computes a validated arithmetic difference
"Summarize the key findings for Odisha" summary Page-distributed representative retrieval, not just the top-k hits
"Which district had the highest gender ratio?" artifact_table (ranking) Full-table scan; the winner is computed by application code, never asserted by the model
"Create a bar chart comparing Karnataka and Odisha literacy rates" artifact_chart Model proposes labels/values only; app hydrates and verifies every cell before the isolated executor renders a PNG
"Build a table of the urban vs rural breakdown" artifact_table Same hydration/verification path, output is table.csv + table.md
"Check whether the population totals in this table are consistent" inconsistency_analysis Deterministic arithmetic tool calls, checked against the same citation-validation gate
"Which source pages support those values?" source_support Rehydrates prior validated claims against current Qdrant state — never trusts conversation prose
"What was France's unemployment rate in 2011?" out_of_scope Refuses cleanly — no hallucinated cross-corpus answer
Upload a PDF, then ask about it any of the above Indexed through the same page-bounded, checksum-bound pipeline; the agent sees the live document catalog
Deep research: "Gender balance and literacy across Karnataka and Odisha" research brief A planner writes 3–5 questions; each runs as a separate validated agent turn; the brief contains only validated claims

Quick start

Reviewer quick start

The submission ZIP includes the three original PDFs and supplied Markdown. A Git checkout requires copying those files into data/source/ as described below. Credentials are never included. Vertex inference and first-time embeddings are billable.

After completing the ADC and .env configuration below, run:

sh scripts/setup.sh --allow-paid-calls

This builds images, checks source-page coverage, initializes an empty index, prepares Linux queue permissions, and waits for healthy services. It preserves an existing valid collection. Open http://localhost:3000. Subsequent starts need only docker compose up -d --wait. The one-shot queue-init container exits successfully; the four application services stay running. Windows users should use WSL2 with Docker Desktop integration enabled.

Credentials

For the complete application, install Docker Engine/Desktop with Compose v2 and the Google Cloud CLI. For host-side development, also install Python 3.12 and uv. The backend requires a Vertex-enabled GCP project and Application Default Credentials (ADC); API-key authentication is unsupported.

gcloud auth application-default login
gcloud auth application-default set-quota-project YOUR_PROJECT_ID
cp .env.example .env

In .env, set GOOGLE_CLOUD_PROJECT, the Gemini model names, and GOOGLE_CREDENTIALS_HOST_PATH to the absolute host ADC JSON path. Keep GOOGLE_APPLICATION_CREDENTIALS=/var/secrets/google/adc.json; it is the intentional container-local mount. Never copy credentials into the repository.

Why Vertex/ADC instead of a free-credit API key? Short version: project-scoped GCP identity without distributing a secret, at the cost of a slower reviewer setup. This tradeoff is made explicitly, not overlooked — see DESIGN.md § Tradeoffs.

Data setup and first initialization

The assignment corpus is ignored because its PDFs and Markdown total about 115 MB. Copy the supplied files into:

data/source/pdf/
  PC11_PCA_Data_Highlights_Karnataka.pdf
  PC11_PCA_Data_Highlights_Odisha.pdf
  PCA Data Highlights MP.pdf
data/source/markdown/
  PC11_PCA_Data_Highlights_Karnataka.md
  PC11_PCA_Data_Highlights_Odisha.md
  PCA Data Highlights MP.md

Pairing and reviewed page-coverage decisions are checked in under data/manifests/. Run the free dry run first; it makes no Gemini calls and no Qdrant writes:

docker compose run --rm --no-deps --build backend uv run --frozen python -m backend.app.ingestion.cli ingest --source-dir /app/data/source --dry-run --report-output /app/data/processed/dry-run-report.json
python3 -m json.tool data/processed/dry-run-report.json

After reviewing zero failures and the estimated inputs, initialize an empty collection intentionally. The initializer stops unless the paid-call flag is present, skips a valid 2,058-point collection, and never deletes a non-matching collection:

docker compose up -d qdrant
docker compose run --rm backend uv run --frozen python scripts/initialize_corpus.py --allow-paid-calls

Initial ingestion invokes billable Vertex embeddings. A fresh docker compose up --build starts the services but cannot answer corpus questions until the assignment files are supplied and initialization is approved.

Ingestion also writes each document's structured table store. For a collection indexed before table stores existed, build them from the indexed chunks with no model calls:

docker compose run --rm backend uv run --frozen python -m backend.app.tables.cli

Without a store, district rankings fall back to scanning chunks.

Run and review

docker compose up --build

The first build and FastEmbed load can take several minutes. Wait for docker compose ps to show four healthy services.

Service URL
Next.js chat UI http://localhost:3000
Streamlit UI (optional, docker compose --profile streamlit up) http://localhost:8501
FastAPI / OpenAPI docs http://localhost:8000/docs
Qdrant dashboard http://localhost:6333/dashboard

Try the example prompts above in one conversation — ask the lookup, then the follow-up comparison, then a chart, to see memory, citations, and artifacts all in the same thread.

Stop without removing indexed data:

docker compose down

Never add --volumes unless deleting the local Qdrant collection is intentional.

Safety and behavior

Original PDFs define identity, checksum, page count, and citation page. Supplied Markdown is preferred; PyMuPDF4LLM is fallback-only. Chunks never cross pages. Twelve unsafe Karnataka chart/map pages remain quarantined as excluded_unverified_visual; raw OCR never reaches embeddings, Qdrant, retrieval, or generation.

Every query gets Vertex dense and local BM25 sparse vectors; Qdrant performs RRF. Retrieval is candidate selection only. LangGraph assesses evidence, builds typed claims, selects exact spans from trusted current-run chunks, validates provenance, and answers or refuses. Comparisons reserve evidence budget for both regions.

AsyncSqliteSaver stores successful conversation state in workspace/checkpoints.sqlite. Run-local errors do not become durable claims. Runtime instructions are discovered from skills/*.md. Artifact proposals are deterministically hydrated with trusted evidence before generated Python reaches the unprivileged, network-disabled executor.

The executor has no cloud credentials, Qdrant, Docker socket, network, or session workspace. Docker isolation is not a hardened hostile multi-tenant sandbox; production use should add a stronger sandbox such as microVM isolation.

Verification and evaluation

A GitHub Actions workflow runs the full offline gate on every push, with no billable calls: format, lint, type check, all 418 Python tests (including the Postgres and Redis integration tests), an offline UI smoke test, a Compose trust-boundary audit that covers both overlays, a git-history secret scan, and a build (never a run) of each service's Docker image. A second job runs the executor under gVisor. Badge at the top of this file reflects the current main branch.

Run the same gate locally:

uv sync --frozen --all-extras
make check
make secret-scan

With the services and local corpus running, the canonical full offline gate adds an exhaustive read-only collection scan and dense/sparse compatibility queries. It performs no ingestion, Qdrant writes, Gemini generation, or Vertex embeddings:

make verify-offline

The Next.js client has its own gate: type check, ESLint, Vitest unit tests (including the proxy allowlist), and a production build. CI runs it as a separate job:

make web-install
make web-check

For UI development against a running backend, make web-dev serves the client with hot reload on http://localhost:3000.

Useful individual commands include make format-check, make lint, make typecheck, make test, make ui-smoke, make qdrant-readonly, docker compose config --quiet, and docker compose exec frontend python -m frontend.security_check.

workspace/checkpoints.sqlite grows without bound — every LangGraph superstep of every turn writes a full state snapshot, and nothing prunes old sessions automatically. make prune-checkpoints-dry-run lists sessions inactive more than 30 days (--older-than-days to change that); make prune-checkpoints actually removes their checkpoints, app_sessions row, and workspace/sessions/<id>/ trace/artifact directory. Neither command touches Qdrant or requires Vertex/GCP configuration.

The real-corpus retrieval evaluator uses a paid Vertex query embedding and is manual:

docker compose exec backend uv run --frozen python evals/run_retrieval.py --cases evals/real_corpus_cases.json --output /app/data/processed/retrieval-evaluation-report.json

The eleven-case live harness is also manual, requires explicit consent, never retries POST /chat, and stores a sanitized report:

docker compose exec backend uv run --frozen python scripts/live_evaluation.py --allow-paid-calls --output /app/data/processed/live-evaluation-report.json

Vertex connectivity alone can be checked with the billable make verify-vertex command.

API and traces

curl -sS -X POST http://localhost:8000/sessions
curl -sS -X POST http://localhost:8000/chat -H 'content-type: application/json' -d '{"session_id":"SESSION_ID","message":"What was Karnataka’s literacy rate in 2011?"}'
curl -sS http://localhost:8000/runs/TRACE_ID/trace

Trace IDs are allocated at request start and remain retrievable for typed failures. The UI's trace view shows allowlisted operational fields only. Files under workspace/ are ignored local state.

Endpoint Purpose
POST /chat/stream Server-sent events: progress per graph node, then one result or error
GET /sessions · GET /sessions/{id}/messages Chat history and a session's display transcript
PATCH /sessions/{id} · DELETE /sessions/{id} Rename; delete transcript, checkpoints, traces, and artifacts
POST /documents/upload · GET /documents/uploads/{job_id} Upload a PDF (202 + background job) and poll it
DELETE /documents/{id} Remove an uploaded document and its vectors (bundled corpus is protected)
POST /research/stream Deep research: plan, per-section progress/section_done, then the saved report
POST /runs/{run_id}/feedback Rate an answer up or down; stored beside its trace and sent to Langfuse when enabled
GET /evaluation/scorecard Latest trust benchmark run (local, else the copy in evals/)
GET /documents/{id}/pages/{n}?highlight=… Render a physical PDF page, highlighting the quote where the PDF has a text layer

Troubleshooting

Symptom Fix
port is already allocated Stop the other process/container using 6333. Do not remove the Qdrant volume.
pytest: executable file not found Run uv run --frozen pytest -q; dev tools are not directly on the production image PATH.
Unknown session Call POST /sessions, then pass its returned ID in valid JSON.
UI stays pending Rebuild frontend/backend and inspect their logs; POST /chat is intentionally not retried.
Backend unhealthy Verify ADC mount, project/location/model values, Qdrant, and executor heartbeat.

Known limitations: excluded visual-page content, progress streaming by graph step rather than by token (answers are only released after citation validation), scanned or vector-text uploads cannot be indexed without OCR review, quote highlighting is unavailable for the bundled PDFs (they have no text layer), loopback-only services without user authentication, and shared-kernel executor isolation unless the gVisor mode is on. See DESIGN.md and FAILURE_ANALYSIS.md for the full, honest accounting — including inputs where this system degrades, why, and what a real fix looks like. Production modes covers the opt-in scaled, sandboxed, and observed deployments.

Production modes

The default docker compose up is deliberately a single host: one API process, SQLite state, and in-process upload indexing. Each production concern is an opt-in mode that keeps every trust guarantee above, and each was verified live against the real stack. Details and failure behavior are in DESIGN.md.

Mode Turn it on What it changes How it was verified
Stateless API replicas make up-scale (needs POSTGRES_PASSWORD) Checkpoints, sessions, and transcripts move to Postgres (AsyncPostgresSaver). The per-session lock becomes a Postgres advisory lock, so a turn on one replica waits for an in-flight turn on another, and a crashed replica cannot wedge a session. A question answered on replica :8000 was followed up on :8001, then both containers were replaced with state intact. Integration tests run two replicas against real Postgres; they fail if the lock is removed.
Durable upload worker Same overlay Uploads are indexed by a Celery worker (Redis broker, append-only persistence) outside the API. Late acknowledgement redelivers a job whose worker died, and transient errors retry with backoff. A 60-page upload was killed mid-index with SIGKILL; Redis redelivered it and it finished with exactly 60 points, no duplicates.
gVisor executor make up-gvisor (Linux with runsc) Generated code runs on gVisor's user-space kernel instead of the host kernel, with every existing control still in place. Isolation checks, all offline cases, and a full queue handoff pass under runsc. The CI job gvisor-executor repeats this on every push.
Langfuse observability LANGFUSE_BASE_URL, LANGFUSE_PUBLIC_KEY, LANGFUSE_SECRET_KEY One trace per turn, grouped by session. Every Gemini call is a generation with token usage and latency. Answers get thumbs up/down scores, and trust scorecards become experiment runs with refusal precision and recall. Prompts and answers stay local unless LANGFUSE_CAPTURE_CONTENT=true. Checked against a self-hosted Langfuse 4.46 server, plus unit tests against the exact spans the SDK exports.

make integration-test runs the Postgres and real-Redis tests against throwaway containers, and CI runs them on every push. Still open: an authenticated gateway with TLS, a microVM per executor job (Firecracker or Kata Containers), and object storage so replicas can span hosts.

Repository layout

backend/app/       FastAPI, ingestion, retrieval, LangGraph, lineage
backend/tests/     Unit and integration tests
web/               Next.js + TypeScript client, allowlisted API proxy, Vitest tests
frontend/          Original Streamlit client (optional Compose profile)
executor/          Isolated worker and code policy
skills/            Runtime Markdown skills
data/manifests/    Pairing and reviewed coverage decisions
data/source/       Locally supplied corpus (ignored)
evals/             Evaluation cases and retrieval evaluator
scripts/           Checks, replays, initialization, live harness
workspace/         Ignored checkpoints, traces, queue, artifacts
docs/              Decisions and reviewer preparation

Further reading

Document What's in it
DESIGN.md The why behind every trust boundary — provenance, retrieval, artifact execution, ranking, and the tradeoffs taken under time pressure
FAILURE_ANALYSIS.md Real inputs where the system broke or degraded, root cause, and the fix — including gaps left open on purpose
docs/DECISIONS.md Binding platform decisions in one page
docs/REQUIREMENTS.md Assignment requirement → implementation map
docs/INTERVIEW_NOTES.md 60-second architecture pitch and likely evaluator questions
docs/VIDEO_SCRIPT.md Five-minute walkthrough script
docs/SUBMISSION_CHECKLIST.md Pre-submission readiness checklist

About

An agentic Census document analysis system supporting cited Q&A, summaries, computations, charts, tables, conversation memory, and safe Python execution.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages