Skip to content

About

Agentic Multi-Source News RAG Assistant — FastAPI + LangGraph + pgvector + React, powered by Guardian, NYT & Tavily, containerized with Docker and deployed on AWS ECS Fargate.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

Source

Ask. Research. Verify. — a production-ready AI news research platform, with Sage as the research guide you talk to. Sage searches live reporting from The Guardian, The New York Times and TheNewsAPI (an aggregator over thousands of outlets), indexes it into a vector store, and answers natural-language questions with grounded, citation-backed answers — summaries, comparisons, timelines and follow-ups — falling back to the open web only when the newsrooms can't answer.

Repository: https://github.com/AJKumarReddy/Agentic-News-app

React + TypeScript + Vite + Tailwind  →  FastAPI + LangGraph + pgvector  →  Guardian · NYT · TheNewsAPI · Tavily · OpenAI

Contents


Architecture

flowchart TB
    U[User Browser] --> F[React Frontend<br/>Vite · Tailwind · light/dark]
    F -- HTTPS REST + SSE --> N[Load balancer<br/>ALB in prod · Vite proxy in dev]
    N --> B[FastAPI<br/>Gunicorn + Uvicorn<br/>Fargate in prod · container in dev]

    B --> AG[LangGraph Agent]
    B --> PG[(PostgreSQL 16 + pgvector<br/>articles · chunks · chats)]
    B --> R[(Redis — cache + locks)]
    B --> SCH[Ingestion<br/>in-process in dev · scheduled task in prod]
    B --> AUD[Speech<br/>answers read aloud · questions transcribed<br/>both opt-in]

    AG --> UN[understand<br/>resolve · route · build queries]
    UN --> SRC[Sources]
    SRC --> G[Guardian Content API]
    SRC --> NY[NY Times API]
    SRC --> TNA[TheNewsAPI<br/>aggregator · metered daily budget]
    UN --> WEB[Tavily web search<br/>gated · cited separately]

    SRC --> FEED[Feed pipeline<br/>exclude · cluster · rank]
    FEED --> C
    SRC --> RAG[RAG Engine]
    RAG --> CH[Chunk 600–1000 tok]
    RAG --> EM[OpenAI embeddings]
    RAG --> HY[Hybrid retrieval<br/>vector + keyword + recency + edition]
    HY --> PG
    HY --> RR[Rerank + source diversity]
    RR --> LLM[OpenAI chat model]
    WEB --> LLM
    LLM --> C[Cited answer<br/>real publisher URLs]
Loading

Key design decisions

Decision Choice Why
Vector DB PostgreSQL + pgvector One database serves relational data and vectors — no second service to run, and writes stay transactional. One container locally, one RDS instance in production. rag/vector_store.py is small enough to swap for Qdrant if scale demands it.
Agent LangGraph, four modes understand resolves the message, then routes to ARTICLE / NEWS / WEB / BOTH. Each mode does only its own work, so an article question never touches the search machinery. Bounded and debuggable — no runaway autonomy.
Scope Deterministic guardrail Task requests (write code, solve maths, ghostwrite, roleplay) are declined by regex in agents/scope.py before any search or model call — a guardrail that asks the model to police itself fails exactly when the model misreads the request. News about those subjects is unaffected.
Query resolution Resolve before searching Follow-ups ("search youtube for related news", "now do a google search") are rewritten against the conversation first. Searching the raw words was the single largest source of wrong answers.
Multi-source Adapter per publisher Every source returns the same NormalizedArticle, so retrieval, chunking, citations and the UI are source-agnostic. Adding a newsroom is one adapter.
Per-article cap No article owns an answer Diversity across publishers was enforced; diversity across articles was not. One Guardian live blog could supply every final chunk, so a seven-story answer cited it seven times and each item linked to a headline about a different story. Each article now contributes at most RAG_MAX_CHUNKS_PER_ARTICLE chunks.
Round-ups Breadth, not depth "Top US news today" is a different retrieval problem from "what did the report say". A round-up (LATEST/SUMMARY/TREND, or any freshness question) retrieves one passage from each of ~10 articles rather than several from a few, and trims each to the article's opening so they all fit the evidence budget. That is what gives each listed story its own citation number pointing at the article that reported it.
Fair ranking Per-source retrieval Publishers expose wildly different text lengths (NYT gives abstracts only). A shared candidate pool silently excluded NYT entirely, so retrieval runs per source and merges.
Web fallback Tavily, gated Results exclude our own publishers' domains, must pass relevance/recency/low-signal-domain gates, and are cited separately. Disabled entirely without TAVILY_API_KEY.
Edition US desk preferred Guardian productionOffice is stored per article and nudges ranking. A nudge, not a filter — a better UK/AUS match still wins.
Freshness API-first for "latest/today" Recency questions are answered from current API results indexed on the fly, never from stale vectors with high semantic scores.
Date ranges Filter first, widen loudly The stated range is always tried first and alone. If it finds nothing the search widens — but the notice names the period that came up empty and the answer opens by saying so, rather than quietly serving results from outside it. Named periods are rolling windows ("this week" = 7 days), never calendar weeks that collapse to a single day on a Monday. The same range bounds publisher retrieval, NYT's Top Stories feed, and web results.
Sections A section is a filing decision, not a subject US political reporting lands in us-news, politics, world or commentisfree depending on the desk, so a section slug is widened to its subject neighbours before it reaches retrieval. Both sides are reduced to letters and digits first — the Guardian stores "US news", the NYT stores "U.S.", and we ask with "us-news". See sources/sections.py.
Dedup Article ID + SHA-256 content hash An article is embedded once; re-embedding only when content or the embedding model changes.
Model tier Chosen per turn, without a model call One model is either too slow for "summarise this article" or too shallow for "why did this happen". llm/routing.py picks FAST / GENERAL / REASONING from the intent the understand step already produced and from how much evidence came back — signals that exist by then, so no round trip is spent deciding. Tiers are settings, not hard-coded names, and reasoning defaults to the general model so no deployment is billed for one it did not choose.
Streaming Server-Sent Events Route decision, pipeline status and answer tokens stream into the UI.
Feed shape Canonical categories, one config Browsing by publisher slug meant finance, money, economy and business were four requests describing one subject, against providers that meter them. Five canonical categories now declare what each covers, what every provider calls it, and what share of the feed it starts with — in sources/categories.py, so weights are edited in one place rather than hunted through the codebase. Free text resolves through the same table, which is how "what's happening in AI?" reaches technology without the caller knowing any provider's vocabulary.
Non-news Filtered before ranking, not before rendering Crosswords and sponsored posts are filtered during normalisation because position is scarce: an advert that reaches the ranker competes on freshness like any article and can take the top slot. Structured signals are trusted (a sponsored section is the publisher's own assertion); headline text is a backstop for puzzles only. Promotional wording is deliberately not matched on headlines — "Meta's advertising revenue fell" is news about advertising, and removing it would be a worse failure than keeping one unlabelled advert. A puzzle instance is told from an article about puzzles by its serial number or its label before a colon. See sources/exclusions.py.
Ranking Importance × relevance × freshness, never published_at DESC Date-sorting puts a five-minute-old council item above an hour-old rate decision. services/ranking.py combines category weight, exponential time decay (six-hour half-life), significance terms, source authority, market impact and cross-source confirmation into one score. Every weight is a field on one dataclass; there are no magic numbers at the call sites. Freshness is strong but bounded — a genuinely major story outranks a fresher trivial one, and a stale story is actively penalised rather than merely faded.
Duplicate coverage Cluster first, then score Six outlets covering one rate decision is one story, not six cards. Articles are grouped by headline-token overlap, the highest-authority member fronts the cluster, and the rest become supporting citations. Ordering matters: scoring before clustering would let six near-identical copies each earn a top slot and crowd out the rest of the day. Independent outlets are counted, not articles — a wire story under three mastheads of one group is one newsroom's work, so syndication cannot fake corroboration.
Request budgets Per publisher, never shared Plans differ by an order of magnitude — 500/day on a Guardian or NYT developer key, 2,500 on TheNewsAPI's Basic plan — so each source declares its own ceiling and spends from its own Redis counter, keyed by source and UTC day. One shared cap would have to be the smallest plan in use, throttling a generous one to a mean one; it would also let a single greedy publisher mute the rest. The default is 0, meaning unmetered: an adapter whose plan nobody has checked is not throttled on a guess. Counting fails open, because a cache outage muting every publisher is worse than briefly overspending, and the publisher's own 402/429 already degrades gracefully.
Ingestion budget Index where the requests buy the most A bulk_efficient flag decides who joins the scheduled sweep, and it is a statement about yield per request, not about quality. On TheNewsAPI's free plan — three articles per request against 100 a day, against the Guardian's fifty — indexing bought ~17× less per request and every request spent sweeping was denied to a reader waiting on a search, so it stayed on the interactive path only. The Basic plan's 25 per request against 2,500 a day clears that bar, so it now sweeps too; the flag is the one line that changes. The sweep covers every desk retrieval can widen into: an un-ingested desk makes that widening a filter over an empty set, which looks broader and finds less.
Speech latency Segments in parallel, cached individually A long answer is split under the provider's input cap, and those segments used to be synthesised in a for loop — so a three-segment answer cost three round trips end to end and the reader waited for the sum. They are independent, so they now go concurrently (bounded by TTS_MAX_CONCURRENCY) and the wait is the slowest one. gather preserves order, which is what keeps the sentences in sequence. Each segment is cached on its own text as well as the whole answer, so a replay or a shared opening reuses work.
Autoplay refusal Not an error A play() the browser withheld for want of a gesture used to land in the same state as a failed request, telling the reader "audio unavailable" and sending them to retry a button that would fail identically. It is now detected separately and answered with a one-time "tap once to enable automatic audio", which clears permanently once anything has played.
Greetings Answered, not searched The understanding step's job is to resolve a fragment into a searchable question, so given "hi" it invents one — which is how a greeting returned a Guardian piece on water storage "understood as" a question nobody asked. GREET is now a mode alongside DECLINE: a deterministic check before the model runs, a fixed reply, no retrieval, no citation, and no cost. Anchored and length-capped, because the opposite failure is worse — "hi, what happened in Gaza today" must still search.

Repository layout

├── frontend/               React + TS + Vite + Tailwind
│   └── src/
│       ├── components/     Sidebar, chat bubbles, cards, citations, Sage panel,
│       │                   mic button, voice + theme toggles
│       ├── pages/          Search (the landing page) · Chat · Article intelligence
│       ├── hooks/          useChat (SSE) · useSpeech · useRecorder · useVoice
│       │                   · useCapabilities · useTheme
│       ├── services/       API client
│       ├── utils/          SSE parser, transcript, publisher labels, storage
│       ├── types/          shared response and state types
│       └── constants/      section taxonomy
├── backend/
│   ├── app/
│   │   ├── api/            chat · news · rag · audio · health routers
│   │   ├── agents/         graph, understand (resolve+route), dateparse, tools
│   │   ├── sources/        NewsSource abstraction · Guardian · NYT · TheNewsAPI
│   │   │                   categories (canonical + weights) · exclusions · quota · registry
│   │   ├── guardian/       Guardian client, normalizer, shared models
│   │   ├── websearch/      Tavily client + quality gates
│   │   ├── rag/            chunker · embeddings · vector store · retrieval · reranker · ingestion
│   │   ├── database/       SQLAlchemy models, session, repositories
│   │   ├── llm/            chat model factory, prompts (grounding rules)
│   │   ├── services/       chat orchestration/SSE, search, ranking, article
│   │   │                   intelligence, speech (TTS/STT), cache
│   │   ├── tasks/          scheduler, ingest_recent, edition backfill
│   │   └── core/           config, JSON logging, security middleware
│   ├── tests/              472 tests
│   └── evaluation/         30-question RAG evaluation harness
├── aws/                    ECS Fargate task definitions + production deployment guide
├── scripts/                health-check.sh · guardian_api_smoke.py
├── .github/workflows/      test.yml (tests) · aws.yml (build → ECR → ECS deploy)
├── docker-compose.yml      local dev: postgres · redis · backend · frontend
└── .env.example

Quick start (local)

Prerequisites: Docker Desktop, plus API keys — Guardian (free), OpenAI, optionally NYT, TheNewsAPI and Tavily (free tiers work; TheNewsAPI defaults here assume its Basic plan — see THENEWSAPI_* in the env table).

git clone https://github.com/AJKumarReddy/Agentic-News-app.git
cd Agentic-News-app
cp .env.example .env        # fill in the keys
docker compose up -d --build

Without Docker — run Postgres and Redis in containers, everything else on the host:

docker compose up -d postgres redis          # DATABASE_URL/REDIS_URL → localhost
cd backend && python -m venv .venv
.venv/Scripts/pip install -r requirements.txt -r requirements-dev.txt
.venv/Scripts/uvicorn app.main:app --reload  # :8000
cd ../frontend && npm install && npm run dev # :5173, proxies /api

Changing tailwind.config.js requires a dev-server restart — Vite hot-reloads CSS but not that config, and a stale config shows up as "class does not exist" or a blank page.

Configuration

Full list in .env.example. The ones that matter:

Variable Purpose
GUARDIAN_API_KEY · NYT_API_KEY · THENEWSAPI_API_KEY publishers; a source with no key is skipped, so the app runs on whichever keys exist
ENABLED_SOURCES active publishers in priority order (guardian,nyt,thenewsapi)
GUARDIAN_DAILY_BUDGET · NYT_DAILY_BUDGET · THENEWSAPI_DAILY_BUDGET requests per UTC day, counted per publisher (500 / 500 / 2,500). A source that runs out degrades to stored articles; the others are unaffected
*_INTERACTIVE_RESERVE of each budget, the slice background ingestion may not touch (150 / 150 / 1,000)
THENEWSAPI_PAGE_SIZE articles per request; the plan hard-caps it (Basic: 25) and rejects more
TTS_ENABLED · TTS_MODEL · TTS_VOICE answer playback; TTS_ENABLED=false removes the control from the UI entirely
TTS_MAX_CONCURRENCY segments of one answer synthesised at once; the wait becomes the slowest segment rather than the sum
STT_ENABLED · STT_MODEL · STT_LANGUAGE spoken questions; naming a language is cheaper and more accurate than auto-detect when a deployment knows its audience
STT_MAX_SECONDS · STT_MAX_BYTES the browser stops recording at the first (60 s); the endpoint refuses past the second (8 MB), so a forgotten open microphone cannot become a 413
OPENAI_API_KEY · CHAT_MODEL · EMBEDDING_MODEL generation and embeddings
CHAT_MODEL_FAST · CHAT_MODEL_REASONING the other two tiers llm/routing.py picks per turn. Reasoning defaults to the general model, so nobody is billed for one they did not choose
TAVILY_API_KEY optional web fallback; empty = newsroom-only, no web request ever made
WEB_SEARCH_THRESHOLD newsroom sources at or below this count trigger a web top-up
DATABASE_URL · REDIS_URL infrastructure — containers locally, RDS and ElastiCache in production, same two variables
INGEST_ENABLED · INGEST_INTERVAL_MINUTES · INGEST_SECTIONS_PER_TICK scheduled pulls (default: on, one section every 5 min — 288 requests/day, under the 500 cap). Off in production — see Scheduled ingestion
VITE_API_BASE_URL build-time arg for the frontend image; /api when one ALB serves both services
PREFERRED_PRODUCTION_OFFICE · EDITION_BOOST Guardian desk to favour (default US)
RAG_INITIAL_TOP_K · RAG_FINAL_TOP_K retrieval funnel (20 → 6)
RAG_MAX_CHUNKS_PER_ARTICLE · RAG_CANDIDATE_OVERFETCH most chunks one article may contribute to an answer (2), and how much wider the database is read to refill the pool after that cap (3×)
RAG_ROUNDUP_TOP_K · RAG_ROUNDUP_CHUNK_CHARS round-up questions retrieve one passage from each of this many articles (10), trimmed to this many characters (1200) so they all fit the evidence budget
RERANKER llm (default) · cohere · none
FRONTEND_URL · EXTRA_CORS_ORIGINS CORS allowlist — never * in production
API_KEY · VITE_API_KEY optional X-API-Key gate for private deployments
RATE_LIMIT_PER_MINUTE · CHAT_RATE_LIMIT_PER_MINUTE · AUDIO_RATE_LIMIT_PER_MINUTE per-IP budgets (30 / 10 / 20). The chat budget covers every path that costs an LLM call, an embedding or a publisher request; audio gets its own, because with autoplay on every turn is a chat request and an audio one and a shared bucket would read as broken chat
ALLOWED_HOSTS · MAX_BODY_BYTES · REQUEST_TIMEOUT_SECONDS optional Host allowlist, request body cap (64 KB) and time-to-first-byte ceiling (120 s)
TRUSTED_PROXY_HOPS proxies in front of the app (default 1 = one ALB/nginx). The client address is read that many entries in from the right of X-Forwarded-For; everything further left is written by the caller. Set 0 when the app is reached directly
ADMIN_API_KEY unlocks the operator endpoints (/api/rag/*, /api/intent) via X-Admin-Key. Unset, they are closed in production and open in development. Never put this in the SPA bundle

Locally these come from .env. In production nothing is read from a file — the ECS task definition lists plain values inline and pulls every secret (GUARDIAN_API_KEY, NYT_API_KEY, THENEWSAPI_API_KEY, TAVILY_API_KEY, OPENAI_API_KEY, DATABASE_URL, REDIS_URL) from SSM Parameter Store by ARN, injected as environment variables at task start. The application code is identical either way; only the source of the values changes.

Never commit .env, API keys, database passwords, AWS credentials or SSH keys — all covered by .gitignore.


How it works

1. Routing — the understand step

One LLM call resolves the message and routes it. Inspect its decision without spending a chat turn:

curl -X POST http://localhost:8000/api/intent -H "Content-Type: application/json" \
  -d '{"message":"what do other outlets say about the merger"}'

It is an operator endpoint, like /api/rag/*: open in development, and in production it answers 404 unless the request carries X-Admin-Key: $ADMIN_API_KEY — a 404 rather than a 403, so the response says nothing about what is there.

Route When Path
ARTICLE About the article you're viewing Reads that article — no search, no filters
NEWS News and current events (default) Publisher fetch → retrieve → rerank
WEB Not news (how-to, definitions, docs), or "search the web" Web search only
BOTH Needs reporting plus outside context Newsroom retrieval, then web

Resolution is the important half. "search for related news on youtube" becomes "related news about UK manufacturers facing cyber-attacks" — the subject comes from the conversation, and the instruction is handled separately. Naming a site always reaches the web and bypasses the low-signal domain filter: ask for YouTube, get YouTube.

The route, intent and resolved question appear as a badge above every answer.

2. The RAG pipeline

publisher APIs → normalize → dedup (id + content hash) → chunk (~800 tok, 15% overlap)
              → embed (batched) → pgvector
              → hybrid retrieve (cosine + Postgres FTS, fused with RRF)
              → recency + edition boost → rerank → source diversity → top ~6
  • Date handling is deterministic ("this week", "last 3 months" → exact ranges), not left to the model.
  • Narrow windows widen rather than fail. A "today" query early in the publishing day matches nothing; retrieval relaxes to 14 days, then to everything, and the UI shows a "Results from Last 14 days" badge — the answer itself never editorialises about the search window.
  • A round-up cites a story per number. [1] [2] [3] in a list of the day's news must each open the article that reported that item. Round-up questions therefore retrieve one passage from each of many articles; focused questions still retrieve several passages from the few that cover the development.
  • No article may be the whole answer. A live blog or a daily round-up is one article covering a dozen unrelated stories, so a broad question matches its chunks over and over. Retrieval caps each article's contribution (default 2 chunks) and reads wider to refill the pool, at every stage — per publisher, and again on what the reranker returns. Uncapped, "top US news today" returned seven stories citing a single Florida-primaries live blog.
  • Incremental only. The whole index is never rebuilt; an article is embedded once and re-embedded only if its content hash or the embedding model changes.

3. Citation integrity

This is the part most likely to mislead a reader, so it's enforced end to end:

  • Web results exclude our publishers' own domains, so they can never duplicate or impersonate indexed journalism.
  • Every source carries a type (publisher / web), its publisher name and machine id, sharing one citation numbering scheme.
  • Newsroom and web evidence get separate context budgets — long full-text chunks were otherwise starving web sources out of the prompt entirely, which caused the model to invent attributions.
  • Attribution must match the citation. If a Guardian article reports what Reuters found, the answer must say "The Guardian reports that Reuters found… [1]" — never "Reuters reported… [1]", which implies Reuters is the cited source.
  • Citation density scales with the answer. A single-article reply cites once; a multi-source synthesis cites per claim.
  • NYT entries are abstracts, not full articles, and the prompt says so, so the model never implies it read the whole piece.
  • The UI renders web citations in amber with a Web · domain label, distinct from newsroom citations.

4. Spoken answers

Answers can be read aloud through the OpenAI speech API. Two independent gates govern it: TTS_ENABLED decides whether the feature exists in a deployment, and each reader's own preference — kept in their browser and off by default — decides whether anything is ever spoken. Nothing is spent until someone opts in.

Always on autoplays each newly completed answer, on the chat page and in the Sage panel alike. Three rules keep that from becoming noise:

  • History never speaks. Messages restored from a stored conversation are marked as already-spoken before they can reach the autoplay path. This is tracked explicitly rather than inferred from render timing, which cannot tell a historical message from a new one.
  • One answer at a time. A new question stops the previous answer; closing the panel stops playback rather than leaving it talking to a closed drawer.
  • A blocked autoplay is not an error. Browsers withhold audio until the page has a user gesture. That is answered with a one-time "tap once to enable automatic audio" rather than an error state, and it clears permanently once anything has played.

Latency comes from the split: a long answer exceeds the provider's per-request input cap, so it is cut between sentences and each segment synthesised concurrently rather than in sequence. Segments are cached individually as well as as a whole, so a replay — or a second answer sharing an opening — reuses the work.

  • Addressed by stored message, never by text. POST /api/audio/speech takes a conversation id and a message id, and resolves them through the same ownership check that guards GET /api/conversations/{id}. Accepting raw text would make the endpoint an open relay for speech billed to our key.
  • The prose is prepared first. app/core/text.py strips citation markers, headings, bullets and link URLs, and drops tables whole — a listener hears "bracket one" as a defect, not as a source. That module also owns the citation regex the agent graph uses for history replay, so both agree on what a citation is.
  • Its own rate-limit bucket. With autoplay on, every turn is a chat request and an audio request; sharing the chat budget would halve usable chat throughput and the 429 would read as a broken chat.
  • Cached on the text, not the message, so the same answer reached from another conversation is free. Redis holds it for an hour — audio dwarfs the JSON around it under an LRU cap — and the browser caches the long tail.

5. Spoken questions

The mirror of playback, and gated the same way: STT_ENABLED decides whether a deployment offers it, the mic button only appears where the browser can actually record, and nothing is captured until the reader presses it.

  • The recording is the request body, not a multipart form — the only field is the audio, which keeps python-multipart out of the dependency list for nothing lost. BodySizeLimitMiddleware caps it per path, and the byte length is re-checked after the read, because a chunked upload declares no Content-Length for the middleware to reject.
  • The container is matched, never trusted. Chrome and Firefox record WebM/Opus and Safari only ever produces MP4, so the browser probes MediaRecorder.isTypeSupported rather than assuming, and the server maps the declared content type against a fixed table — the filename handed to the API is ours, never a string the caller chose. Anything else is a 415.
  • No ownership check, and no cache. Unlike /speech it reads nothing back to the caller, it only hands them their own words; and a recording is a few hundred kilobytes that will never be sent again, so caching it would evict the article entries that keep the site up when a publisher is unreachable.
  • Silence is not a failure. Stop pressed without speaking returns 204, and the composer simply returns to idle. A blocked microphone is reported as the permission problem it is, not as a generic error.
  • Transcribed text is sanitised and dropped into the composer, never sent on its own. Speech-to-text mishears exactly the proper nouns a news question turns on, and a wrong question sent automatically costs a retrieval and an answer to undo — so the reader reads what was heard, edits if needed, and presses send. The caret lands after the insert, so typing continues the sentence.

The feed pipeline

What reaches the homepage, and the order the stages run in. The order is not incidental — each stage exists because doing it later produces a specific, observed failure.

publishers (Guardian · NYT · TheNewsAPI)
        │   fanned out concurrently; a failed source degrades to stored articles
        ▼
   normalize            one NormalizedArticle shape, whatever the provider
        │
        ▼
   EXCLUDE              crosswords, puzzles, sponsored, affiliate
        │               ── before ranking: position is scarce, and an advert
        │                  that reaches the ranker competes on freshness
        ▼
   CLUSTER              one event = one card, other outlets become citations
        │               ── before scoring: six copies would each earn a slot
        ▼
   SCORE                category + freshness + significance + authority
        │                  + corroboration + market impact − staleness
        ▼
   order and render

Why each stage sits where it does

Exclusions before ranking. A sponsored post is fresh, well-formed and carries a plausible headline. Every signal the ranker reads says "promote this". Filtering at render time would leave it holding a priority slot that a real story should have had.

Clustering before scoring. Scoring first means six near-identical reports of one rate decision each score highly on their own merits and take the top six positions, pushing the rest of the day off the page. Clustering first turns that into one strong story — and the fact that six independent newsrooms ran it becomes evidence of importance rather than repetition.

Independent outlets, not article count. A wire story republished under three mastheads of one group is one newsroom's work. Counting articles would let syndication manufacture the corroboration signal, which is precisely the signal that is otherwise hardest to game: an outlet can write any headline it likes, but it cannot make five others cover the same event.

Freshness strong but bounded. published_at DESC is the obvious sort and the wrong one — it puts a five-minute-old parish-council item above an hour-old supreme court ruling. Decay is exponential with a six-hour half-life (15 min → 0.97, 2 h → 0.79, 12 h → 0.25, 48 h → 0.004), which keeps a morning story alive through the working day while a two-day-old one is effectively gone. Past 72 hours an article is actively penalised, not merely faded: a rolling feed that still leads with last week is worse than a shorter one.

Undated is not fresh. An article with no timestamp is scored as stale rather than as "now". The alternative lets every article missing a date lead the feed.

Market impact is not asserted. It applies only to stories already categorised as business, and only from terms actually present in the headline and standfirst. Inventing market relevance for a sports story would be a lie the ranking tells itself.

The fifth category is deliberately broad. "High-Impact Other" exists so the homepage cannot collapse into politics-finance-tech-sports. An earthquake, a court ruling, a public-health event or a scientific discovery reaches the top on its own significance, regardless of its category's normal allocation.

Category weights

Starting shares, not quotas — a major story is expected to override them:

Category Weight Fetched first
Politics & Government 32% yes
Business, Finance & Economy 22% yes
Technology 18% yes
High-Impact Other 15% no
Sports 13% no

default_weights() is a function rather than a constant so a later personalisation layer can compose without touching the callers:

effective = default_weights() + user_interest + query_intent + breaking_event

News sources

Source Coverage Notes
The Guardian Full article bodies, all sections, deep archive The richest evidence; productionOffice enables US-desk preference
The New York Times Headlines, abstracts and lead paragraphs The API exposes no article bodies — evidence is short by design
TheNewsAPI Thousands of outlets, US-weighted An aggregator, not a masthead. Descriptions and snippets only; each article is cited under its own publisher. Metered: 2,500 requests/day, 25 articles each on the Basic plan, split between the sweep and live search
Tavily (web) Everything else Supplementary only, gated and cited separately

NYT specifics worth knowing:

  • NYT enables each API per key. A key valid for Top Stories may be rejected by Article Search. The adapter detects a 401 once, remembers it, and falls back to Top Stories rather than dropping NYT entirely — enable "Article Search API" for your app at developer.nytimes.com to unlock keyword search of the archive.
  • A keyword-less section browse uses Top Stories directly, since Article Search would otherwise be asked for the literal word "news".
  • Rate limits are tight (~5 req/min, 500/day); the per-minute floor is a throttle in the adapter, the daily cap is its own NYT_DAILY_BUDGET counter, and responses are cached on top of both.
  • Its multimedia field has shipped as a list of objects, a list of strings and a dict — all three are handled.

TheNewsAPI specifics worth knowing:

  • It relays other newsrooms, so source carries the publisher that actually reported the story ("cnn.com", "salon.com") while source_id stays thenewsapi. Citing the aggregator would misattribute every article.
  • The plan is metered per day — Basic: 2,500 requests/day and 25 articles per request (free: 100 and 3), on its own counter, separate from the Guardian's 500 and the NYT's — against a scheduler that ticks every five minutes. A daily counter in Redis is shared across workers, and background ingestion is cut off at THENEWSAPI_DAILY_BUDGET - THENEWSAPI_INTERACTIVE_RESERVE so it can never drain the quota a reader's search needs. Running out degrades to stored articles, exactly like an unreachable publisher. Moving plans is three settings, not a code change; bulk_efficient in the adapter is the one judgement call that goes with them.
  • /api/health reports it from the budget rather than by calling the API — a real probe on a 60-second cache would spend ~1,400 requests/day, over half the Basic plan and many times the free one.
  • Articles arrive tagged ["general", "politics"] and similar; the category asked for wins, then anything more specific than the catch-all, so a section browse files its results where the section filter will find them.
  • Its ten categories are mapped into the app's Guardian-flavoured section vocabulary on the way in ("tech" is stored as "Technology").
  • Because it aggregates hundreds of image CDNs, the frontend CSP allows img-src https:. See the comment in frontend/nginx.conf.

Because NYT and TheNewsAPI chunks are an order of magnitude shorter than full-text ones, retrieval runs per source and merges, and a diversity pass guarantees each publisher a foothold. Without it, answers silently became single-source.

Scheduled ingestion

One loop, one interval, every environment — including production. A tick refreshes one desk and cycles through the 13, so the index stays current without an external cron.

Interval Scope per tick Full cycle Requests/day per publisher
Local and production 7 minutes one desk, rotating 91 minutes ~205

The rotation covers 13 Guardian desks, weightiest category first (sources/categories.py → ingest_sections()) — every desk the five canonical categories cover, not one per category. That is a correctness requirement rather than thoroughness: retrieval widens a slug into its subject neighbours, so a question about us-news also looks under politics, world and commentisfree, and a desk that is never ingested makes that widening a filter over an empty set. Ordering by weight is what makes a cut-short cycle degrade gracefully — the desks the feed leads with come round first after a restart.

Every publisher marked bulk_efficient is swept — the Guardian, the NYT and, since the move to TheNewsAPI's Basic plan, TheNewsAPI as well. The flag asks one question: does a request buy enough articles to be worth spending a metered budget on indexing rather than on a search someone is waiting for? At three articles a call it did not; at 25 it does.

Each publisher's cap is enforced on its own counter (sources/quota.py), so the sweep is charged per source and one publisher running out never mutes the others.

The ceiling that binds is not the plan's. Background work is cut off at DAILY_BUDGET − INTERACTIVE_RESERVE — 350 for the Guardian and the NYT — which is what keeps a sweep from spending the requests a reader is waiting on. One desk every 7 minutes is ~205/day, comfortably inside it. Two desks on that interval is 411, and past the ceiling the sweep does not slow down: it stops for the rest of the UTC day and the index goes stale. Keep (1440 / INGEST_INTERVAL_MINUTES) × INGEST_SECTIONS_PER_TICK under 350 when tuning.

# manual run — sweeps all 13 desks at once, for a cold index
docker compose exec backend python -m app.tasks.ingest_recent

Under multiple Gunicorn workers — and under multiple Fargate replicas — a short-lived Redis lock ensures exactly one performs each tick, so publisher API usage isn't multiplied by the worker count. The rotation is derived from the clock rather than in-process state, so every worker resolves the same slice for a tick and a restart resumes the cycle in place. The lock fails open: a Redis outage still ingests rather than silently freezing the index. A failed run is logged and retried on the next tick.

Checking it actually runs. Filter the backend log for ingest. scheduled_ingest_tick and scheduled_ingest_section every 7 minutes is the healthy state; a stream of scheduled ingestion disabled means INGEST_ENABLED is not true. That distinction is worth knowing because a scheduler that never runs reports nothing at all — production once sat that way, its index fed only by the incidental indexing a freshness question triggers.

An earlier design ran ingestion as an EventBridge one-shot Fargate task instead, to avoid tying cron work to long-running replicas. It is not used: a direct invocation sweeps all 13 desks in one run rather than a rotating slice, so the interval multiplies by 13 and any short cadence lands far past the ceiling. See aws/ECS_PIPELINE.md §9.

Interface

Three pages, all responsive, in light or dark theme (follows your OS until you choose, then persists per browser):

  • Chat — streaming answers with a route badge, inline citation chips and a grouped source list
  • Search (the landing page) — ten newsroom desks, date range, sort, per-publisher filter chips, real pagination (Page 1 of 50 · 131,819 results). The desk list is deliberately short: eighteen entries across three groups gave Fashion and Books the same weight as Politics in a rail people scan for the news, and narrower subjects are ordinary search queries anyway
  • Article intelligence — AI summary, key points, entities, topics, important dates, related coverage, and "Ask AI about this article"

Sage travels with the reader. Search is the landing page, so a question usually arrives while someone is looking at articles rather than at the chat. A floating panel answers beside the page instead of navigating away from it — the same useChat hook, the same streaming, the same citations as the full page, and the thread it starts is a normal conversation that "Open full view" carries to /chat by id. It is never rendered on /chat, where Sage is already the page.

Voice on both sides, both opt-in. A mic button records a question and sends it to /api/audio/transcribe; answers can be read aloud through /api/audio/speech. Each is gated twice — by what the deployment can serve (GET /api/capabilities, so a control the backend cannot honour is never rendered) and by the reader's own preference, kept in their browser and off by default. Recording is also hidden outside a secure context, since getUserMedia cannot work over plain HTTP.

Chats are scoped to an anonymous per-browser id (X-Client-Id), so one visitor never sees another's history; individual chats and the whole history can be deleted from the sidebar.

Backend API

Endpoint Description
POST /api/chat agentic chat; {"message", "conversation_id?", "article_id?", "stream": true} → SSE (state, route, status, token, sources, notice, done, error) or JSON with stream:false
POST /api/intent routing decision only — no search, no answer, nothing written. Operator-gated (X-Admin-Key)
GET /api/news/search multi-source search: q, from_date, to_date, section, order_by, page, page_size, sources
GET /api/news/sources active publishers, for the UI's filter
GET /api/news/article/{id} normalized article (index first, then publisher API)
GET /api/news/article/{id}/intelligence AI analysis + related coverage
POST /api/rag/retrieve hybrid retrieval with metadata filters. Operator-gated, like every /api/rag/* route
POST /api/rag/ingest index by article ids and/or a search query
GET /api/conversations · GET /api/conversations/{id} chat history, scoped to X-Client-Id
DELETE /api/conversations/{id} · DELETE /api/conversations delete one chat / clear history
DELETE /api/conversations/{id}/article unpin the article a chat is anchored to, so it stops answering about that piece
POST /api/audio/speech audio for one stored answer; {"conversation_id", "message_id"} → audio/mpeg, 204 when nothing is speakable
POST /api/audio/transcribe text for one spoken question; the raw recording as the body → {"text"}, 204 when nothing was said
GET /api/capabilities which optional features this deployment can serve, so the UI hides what it cannot
GET /api/health database, vector extension, cache and each publisher

Testing

cd backend && pytest -q     # 472 tests
cd frontend && npm test     # 114 tests, 20 files

Covers the Guardian, NYT and TheNewsAPI adapters (mocked HTTP), the daily request budget and its page-size clamp, feed ranking and clustering, non-news exclusions, canonical categories, greeting short-circuit, chunking, dedup, RRF fusion, edition boost, source diversity, reranking, routing and resolution, date parsing and filtering, scope refusals, the article pin, conversation continuity, the scheduler's lock, security middleware, abuse handling, speech text preparation, playback ownership and transcription, API contracts, and on the frontend the SSE parser, citation components, the Sage panel, the recorder and the voice controls.

RAG evaluation — 30 questions against a running stack with real keys:

cd backend && python evaluation/run_eval.py --base-url http://localhost:8000

Checks citation presence, real publisher URLs, honest refusals on unanswerable questions (anti-hallucination), routing accuracy and follow-up context retention.


Deployment

Production runs on AWS ECS Fargate, deployed by GitHub Actions. Docker Compose is the local development stack and is not used in production.

Production: GitHub Actions → ECR → ECS Fargate

Developer ──► git push main ──► GitHub Actions
                                    ├── test.yml  (pytest · vitest · docker builds)
                                    └── aws.yml   (deploy — gated on tests)
                                                                │
                                            Docker images ──► Amazon ECR
                                                                │
                                                        Amazon ECS Fargate
                                                                │
                                                  Application Load Balancer (HTTPS)
                                                     ├── /api/*  ──► backend  :8000
                                                     └── /*      ──► frontend :80

Behind the backend service:

Backend
 ├── Amazon RDS PostgreSQL + pgvector      DATABASE_URL
 ├── Amazon ElastiCache Redis              REDIS_URL
 ├── SSM Parameter Store / Secrets Manager all API keys, injected at task start
 ├── Amazon CloudWatch Logs                /ecs/guardian-backend · /ecs/guardian-frontend
 └── Guardian · NYT · TheNewsAPI · Tavily · OpenAI
Piece Where
Deploy workflow (build both images, push to ECR, register + roll both services) .github/workflows/aws.yml
Fargate task definitions (awsvpc, backend :8000, frontend :80, awslogs, SSM secrets) aws/taskdef-backend.json · aws/taskdef-frontend.json
Full setup: VPC, security groups, RDS, ElastiCache, IAM, ALB, services, scheduler, deployer policy aws/ECS_PIPELINE.md

Notes that matter in production:

  • Secrets never live in the repo, images or task definitions — the task definitions carry SSM parameter ARNs and ECS injects the values at task start.
  • Postgres and Redis are managed services, not containers in the task. The application is unchanged: it still reads DATABASE_URL and REDIS_URL.
  • ALB idle timeout must be ~300 s — /api/chat streams over SSE and the 60 s default cuts long answers off mid-stream.
  • INGEST_ENABLED=true on the services; ingestion is the in-process loop, deduplicated across replicas by a Redis tick lock. There is no EventBridge schedule.
  • RDS and ElastiCache are never publicly accessible — private data subnets, security groups that only accept traffic from the ECS tasks.
  • The frontend is baked at build time. The workflow passes VITE_API_BASE_URL (default /api) into the image, so changing the API origin means a rebuild, not an env var flip.
  • The image in each task definition is a placeholder. The deploy overwrites it with the commit-SHA tag on every run and registers a new revision.
  • Task definitions are versioned with the code. Changing an env var or secret ARN is an edit to aws/taskdef-*.json plus a push — no console step.

Order for a first deployment: VPC and security groups → ECR → RDS and ElastiCache → SSM parameters → IAM roles → cluster, log groups and task definitions → ALB → ECS services → scheduled ingestion → deployer credentials. Each step is a copy-pasteable block in aws/ECS_PIPELINE.md.

Both task definitions are set to us-east-2 — the ECR image URI, the SSM parameter ARNs and awslogs-region. Deploying elsewhere is one sed over the two files plus AWS_REGION in the workflow. No account id, VPC, subnet, security group or ARN is committed: the account id stays an <ACCOUNT_ID> placeholder, substituted at deploy time from sts:GetCallerIdentity.

Rough cost: ALB + 2 Fargate services + RDS + ElastiCache ≈ $80–120/month at the smallest sensible sizes.

Local: Docker Compose

docker compose up -d --build

Services: postgres (pgvector), redis, backend, frontend. Postgres and Redis bind to loopback only.

Where your data lives: the Docker volume project_postgres-data (/var/lib/docker/volumes/project_postgres-data/_data, inside the Docker VM). It survives docker compose down; only down -v destroys it. Back up with:

docker exec project-postgres-1 pg_dump -U guardian guardian_news > backup.sql

Article text and embeddings are rebuildable from the APIs — conversations are the only irreplaceable data.

CI/CD

Tests — test.yml runs on every push and pull request: backend pytest + import validation, frontend vitest + production build, both Docker image builds.

Deployment — aws.yml on every push to main:

  1. test.yml runs first as a reusable workflow; a failing test stops the deploy.
  2. A matrix job runs once per service — build the image from its own context, push it to ECR tagged with the commit SHA, render that image into the task definition from aws/, register the revision, and update the service.
  3. wait-for-service-stability holds the job open until the rollout settles. New tasks must pass the target-group health check (/api/health for the backend) before the old ones drain.

Repository secrets required: AWS_ACCESS_KEY_ID and AWS_SECRET_ACCESS_KEY, for a deployer IAM user scoped to ECR push + ecs:RegisterTaskDefinition/UpdateService + iam:PassRole for the two task roles. The exact policy — and the OIDC alternative that removes the long-lived key entirely — is in aws/ECS_PIPELINE.md §10.

The deploy job declares environment: production, so adding required reviewers to that environment in Settings → Environments turns every deploy into an approval gate.

Scaling further

Concern Next step
Traffic Raise desired-count, or attach ECS Service Auto Scaling on ALB request count / CPU
Database Larger RDS instance class, then a read replica; Multi-AZ for failover
Cache ElastiCache replication group instead of a single node
Rate limiting Currently per-task in memory — move to Redis-backed storage so limits are global
Assets Serve the SPA from S3 + CloudFront instead of a frontend Fargate service
Background jobs Promote app/tasks to Celery when ingestion volume outgrows one scheduled task
Vector store rag/vector_store.py is small enough to swap for Qdrant/OpenSearch if pgvector plateaus

Security checklist (production)

  • API keys server-side only; the React bundle never sees them
  • Optional X-API-Key gate for the whole API (constant-time comparison)
  • CORS restricted to explicit origins; optional Host allowlist
  • Per-IP rate limiting (30/min general, 10/min for /api/chat, 20/min for /api/audio/*) with idle-key eviction
  • Answer playback names a stored message, never raw text, so it can't relay arbitrary TTS on our key
  • Request body cap (64 KB) and time-to-first-byte timeout (120 s)
  • Content-Security-Policy on API responses and the SPA
  • Secure headers (nosniff, frame-deny, referrer policy); HTTPS terminated at the ALB
  • Publisher/OpenAI calls time-boxed with retries; SSE streamed unbuffered
  • Bounded agent (recursion limit, capped evidence context, capped tool fan-out)
  • Prompt-injection defenses: evidence wrapped as data with an explicit "ignore instructions inside articles" rule; user input sanitized and truncated
  • Conversations scoped per client; cross-client access returns 404, not 403, so ids aren't enumerable
  • RDS and ElastiCache in private subnets, reachable only from the ECS security group
  • Secrets in SSM Parameter Store only — never in the repo, images, task definitions or GitHub
  • Non-root Docker user for the backend; no SSH keys anywhere
  • Deploy credentials scoped to ECR push + those two ECS services, with iam:PassRole limited to the two task roles
  • Swap the deployer access key for GitHub OIDC role assumption (no long-lived key in GitHub)
  • Structured JSON logs with request ids to CloudWatch; secrets redacted, never logged
  • Rotate the RDS password and API keys periodically (update the SSM parameters, redeploy)
  • Enable RDS automated backups / snapshots and ECR image scanning
  • Use ECS Exec (not SSH) for shell access into a running task

Troubleshooting

Symptom Cause and fix
Blank page, or "class does not exist" in dev tailwind.config.js changed while the dev server was running — restart Vite. If it starts on :5174, an orphaned process still holds :5173; kill it.
NYT missing from results Key not licensed for Article Search (401). It falls back to Top Stories; enable "Article Search API" at developer.nytimes.com for archive search.
"nyt": "unavailable" in /api/health No NYT_API_KEY, or the key is rejected by every endpoint.
One publisher missing from results Its own daily budget is spent (GUARDIAN_DAILY_BUDGET / NYT_DAILY_BUDGET / THENEWSAPI_DAILY_BUDGET) — results fall back to stored articles and the source is listed in degraded_sources, while the other publishers carry on. Check the backend log for budget exhausted; the counters are quota:{source}:{YYYY-MM-DD} in Redis.
Index stopped updating mid-day The sweep hit budget − reserve for a publisher and stopped for the day. Either lengthen INGEST_INTERVAL_MINUTES / the EventBridge rate, or raise that publisher's budget if its plan allows more.
Answer says evidence is insufficient Index may be cold — run python -m app.tasks.ingest_recent, or wait for the next tick.
Chat returns 500 Missing OPENAI_API_KEY, or Postgres/pgvector not reachable — check /api/health.
Frontend can't reach the API VITE_API_BASE_URL mismatch, or the origin isn't in FRONTEND_URL/EXTRA_CORS_ORIGINS.
No mic button, or no playback control GET /api/capabilities reports the feature off — no OPENAI_API_KEY, or STT_ENABLED/TTS_ENABLED false. Recording additionally needs a secure context: over plain HTTP (not localhost) getUserMedia cannot work, so the button is hidden rather than rendered to fail.
Answers never speak although voice is on Autoplay was withheld for want of a user gesture — the panel says "tap once to enable automatic audio", and it clears permanently after the first playback. Restored history is deliberately silent.
Chat stream cuts off in production ALB idle timeout — raise it to 300 s. See aws/ECS_PIPELINE.md.
ECS tasks stop right after starting Usually a missing SSM parameter or no route to ECR — check the stopped reason and /ecs/guardian-backend.

License and attribution

For educational and portfolio use. Content is subject to each publisher's terms — the Guardian Open Platform and the NYT Developer terms. Non-commercial tiers require retaining attribution and linking back to source articles, which the citation system does by design.

About

Agentic Multi-Source News RAG Assistant — FastAPI + LangGraph + pgvector + React, powered by Guardian, NYT & Tavily, containerized with Docker and deployed on AWS ECS Fargate.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages