Ask the ocean a question in plain English. Get an answer you can check.
Overview · Architecture · Tech Stack · Modules · Guard · It Refuses · Integrity · Workflow · Quick Start · Deployment · Data · Docs
Vortex turns a question in plain English into a plan of database operators, executes it against 55,795 ARGO ocean float profiles, and answers with grounded prose and charts.
The design goal is not fluency - it is an answer you can check:
- 🔢 Every number is computed in PostgreSQL, never by the language model
- ✅ Prose is verified against those numbers by deterministic regex before you see it
- 🚫 A question the data cannot support is refused with the reason, not answered emptily
- 📋 Every answer carries its plan, its timing, its assumptions and its sources
Architecture follows TAIJI (arXiv:2505.11270) specialised to ARGO, with MCP-Guard (arXiv:2508.10991) defence and Memento (arXiv:2508.16153) case-based memory.
| Layer | Technology |
|---|---|
| 🎨 Frontend | React 19 · TypeScript · Vite · Tailwind CSS v4 · React Router 7 |
| 📊 Charts | Vega-Lite specifications, rendered client-side - never server-drawn images |
| ⚙️ Backend | FastAPI · uvicorn · five tool servers as isolated worker subprocesses |
| 🗄️ Database | PostgreSQL 16 + PostGIS - array storage, one row per profile |
| 🧠 Semantic search | sentence-transformers E5-base (768-d) over a filter-aware HNSW index |
| 💬 Synthesis | Groq-hosted LLM, with non-model grounding checks |
| ☁️ Deployment | Vercel (frontend) + AWS EC2 (bridge and Postgres on one instance) |
| # | Module | What it does |
|---|---|---|
| 1 | 🗣️ NL Parser | Resolves a question into typed slots - region, time span, parameter, depth - and refuses spans the deployment never ingested |
| 2 | 💰 Cost-Based Planner | Enumerates candidate operator DAGs and prices each one against calibrated latency constants before executing any |
| 3 | 🔧 Five MCP Tool Servers | structured · metadata · analysis · semantic · visualization, each an isolated worker process |
| 4 | 🌉 MCP Bridge | REST proxy with a worker pool and risk-based execution |
| 5 | 🛡️ MCP-Guard | Layered detection enforced at three points in the pipeline |
| 6 | 🧠 Case-Based Memory | Five banks that learn from past questions - storing no numerals, ever |
| 7 | ✍️ Grounded Synthesis | Prose whose every numeric claim is checked against the executed result |
| 8 | 💻 Web UI | Chat, dashboard, editable assumptions, receipts and export |
💡 The semantic server runs out of process for a concrete reason: loading torch in a process that holds a psycopg connection causes an access violation. That is why it is a bridge worker and not a function call.
Detection runs in stages - a pattern scan, then a trained classifier. Enforcement happens at three separate points, because an attack has three separate routes in:
| Point | Where | Why it exists |
|---|---|---|
| P1 | Ingress - the user's question | Before anything reads it |
| P2 | Egress - the composed answer | Before it leaves the system |
| P3 | Retrieved documents | A document can carry an instruction aimed at the model reading it |
A blocked request names the stage that caught it rather than failing silently.
Real responses from the deployed system - not illustrations:
|
✅ A question it can answer
|
🚫 A year that was never ingested
|
|
🛡️ A prompt injection
| |
An empty result reads as "the ocean was not measured." A refusal that names the gap reads as "we did not ingest that year." Those are different claims, and only one of them is true.
Each of these exists because its absence produced a wrong number during development.
| Guarantee | Why |
|---|---|
| 🎯 Grounding is deterministic | A regex, not a model. The second LLM call is a retry, not a verification pass |
| 🙈 Case banks store no numerals | Digits are masked at write time - the grounding check cannot tell a number the model computed from one it copied out of a remembered answer |
| 📅 Coverage reports years held, not the span | An ingest is not obliged to be contiguous, and this one is not |
| 🧮 Statistics are exact, computed in SQL | rows is a page, not a sample; row_count and stats cover every matching row |
| 🔗 The working set carries its predicate | Downstream operators re-derive it inside Postgres instead of shipping ids out and back |
| ❄️ Cold cache is the headline number | The operator cache defaults off, and benchmark harnesses force it off |
| ⚖️ Cost constants are never hand-set | They are calibrated against the live database |
| 🏷️ Scraped context is labelled, never measured | The ENSO index beside an answer carries its source and fetch date, and appears in no other field |
📖 The stories behind three of these
The database looked slow. Postgres returned 50,000 rows in 152 ms while the operator reported 1,269 ms. The cost was shipping 69,227 ids out to Python and back - which is why the working set carries its predicate instead of its contents.
A calibration reported MAPE 0.0%. That was not a perfect model. The operator cache had
served all five repeats of every query and reported actual = 0 ms for all ten. A default
of true meant a dozen harnesses had to remember to turn it off, and one of them not
remembering was indistinguishable from success. It now defaults to off.
A rollup was "obviously" the fix. Per-profile aggregation measured 142 ms against
123 ms raw - no win at all. profile_measurements is a view; levels are stored as
arrays, one row per profile, and the view UNNESTs them.
- 🛡️ Question arrives at the bridge and is checked by the guard (P1)
- 🗣️ Parser resolves it into typed slots, or refuses an unsupported span
- 💰 Planner enumerates candidate operator DAGs and costs each one
- 🔧 Chosen plan executes as tool calls across the five MCP servers
- 🗄️ Operators push predicates into Postgres rather than shipping row ids
- 🛡️ Retrieved documents are guard-checked before the model sees them (P3)
- ✍️ Synthesis composes prose; every number in it is verified against the result
- 🛡️ Answer is guard-checked on the way out (P2) and returned with plan, timing and sources
docker start vortex-postgres # PostgreSQL 16 + PostGIS
PYTHONPATH=src python -m vortex.bridge.api # bridge + workers, port 8765
cd web/app && npm install && npm run build # tsc --noEmit, then vite buildThen open http://127.0.0.1:8765/ - the bridge serves the built UI from its own origin.
⚠️ There is no venv in this repo andPYTHONPATH=srcis required. A checkout that has never run the frontend build serves a single-file fallback UI instead - it looks fine and lacks every route added since. The bridge logsfrontend_fallbackwhen that happens.
🧪 Verifying a change
PYTHONPATH=src python -m pytest -q # 532 tests
PYTHONPATH=src python -m ruff check src/ scripts/ tests/
cd web/app && npx vitest run # 28 fixture-shape testsLatency work also needs scripts/measure_cost_model.py and scripts/measure_plan_stability.py.
Guard or corpus changes need scripts/measure_guard.py and scripts/attack_adaptive.py.
Rules that are not negotiable - breaking any of these has already produced a wrong number:
- A memory-starved run is not a result. Harnesses flag below ~3.5–4 GB free
- Cold cache is the headline number; report hit rate beside latency, never inside it
- Never hand-set a cost constant - use
scripts/calibrate_cost_model.py - Never load an embedding model in a process holding a psycopg connection
- Credentials never enter the config tree; config is snapshotted to disk
Frontend on Vercel, backend and database together on one EC2 instance:
Browser ──HTTPS──> Vercel ──rewrites──> EC2 :80 nginx ──> uvicorn :8765 ──> Postgres :5433
static build rate limit, 5 workers loopback only
SSE unbuffered
Rewrites keep the browser same-origin, so there is no CORS to configure and no TLS certificate needed on the instance - a direct HTTPS→HTTP call would be blocked as mixed content.
🔒 Postgres must stay on the same box as the bridge. The planner's cost model is calibrated to localhost latency, and a network round-trip corrupts it - so never RDS, Neon or Supabase.
Full runbook: deploy/DEPLOY-EC2.md
| Profiles | 55,795 |
| Region | Indian Ocean |
| Years held | 2019 · 2025–2026 |
| Storage | Level arrays, one row per profile |
Coverage is deliberately not contiguous, and the system reports the years it actually holds rather than the range they span.
⚠️ Benchmark figures in the engineering notes are stale as of 2026-08-20 and are kept only because they remain reproducible from the artifacts that produced them. The provider retired every configured Llama model, and case-based synthesis has since changed the synthesis prompt. Seedocs/OPEN-WORK.md.
| Document | What's in it |
|---|---|
docs/ENGINEERING.md |
The long form - what was measured, what was tried, what the measurements killed |
docs/OPEN-WORK.md |
What is left, and the tried, measured, rejected table |
docs/decisions/ |
28 ADRs - why each thing is the way it is |
docs/API.md |
The bridge's REST surface |
CLAUDE.md |
Working brief for an AI assistant picking this up |
🌊 Try it live · License: MIT © 2026 Team GDHTM
