A deny-by-default scope firewall for customer-facing LLMs, with a deterministic information-flow-control core and an adaptive red-team that proves defense under attack.
Built for the London Cybersecurity Hackathon — AI Security track.
ScopeGuard is not another guardrail wrapper. Pattern-matching defenses lose — the attack surface is infinite. Instead, this system enforces a deterministic policy outside the model: every value carries a trust label, labels propagate, and privileged actions execute only if a capability check passes. Free-text answers additionally pass probabilistic intent and grounding gates. Then an autonomous red-team (
Breaker) attacks both an undefended baseline and the protected system with a static suite and an adaptive mutation loop, producing a before/after attack-success-rate that holds up under adaptation.
Every brand ships an LLM to customers now. Almost none ship a security layer in front of it — and the failure mode isn't hypothetical:
Screenshots of real news coverage (BBC, The Markup, CTV News) of actual chatbot failures — Air Canada was held liable for its bot's bad advice, a Chevrolet dealership's bot "agreed" to sell a car for $1, DPD's bot swore at a customer, and an NYC-backed small-business chatbot told users they could break the law. These are the concrete stakes an unguarded customer-facing LLM creates.
ScopeGuard is a layered reference monitor around a customer-facing LLM, following the 2025–26 security consensus (CaMeL, Microsoft FIDES, PAuth, Costa et al. SaTML 2026). The core differentiator is information-flow-control (IFC) with capabilities: untrusted data is labeled, labels propagate, and privileged actions are gated by deterministic checks, not LLM judges.
┌─────────────── ScopeGuard Gateway (:8000) ──────────────────┐
│ │
User/Breaker ──→ L0 Trust Labeling
│ Every value: TRUSTED | USER | UNTRUSTED
│ │
│ L1 Cheap Deterministic Gates
│ Rate limits, token budgets, LLM Guard input scan
│ │
│ L2 Deterministic Control-Flow Integrity
│ • Dual-LLM action selector (provable subset)
│ • Capability checks (action allowed iff input integrity ≥ required)
│ • No untrusted-to-privileged-sink data flows
│ │
│ L3 Probabilistic Answer Gates (free-text only)
│ • Intent classification (deny-by-default)
│ • Grounding (answer must be supported by corpus + serve intent)
│ • Output scan (leak/PII/off-brand detection)
│ │
│ L4 Audit & Constant Refusal
│ Uniform refusal message (reason logged internally only)
│ Taint traces written to audit log
└────→
Supabase (decision log + traces) ← Next.js Dashboard (Vercel)
Every value entering the system is assigned a label in a small trust lattice (after Denning 1976 / Myers–Liskov 1997):
- TRUSTED — system prompt, policy rules, approved corpus metadata
- USER — the customer's input message
- UNTRUSTED — retrieved KB chunks, external API responses, any untrusted source
Labels propagate: any value derived from UNTRUSTED input inherits UNTRUSTED. This is implemented as a lightweight taint-tracking wrapper (TaintedValue[T]), not a full CaMeL interpreter — but the same principle: untrusted data can never influence program control flow.
Privileged actions (placing an order, reading customer data, emitting an answer) execute only if:
- All input values meet the required integrity level (declared in
policy.yaml) - No higher-confidentiality value flows to a lower-trust sink (exfiltration guard)
Example: place_order requires all order fields to be USER-or-higher integrity. An order assembled from UNTRUSTED KB text is refused deterministically — no LLM judgment involved.
For privileged actions, the bot cannot emit free-form instructions. Instead:
- A Privileged planner (P-LLM) sees only
TRUSTED/USERinputs and produces a plan over a fixed, typed action schema (Gemini function calling with a locked tool list) - A Quarantined model (Q-LLM) processes
UNTRUSTEDcontent and returns typed/symbolic values only (Gemini structured output), never natural-language instructions - The capability check then permits or denies the action deterministically
This gives provable control-flow integrity for the action subset and structurally defeats indirect prompt injection on that path.
Not every reply is a privileged action; menu, allergen, and hours questions are free text. These additionally pass:
- Intent — Gemini judge classifies whether the answer aligns with the user's original request (deny-by-default)
- Grounding — the draft answer must be supported by approved corpus docs AND serve the classified intent (Task-Shield style)
- Output Scan — LLM Guard flags leakage, PII, or off-brand content
These gates are honestly labeled as probabilistic in the demo — the point is that the high-risk path (privileged actions) is deterministic and only the low-risk path is probabilistic.
ScopeGuard includes a fully-functional demo bot called "Not the Greggs" — a self-aware parody name, explicitly not affiliated with the real UK bakery chain. The bot handles customer queries about:
- Ordering sausage rolls, vegan rolls, pastries, and other baked goods
- Checking the menu, allergens, and ingredients
- Store hours and locations
- Complaints and feedback
The bot runs behind ScopeGuard's gateway, demonstrating how attacks fail deterministically while benign requests pass through and receive sensible answers.
Tested locally end-to-end with real credentials (not yet deployed to Vercel/Modal):
- Gemini (a rolling
-latestalias, e.g.gemini-flash-latest) for the planner, judge, and quarantine model — pinned dated model IDs (gemini-2.5-flash, etc.) can 404 for new API keys once Google deprecates them, so a rolling alias is the safer default - Supabase (Postgres) for audit logging and taint traces
- Next.js dashboard, run locally with
npm run dev - A "ScopeGuard: ON/OFF" toggle on
/demothat replays the same message against the undefended baseline (/chat/baseline, demo-only, bypasses every gate) and the protected gateway (/chat) side by side
Results from a real run against the static attack suite (20 attacks, real Gemini judgment, not the offline fake-LLM used in pytest):
- Baseline ASR: varied 5–15% across runs (the undefended bot is non-deterministic; some attacks it resists on its own, most it doesn't)
- Protected ASR: 0/20
- Adaptive ASR: 0/20 after mutation rounds, in the run that also hit 0/20 protected
- Benign false-positive rate: ~10% (down from an initial 30% after fixing a missing intent category, a retrieval gap, and grounding being wrongly applied to transactional replies — see git history; residual FPR is inherent LLM judge variance, not a known bug)
- Python 3.11+
- Node.js 18+ (for the dashboard)
- Google Gemini API key (from Google AI Studio)
- Supabase account (optional, for audit logging and dashboard; tests run without it)
Copy .env.example to .env and fill in your credentials:
cp .env.example .envRequired variables:
| Variable | Purpose | Where to get it |
|---|---|---|
GEMINI_API_KEY |
Your Gemini API key | Google AI Studio |
JUDGE_MODEL |
Gemini Flash-tier model ID for judges | Google AI Studio -- use a rolling alias like gemini-flash-latest rather than a pinned dated ID, which Google can deprecate for new keys |
TARGET_MODEL |
Gemini model ID for the bot | Same as above, e.g. gemini-flash-latest |
SUPABASE_URL |
(Optional) Supabase project URL | Supabase dashboard |
SUPABASE_KEY |
(Optional) Supabase service role key | Supabase dashboard |
NEXT_PUBLIC_SUPABASE_URL |
(Optional) Supabase public URL | Supabase dashboard |
NEXT_PUBLIC_SUPABASE_ANON_KEY |
(Optional) Supabase anon key | Supabase dashboard |
-
Install dependencies:
pip install -e ".[dev]" -
Start the gateway:
uvicorn scopeguard.gateway:app --reload --port 8000
The gateway is now running at
http://localhost:8000. Check/healthto confirm.
Run the headline red-team evaluation (static + adaptive attack suite, baseline vs. protected, FPR):
python scripts/run_eval.pyThis outputs:
- Attack-success-rate (ASR) on the undefended baseline
- ASR on the protected gateway (static attacks)
- ASR after adaptive mutation rounds
- False-positive rate on benign requests
- Coverage of deterministic enforcement
The results are also written to Supabase (if configured) for the dashboard.
-
Install dependencies:
cd web npm install -
Start the development server:
npm run dev
Open
http://localhost:3000/demoin your browser. You'll see:- A full-screen chat interface with the "Not the Greggs" bot
- A toggle switch to flip between undefended and ScopeGuard-protected modes
- Side-by-side comparison of the bot's responses and which defense layer (if any) caught the attack
All unit tests run offline using a fake-Gemini fixture:
pytest -q # Run all tests
pytest tests/test_ifc_labels.py # Test taint propagation (IFC invariants)
pytest tests/test_capabilities.py # Test capability checks + exfiltration guard
pytest -v # Verbose outputThe fake LLM is deterministic and controlled via tests/conftest.py, so no API keys or network access is needed. This allows CI/CD to run the full test suite in GitHub Actions.
scopeguard/
├── README.md # You are here
├── pyproject.toml
├── .env.example
├── docker-compose.yml
├── Dockerfile
├── policy.yaml # Intents, capabilities, integrity requirements, refusal message
├── modal_app.py # Serverless deployment (Modal)
│
├── src/scopeguard/
│ ├── config.py # Loads policy.yaml + env vars
│ ├── models.py # Pydantic: RequestContext, Decision, TaintedValue, CapabilityPolicy
│ ├── llm_client.py # Google Gemini wrapper (all model calls route here)
│ ├── gateway.py # FastAPI reverse proxy + layered pipeline.
│ │ # /chat = protected; /chat/baseline = demo-only,
│ │ # bypasses every gate for live before/after comparison
│ │
│ ├── ifc/ # ← Information-flow-control core
│ │ ├── labels.py # Label lattice + TaintedValue[T] + propagation
│ │ └── capabilities.py # Capability checks, exfiltration guard, Trusted Action
│ │
│ ├── planner/ # ← Deterministic control-flow core
│ │ ├── action_selector.py # P-LLM plan over fixed typed actions (Gemini function calling)
│ │ ├── quarantine.py # Q-LLM typed extraction from untrusted KB (structured output)
│ │ └── actions.py # Locked action schema (place_order, lookup_menu, etc.)
│ │
│ ├── gates/ # Layered defense gates (L0–L4)
│ │ ├── base.py # Gate ABC + pipeline runner
│ │ ├── rate_cost.py # L1: Rate limits, token budgets
│ │ ├── input_scan.py # L1: LLM Guard + Gemini safety checks
│ │ ├── intent.py # L3: Intent classification (deny-by-default)
│ │ ├── grounding.py # L3: Task-Shield grounding
│ │ └── output_scan.py # L3: Output leak/PII/off-brand detection
│ │
│ ├── rag/ # Retrieval-augmented generation (minimal)
│ │ ├── retriever.py # Tiny retriever over kb/ (labels output UNTRUSTED)
│ │ └── spotlight.py # Datamarking/spotlighting for untrusted content
│ │
│ ├── bot/
│ │ └── sausagebot.py # The deliberately-vulnerable demo bot, branded "Not the Greggs"
│ │
│ ├── audit/
│ │ └── logger.py # Decision + taint trace → Supabase (PII redacted)
│ │
│ └── redteam/ # ← Autonomous red-team framework
│ ├── attacks.py # Static 20-attack suite (data)
│ ├── multiturn.py # Multi-turn / branch-steering attacks
│ ├── adaptive.py # Mutation loop (PISmith-style) over failed attacks
│ ├── breaker.py # LangGraph runner: baseline vs. protected
│ └── report.py # ASR / adaptive-ASR / FPR / coverage reporting
│
├── kb/ # Knowledge base corpus
│ ├── menu.md
│ ├── policy.md
│ ├── stores.md
│ └── poisoned_menu.md # Indirect-injection test corpus
│
├── web/ # Next.js 14 dashboard (App Router)
│ ├── package.json
│ ├── app/
│ │ ├── layout.tsx
│ │ ├── page.tsx # Analytics dashboard (dark "reference monitor" theme)
│ │ ├── demo/page.tsx # "Not the Greggs" live chat demo -- its own light
│ │ │ # bakery theme, ScopeGuard ON/OFF toggle, product list
│ │ ├── api/demo-chat/ # Server-side proxy from the demo page to the gateway
│ │ └── trace/[id]/page.tsx # Per-request taint trace viewer
│ ├── components/ # Custom Tailwind components (StatCard, AsrChart/Recharts)
│ ├── lib/supabase.ts # Supabase JS client
│ └── tailwind.config.ts # Two palettes: dashboard dark theme + demo-page bakery theme
│
├── tests/
│ ├── conftest.py # Fake-Gemini fixture (deterministic, no network)
│ ├── benign_set.py # Golden benign requests (FPR testing)
│ ├── test_ifc_labels.py # Taint propagation invariants
│ ├── test_capabilities.py # Trusted Action + exfiltration guard
│ ├── test_action_selector.py # Planner emits only typed actions
│ ├── test_rate_cost.py # L1 rate limiting
│ ├── test_input_scan.py # L1 input scanning
│ ├── test_intent.py # L3 intent gate
│ ├── test_grounding.py # L3 grounding gate
│ ├── test_output_scan.py # L3 output scanning
│ └── test_gateway_integration.py # End-to-end pipeline
│
├── scripts/
│ ├── run_baseline.py # Baseline bot attack evaluation
│ ├── run_protected.py # Protected gateway attack evaluation
│ └── run_eval.py # Full eval (static + adaptive + FPR + coverage)
│
└── .github/workflows/ci.yml # GitHub Actions (pytest, no keys)
The policy.yaml file declares all security and capability rules:
brand: "Not the Greggs"
allowed_intents: [order, menu, hours, allergen, store, complaint, smalltalk]
# Deterministic control-flow: privileged actions + integrity requirements
capabilities:
- action: place_order
min_input_integrity: user # Order fields must be USER-or-higher, never UNTRUSTED
allowed_sinks: [order_system]
- action: lookup_menu
min_input_integrity: untrusted # Menu lookup can read the UNTRUSTED KB
allowed_sinks: [user_reply]
# Grounding rules for probabilistic gates
grounding:
require_source: true
approved_corpus: [kb/menu.md, kb/policy.md, kb/stores.md]
min_support: 0.6
# Data classes to block in output scan
blocked_data_classes: [pii, credentials, internal_pricing, other_customer_data]
# Limits and quotas
limits:
max_tokens_per_session: 4000
max_requests_per_minute: 10
max_output_tokens: 500
# Uniform refusal (same for every denial — mitigates denial-feedback leakage)
refusal_message: >-
I can only help with Not the Greggs orders, our menu, allergens, opening hours,
and store info — I can't help with that one!
# LLM config
judge:
model: "${JUDGE_MODEL}"
temperature: 0.0
target:
model: "${TARGET_MODEL}"
temperature: 0.3Edit this file to:
- Add new allowed intents or actions
- Tighten integrity requirements on existing actions
- Expand the approved corpus
- Adjust rate limits and token budgets
- Customize the refusal message (keeping it uniform for all denials)
Every design decision in ScopeGuard implements a lightweight version of a named 2025–26 research result. This is what makes it research-grade, not just another guardrail demo.
| Component | Backing Work | What We Implement |
|---|---|---|
| Trust labels + taint propagation | CaMeL (arXiv:2503.18813); Denning lattice 1976; Myers–Liskov DLM 1997 | TaintedValue[T] lattice; untrusted-derived values stay untrusted |
| Capability checks / Trusted Action | Microsoft FIDES; CaMeL capabilities; PAuth (arXiv:2603.17170); Costa et al. SaTML 2026 (arXiv:2505.23643) | Action allowed iff input integrity ≥ required; exfiltration guard |
| Dual-LLM Action-Selector | Willison Dual-LLM 2023; CaMeL; Single-Shot Planning (arXiv:2601.09923) | P-LLM plans over trusted query; Q-LLM returns typed values via structured output → provable CFI |
| Grounding / task alignment | Task Shield, Jia et al., ACL 2025 | Draft must serve the user's original in-scope task |
| Spotlighting of untrusted content | Spotlighting, Microsoft, arXiv:2403.14720 | Datamark untrusted KB in prompts |
| Design patterns | Design Patterns for Securing LLM Agents, Beurer-Kellner et al., arXiv:2506.08837 | Action-Selector + Context-Minimization + Dual-LLM patterns |
| Adaptive red-team | PISmith RL red-teaming (arXiv:2603.13026); Adaptive Evaluation (arXiv:2606.26479) | Mutation loop over failed attacks; report adaptive ASR |
| Side-channel-aware refusal | Denial-feedback leakage (arXiv:2604.04035); Operationalizing CaMeL | Uniform constant refusal; reason logged internally |
Each of these is named as "future work" in the papers above. We structure the code so they're additive:
- Formal verification of the policy layer — Keep
policy.yamldeclarative so a solver could later prove "noUNTRUSTEDvalue reachesplace_order" - Side-channel hardening — Result-types instead of exceptions, uniform refusal (done); constant-time execution and jitter (roadmap)
- Branch-steering defence — Single-shot planning mitigates the basic case; full multi-turn branch-steering is roadmap; we include the attack in
Breakerto show residual risk honestly - Certified robustness bounds — Move from "0/20 on our suite" toward provable ASR bounds on the deterministic subset (field direction per Costa et al.)
- Interpretability-based detection — Activation/refusal methods (COSMIC, ACL 2025) as a future L1 gate
- Transactional rollback safety — (SAFEFLOW, arXiv:2506.07564) for atomic ordering operations
docker-compose upThis brings up the gateway on :8000 and (if configured) the dashboard backend on :3000.
Deploy the gateway to Modal:
modal deploy modal_app.pyThis runs the FastAPI app and the red-team evaluation on Modal's serverless compute.
Deploy the Next.js frontend to Vercel:
cd web
vercel deployOr use the Vercel GitHub integration for automatic deployments.
web/app/demo/page.tsx — a menu panel mirroring kb/menu.md next to the chat, with preset buttons pulled straight from the real attack suite (redteam/attacks.py) and benign set. The shield toggle replays the same message against /chat/baseline (undefended) or /chat (protected) so you can watch the same input succeed, then get blocked, without leaving the page.
web/app/page.tsx — reads live from Supabase's decisions and reports tables (written by the gateway's audit logger and scripts/run_eval.py respectively). The numbers above are from a real run against real Gemini judgment, not the offline fake-LLM used in pytest.
To reproduce this yourself: start the backend (uvicorn scopeguard.gateway:app --reload), start the frontend (cd web && npm run dev), open http://localhost:3000/demo for the chat or http://localhost:3000 for the dashboard.
GitHub Actions runs the full test suite on every push:
pytest -qThe CI environment has no API keys — all tests use the fake-Gemini fixture and run offline. This ensures the core IFC invariants and capability checks are green before any live-model testing.
- Python 3.11+, fully typed with
pydantic v2 - FastAPI for the gateway, async throughout
- Google Gemini via
google-genaiSDK; all calls route throughsrc/scopeguard/llm_client.py - Structured output (
response_schema) for judges and Q-LLM extraction - Function calling for the P-LLM action-selector
- No secrets in code/logs; redact PII before audit writes; load config via
config.pyonly - Deny-by-default; any exception → deny + log (never fail open)
- Result types, not exceptions on the decision path (side-channel hygiene)
Linting & formatting:
ruff check .
black --check .Explicitly out of scope to keep the hackathon build focused and clean:
- No fine-tuning of the LLM
- No full CaMeL interpreter (only lightweight
TaintedValue[T]) - No multi-tenant or authentication (single bot per deployment)
- No general-purpose policy engine (only
policy.yamlfor one bot) - No second bot or ensemble methods
- No custom vector DB (minimal RAG only)
- No timing-channel or side-channel hardening beyond result-types and uniform refusal (roadmap)
- No formal verification toolchain (roadmap — but
policy.yamlis designed to be checkable)
If tempted to add any of these, stop and ship the before/after attack-success-rate number first.
This is a hackathon project. If you're building on it:
- This README is the authoritative architecture/research doc -- the original internal build spec has been folded in above and removed from the repo
- Follow test-driven development: write
test_*.pyfirst, then implement - Keep the IFC core (labels, capability checks, action-selector) clean and small — this is the differentiator
- Use the fake-Gemini fixture in tests; never require live API keys for
pytest - Red-team your own changes — add new attacks to
attacks.pyif you add new functionality
Hackathon project. Please refer to any accompanying LICENSE file in the repository.
- Architecture, research grounding, and roadmap: this README (Architecture + Research Grounding sections above)
- Demo: Run
http://localhost:3000/demoafter starting the backend and frontend - Evaluation:
python scripts/run_eval.pyfor before/after ASR, adaptive rounds, and FPR
Built during the London Cybersecurity Hackathon, AI Security track.


