Skip to content

Repository files navigation

ScopeGuard

A deny-by-default scope firewall for customer-facing LLMs, with a deterministic information-flow-control core and an adaptive red-team that proves defense under attack.

Built for the London Cybersecurity Hackathon — AI Security track.

ScopeGuard is not another guardrail wrapper. Pattern-matching defenses lose — the attack surface is infinite. Instead, this system enforces a deterministic policy outside the model: every value carries a trust label, labels propagate, and privileged actions execute only if a capability check passes. Free-text answers additionally pass probabilistic intent and grounding gates. Then an autonomous red-team (Breaker) attacks both an undefended baseline and the protected system with a static suite and an adaptive mutation loop, producing a before/after attack-success-rate that holds up under adaptation.

Why this matters

Every brand ships an LLM to customers now. Almost none ship a security layer in front of it — and the failure mode isn't hypothetical:

Real chatbot incidents: Air Canada ordered to pay a customer misled by its chatbot, a DPD bot swearing at a customer, a Chevrolet dealer bot agreeing to sell a car for $1, and an NYC city chatbot telling businesses to break the law

Screenshots of real news coverage (BBC, The Markup, CTV News) of actual chatbot failures — Air Canada was held liable for its bot's bad advice, a Chevrolet dealership's bot "agreed" to sell a car for $1, DPD's bot swore at a customer, and an NYC-backed small-business chatbot told users they could break the law. These are the concrete stakes an unguarded customer-facing LLM creates.


Architecture

ScopeGuard is a layered reference monitor around a customer-facing LLM, following the 2025–26 security consensus (CaMeL, Microsoft FIDES, PAuth, Costa et al. SaTML 2026). The core differentiator is information-flow-control (IFC) with capabilities: untrusted data is labeled, labels propagate, and privileged actions are gated by deterministic checks, not LLM judges.

                  ┌─────────────── ScopeGuard Gateway (:8000) ──────────────────┐
                  │                                                              │
   User/Breaker ──→ L0 Trust Labeling
                  │    Every value: TRUSTED | USER | UNTRUSTED
                  │    │
                  │ L1 Cheap Deterministic Gates
                  │    Rate limits, token budgets, LLM Guard input scan
                  │    │
                  │ L2 Deterministic Control-Flow Integrity
                  │    • Dual-LLM action selector (provable subset)
                  │    • Capability checks (action allowed iff input integrity ≥ required)
                  │    • No untrusted-to-privileged-sink data flows
                  │    │
                  │ L3 Probabilistic Answer Gates (free-text only)
                  │    • Intent classification (deny-by-default)
                  │    • Grounding (answer must be supported by corpus + serve intent)
                  │    • Output scan (leak/PII/off-brand detection)
                  │    │
                  │ L4 Audit & Constant Refusal
                  │    Uniform refusal message (reason logged internally only)
                  │    Taint traces written to audit log
                  └────→
              Supabase (decision log + traces) ← Next.js Dashboard (Vercel)

Trust Labels & Taint Propagation

Every value entering the system is assigned a label in a small trust lattice (after Denning 1976 / Myers–Liskov 1997):

  • TRUSTED — system prompt, policy rules, approved corpus metadata
  • USER — the customer's input message
  • UNTRUSTED — retrieved KB chunks, external API responses, any untrusted source

Labels propagate: any value derived from UNTRUSTED input inherits UNTRUSTED. This is implemented as a lightweight taint-tracking wrapper (TaintedValue[T]), not a full CaMeL interpreter — but the same principle: untrusted data can never influence program control flow.

Capability Checks (Deterministic Enforcement)

Privileged actions (placing an order, reading customer data, emitting an answer) execute only if:

  1. All input values meet the required integrity level (declared in policy.yaml)
  2. No higher-confidentiality value flows to a lower-trust sink (exfiltration guard)

Example: place_order requires all order fields to be USER-or-higher integrity. An order assembled from UNTRUSTED KB text is refused deterministically — no LLM judgment involved.

Dual-LLM Action-Selector (Provable Control Flow)

For privileged actions, the bot cannot emit free-form instructions. Instead:

  • A Privileged planner (P-LLM) sees only TRUSTED/USER inputs and produces a plan over a fixed, typed action schema (Gemini function calling with a locked tool list)
  • A Quarantined model (Q-LLM) processes UNTRUSTED content and returns typed/symbolic values only (Gemini structured output), never natural-language instructions
  • The capability check then permits or denies the action deterministically

This gives provable control-flow integrity for the action subset and structurally defeats indirect prompt injection on that path.

Probabilistic Gates (Free-Text Answers)

Not every reply is a privileged action; menu, allergen, and hours questions are free text. These additionally pass:

  1. Intent — Gemini judge classifies whether the answer aligns with the user's original request (deny-by-default)
  2. Grounding — the draft answer must be supported by approved corpus docs AND serve the classified intent (Task-Shield style)
  3. Output Scan — LLM Guard flags leakage, PII, or off-brand content

These gates are honestly labeled as probabilistic in the demo — the point is that the high-risk path (privileged actions) is deterministic and only the low-risk path is probabilistic.


The Demo Bot: "Not the Greggs"

ScopeGuard includes a fully-functional demo bot called "Not the Greggs" — a self-aware parody name, explicitly not affiliated with the real UK bakery chain. The bot handles customer queries about:

  • Ordering sausage rolls, vegan rolls, pastries, and other baked goods
  • Checking the menu, allergens, and ingredients
  • Store hours and locations
  • Complaints and feedback

The bot runs behind ScopeGuard's gateway, demonstrating how attacks fail deterministically while benign requests pass through and receive sensible answers.

Real Deployment

Tested locally end-to-end with real credentials (not yet deployed to Vercel/Modal):

  • Gemini (a rolling -latest alias, e.g. gemini-flash-latest) for the planner, judge, and quarantine model — pinned dated model IDs (gemini-2.5-flash, etc.) can 404 for new API keys once Google deprecates them, so a rolling alias is the safer default
  • Supabase (Postgres) for audit logging and taint traces
  • Next.js dashboard, run locally with npm run dev
  • A "ScopeGuard: ON/OFF" toggle on /demo that replays the same message against the undefended baseline (/chat/baseline, demo-only, bypasses every gate) and the protected gateway (/chat) side by side

Results from a real run against the static attack suite (20 attacks, real Gemini judgment, not the offline fake-LLM used in pytest):

  • Baseline ASR: varied 5–15% across runs (the undefended bot is non-deterministic; some attacks it resists on its own, most it doesn't)
  • Protected ASR: 0/20
  • Adaptive ASR: 0/20 after mutation rounds, in the run that also hit 0/20 protected
  • Benign false-positive rate: ~10% (down from an initial 30% after fixing a missing intent category, a retrieval gap, and grounding being wrongly applied to transactional replies — see git history; residual FPR is inherent LLM judge variance, not a known bug)

Quick Start

Prerequisites

  • Python 3.11+
  • Node.js 18+ (for the dashboard)
  • Google Gemini API key (from Google AI Studio)
  • Supabase account (optional, for audit logging and dashboard; tests run without it)

Environment Setup

Copy .env.example to .env and fill in your credentials:

cp .env.example .env

Required variables:

Variable Purpose Where to get it
GEMINI_API_KEY Your Gemini API key Google AI Studio
JUDGE_MODEL Gemini Flash-tier model ID for judges Google AI Studio -- use a rolling alias like gemini-flash-latest rather than a pinned dated ID, which Google can deprecate for new keys
TARGET_MODEL Gemini model ID for the bot Same as above, e.g. gemini-flash-latest
SUPABASE_URL (Optional) Supabase project URL Supabase dashboard
SUPABASE_KEY (Optional) Supabase service role key Supabase dashboard
NEXT_PUBLIC_SUPABASE_URL (Optional) Supabase public URL Supabase dashboard
NEXT_PUBLIC_SUPABASE_ANON_KEY (Optional) Supabase anon key Supabase dashboard

Local Backend

  1. Install dependencies:

    pip install -e ".[dev]"
  2. Start the gateway:

    uvicorn scopeguard.gateway:app --reload --port 8000

    The gateway is now running at http://localhost:8000. Check /health to confirm.

Evaluation Script

Run the headline red-team evaluation (static + adaptive attack suite, baseline vs. protected, FPR):

python scripts/run_eval.py

This outputs:

  • Attack-success-rate (ASR) on the undefended baseline
  • ASR on the protected gateway (static attacks)
  • ASR after adaptive mutation rounds
  • False-positive rate on benign requests
  • Coverage of deterministic enforcement

The results are also written to Supabase (if configured) for the dashboard.

Dashboard (Next.js)

  1. Install dependencies:

    cd web
    npm install
  2. Start the development server:

    npm run dev

    Open http://localhost:3000/demo in your browser. You'll see:

    • A full-screen chat interface with the "Not the Greggs" bot
    • A toggle switch to flip between undefended and ScopeGuard-protected modes
    • Side-by-side comparison of the bot's responses and which defense layer (if any) caught the attack

Tests (No API Keys Required)

All unit tests run offline using a fake-Gemini fixture:

pytest -q                  # Run all tests
pytest tests/test_ifc_labels.py  # Test taint propagation (IFC invariants)
pytest tests/test_capabilities.py  # Test capability checks + exfiltration guard
pytest -v                  # Verbose output

The fake LLM is deterministic and controlled via tests/conftest.py, so no API keys or network access is needed. This allows CI/CD to run the full test suite in GitHub Actions.


Code Structure

scopeguard/
├── README.md                    # You are here
├── pyproject.toml
├── .env.example
├── docker-compose.yml
├── Dockerfile
├── policy.yaml                  # Intents, capabilities, integrity requirements, refusal message
├── modal_app.py                 # Serverless deployment (Modal)
│
├── src/scopeguard/
│   ├── config.py                # Loads policy.yaml + env vars
│   ├── models.py                # Pydantic: RequestContext, Decision, TaintedValue, CapabilityPolicy
│   ├── llm_client.py            # Google Gemini wrapper (all model calls route here)
│   ├── gateway.py               # FastAPI reverse proxy + layered pipeline.
│   │                            #   /chat = protected; /chat/baseline = demo-only,
│   │                            #   bypasses every gate for live before/after comparison
│   │
│   ├── ifc/                     # ← Information-flow-control core
│   │   ├── labels.py            # Label lattice + TaintedValue[T] + propagation
│   │   └── capabilities.py      # Capability checks, exfiltration guard, Trusted Action
│   │
│   ├── planner/                 # ← Deterministic control-flow core
│   │   ├── action_selector.py   # P-LLM plan over fixed typed actions (Gemini function calling)
│   │   ├── quarantine.py        # Q-LLM typed extraction from untrusted KB (structured output)
│   │   └── actions.py           # Locked action schema (place_order, lookup_menu, etc.)
│   │
│   ├── gates/                   # Layered defense gates (L0–L4)
│   │   ├── base.py              # Gate ABC + pipeline runner
│   │   ├── rate_cost.py         # L1: Rate limits, token budgets
│   │   ├── input_scan.py        # L1: LLM Guard + Gemini safety checks
│   │   ├── intent.py            # L3: Intent classification (deny-by-default)
│   │   ├── grounding.py         # L3: Task-Shield grounding
│   │   └── output_scan.py       # L3: Output leak/PII/off-brand detection
│   │
│   ├── rag/                     # Retrieval-augmented generation (minimal)
│   │   ├── retriever.py         # Tiny retriever over kb/ (labels output UNTRUSTED)
│   │   └── spotlight.py         # Datamarking/spotlighting for untrusted content
│   │
│   ├── bot/
│   │   └── sausagebot.py        # The deliberately-vulnerable demo bot, branded "Not the Greggs"
│   │
│   ├── audit/
│   │   └── logger.py            # Decision + taint trace → Supabase (PII redacted)
│   │
│   └── redteam/                 # ← Autonomous red-team framework
│       ├── attacks.py           # Static 20-attack suite (data)
│       ├── multiturn.py         # Multi-turn / branch-steering attacks
│       ├── adaptive.py          # Mutation loop (PISmith-style) over failed attacks
│       ├── breaker.py           # LangGraph runner: baseline vs. protected
│       └── report.py            # ASR / adaptive-ASR / FPR / coverage reporting
│
├── kb/                          # Knowledge base corpus
│   ├── menu.md
│   ├── policy.md
│   ├── stores.md
│   └── poisoned_menu.md         # Indirect-injection test corpus
│
├── web/                         # Next.js 14 dashboard (App Router)
│   ├── package.json
│   ├── app/
│   │   ├── layout.tsx
│   │   ├── page.tsx             # Analytics dashboard (dark "reference monitor" theme)
│   │   ├── demo/page.tsx        # "Not the Greggs" live chat demo -- its own light
│   │   │                        #   bakery theme, ScopeGuard ON/OFF toggle, product list
│   │   ├── api/demo-chat/       # Server-side proxy from the demo page to the gateway
│   │   └── trace/[id]/page.tsx  # Per-request taint trace viewer
│   ├── components/              # Custom Tailwind components (StatCard, AsrChart/Recharts)
│   ├── lib/supabase.ts          # Supabase JS client
│   └── tailwind.config.ts       # Two palettes: dashboard dark theme + demo-page bakery theme
│
├── tests/
│   ├── conftest.py              # Fake-Gemini fixture (deterministic, no network)
│   ├── benign_set.py            # Golden benign requests (FPR testing)
│   ├── test_ifc_labels.py       # Taint propagation invariants
│   ├── test_capabilities.py     # Trusted Action + exfiltration guard
│   ├── test_action_selector.py  # Planner emits only typed actions
│   ├── test_rate_cost.py        # L1 rate limiting
│   ├── test_input_scan.py       # L1 input scanning
│   ├── test_intent.py           # L3 intent gate
│   ├── test_grounding.py        # L3 grounding gate
│   ├── test_output_scan.py      # L3 output scanning
│   └── test_gateway_integration.py  # End-to-end pipeline
│
├── scripts/
│   ├── run_baseline.py          # Baseline bot attack evaluation
│   ├── run_protected.py         # Protected gateway attack evaluation
│   └── run_eval.py              # Full eval (static + adaptive + FPR + coverage)
│
└── .github/workflows/ci.yml     # GitHub Actions (pytest, no keys)

Policy Configuration

The policy.yaml file declares all security and capability rules:

brand: "Not the Greggs"
allowed_intents: [order, menu, hours, allergen, store, complaint, smalltalk]

# Deterministic control-flow: privileged actions + integrity requirements
capabilities:
  - action: place_order
    min_input_integrity: user       # Order fields must be USER-or-higher, never UNTRUSTED
    allowed_sinks: [order_system]
  - action: lookup_menu
    min_input_integrity: untrusted  # Menu lookup can read the UNTRUSTED KB
    allowed_sinks: [user_reply]

# Grounding rules for probabilistic gates
grounding:
  require_source: true
  approved_corpus: [kb/menu.md, kb/policy.md, kb/stores.md]
  min_support: 0.6

# Data classes to block in output scan
blocked_data_classes: [pii, credentials, internal_pricing, other_customer_data]

# Limits and quotas
limits:
  max_tokens_per_session: 4000
  max_requests_per_minute: 10
  max_output_tokens: 500

# Uniform refusal (same for every denial — mitigates denial-feedback leakage)
refusal_message: >-
  I can only help with Not the Greggs orders, our menu, allergens, opening hours,
  and store info — I can't help with that one!

# LLM config
judge:
  model: "${JUDGE_MODEL}"
  temperature: 0.0
target:
  model: "${TARGET_MODEL}"
  temperature: 0.3

Edit this file to:

  • Add new allowed intents or actions
  • Tighten integrity requirements on existing actions
  • Expand the approved corpus
  • Adjust rate limits and token budgets
  • Customize the refusal message (keeping it uniform for all denials)

Research Grounding

Every design decision in ScopeGuard implements a lightweight version of a named 2025–26 research result. This is what makes it research-grade, not just another guardrail demo.

Implemented

Component Backing Work What We Implement
Trust labels + taint propagation CaMeL (arXiv:2503.18813); Denning lattice 1976; Myers–Liskov DLM 1997 TaintedValue[T] lattice; untrusted-derived values stay untrusted
Capability checks / Trusted Action Microsoft FIDES; CaMeL capabilities; PAuth (arXiv:2603.17170); Costa et al. SaTML 2026 (arXiv:2505.23643) Action allowed iff input integrity ≥ required; exfiltration guard
Dual-LLM Action-Selector Willison Dual-LLM 2023; CaMeL; Single-Shot Planning (arXiv:2601.09923) P-LLM plans over trusted query; Q-LLM returns typed values via structured output → provable CFI
Grounding / task alignment Task Shield, Jia et al., ACL 2025 Draft must serve the user's original in-scope task
Spotlighting of untrusted content Spotlighting, Microsoft, arXiv:2403.14720 Datamark untrusted KB in prompts
Design patterns Design Patterns for Securing LLM Agents, Beurer-Kellner et al., arXiv:2506.08837 Action-Selector + Context-Minimization + Dual-LLM patterns
Adaptive red-team PISmith RL red-teaming (arXiv:2603.13026); Adaptive Evaluation (arXiv:2606.26479) Mutation loop over failed attacks; report adaptive ASR
Side-channel-aware refusal Denial-feedback leakage (arXiv:2604.04035); Operationalizing CaMeL Uniform constant refusal; reason logged internally

Roadmap (Honest Scope Notes)

Each of these is named as "future work" in the papers above. We structure the code so they're additive:

  1. Formal verification of the policy layer — Keep policy.yaml declarative so a solver could later prove "no UNTRUSTED value reaches place_order"
  2. Side-channel hardening — Result-types instead of exceptions, uniform refusal (done); constant-time execution and jitter (roadmap)
  3. Branch-steering defence — Single-shot planning mitigates the basic case; full multi-turn branch-steering is roadmap; we include the attack in Breaker to show residual risk honestly
  4. Certified robustness bounds — Move from "0/20 on our suite" toward provable ASR bounds on the deterministic subset (field direction per Costa et al.)
  5. Interpretability-based detection — Activation/refusal methods (COSMIC, ACL 2025) as a future L1 gate
  6. Transactional rollback safety — (SAFEFLOW, arXiv:2506.07564) for atomic ordering operations

Deployment

Local Docker

docker-compose up

This brings up the gateway on :8000 and (if configured) the dashboard backend on :3000.

Serverless (Modal)

Deploy the gateway to Modal:

modal deploy modal_app.py

This runs the FastAPI app and the red-team evaluation on Modal's serverless compute.

Dashboard (Vercel)

Deploy the Next.js frontend to Vercel:

cd web
vercel deploy

Or use the Vercel GitHub integration for automatic deployments.


Screenshots & Demo

"Not the Greggs" — the live demo page

The "Not the Greggs" demo chat page: a shop-front product listing (sausage rolls, bakes, doughnuts, coffee) beside a chat panel with a ScopeGuard ON/OFF toggle

web/app/demo/page.tsx — a menu panel mirroring kb/menu.md next to the chat, with preset buttons pulled straight from the real attack suite (redteam/attacks.py) and benign set. The shield toggle replays the same message against /chat/baseline (undefended) or /chat (protected) so you can watch the same input succeed, then get blocked, without leaving the page.

The dashboard

The ScopeGuard analytics dashboard: ASR baseline 3/20, ASR protected 0/20, ASR adaptive 0/20, benign pass 9/10, an attack-success-rate bar chart, a blocked-by-category breakdown, and a table of recent decisions with per-request trace links

web/app/page.tsx — reads live from Supabase's decisions and reports tables (written by the gateway's audit logger and scripts/run_eval.py respectively). The numbers above are from a real run against real Gemini judgment, not the offline fake-LLM used in pytest.

To reproduce this yourself: start the backend (uvicorn scopeguard.gateway:app --reload), start the frontend (cd web && npm run dev), open http://localhost:3000/demo for the chat or http://localhost:3000 for the dashboard.


CI/CD

GitHub Actions runs the full test suite on every push:

pytest -q

The CI environment has no API keys — all tests use the fake-Gemini fixture and run offline. This ensures the core IFC invariants and capability checks are green before any live-model testing.


Coding Standards

  • Python 3.11+, fully typed with pydantic v2
  • FastAPI for the gateway, async throughout
  • Google Gemini via google-genai SDK; all calls route through src/scopeguard/llm_client.py
  • Structured output (response_schema) for judges and Q-LLM extraction
  • Function calling for the P-LLM action-selector
  • No secrets in code/logs; redact PII before audit writes; load config via config.py only
  • Deny-by-default; any exception → deny + log (never fail open)
  • Result types, not exceptions on the decision path (side-channel hygiene)

Linting & formatting:

ruff check .
black --check .

Non-Goals (Scope Discipline)

Explicitly out of scope to keep the hackathon build focused and clean:

  • No fine-tuning of the LLM
  • No full CaMeL interpreter (only lightweight TaintedValue[T])
  • No multi-tenant or authentication (single bot per deployment)
  • No general-purpose policy engine (only policy.yaml for one bot)
  • No second bot or ensemble methods
  • No custom vector DB (minimal RAG only)
  • No timing-channel or side-channel hardening beyond result-types and uniform refusal (roadmap)
  • No formal verification toolchain (roadmap — but policy.yaml is designed to be checkable)

If tempted to add any of these, stop and ship the before/after attack-success-rate number first.


Contributing

This is a hackathon project. If you're building on it:

  1. This README is the authoritative architecture/research doc -- the original internal build spec has been folded in above and removed from the repo
  2. Follow test-driven development: write test_*.py first, then implement
  3. Keep the IFC core (labels, capability checks, action-selector) clean and small — this is the differentiator
  4. Use the fake-Gemini fixture in tests; never require live API keys for pytest
  5. Red-team your own changes — add new attacks to attacks.py if you add new functionality

License

Hackathon project. Please refer to any accompanying LICENSE file in the repository.


Questions?

  • Architecture, research grounding, and roadmap: this README (Architecture + Research Grounding sections above)
  • Demo: Run http://localhost:3000/demo after starting the backend and frontend
  • Evaluation: python scripts/run_eval.py for before/after ASR, adaptive rounds, and FPR

Built during the London Cybersecurity Hackathon, AI Security track.

About

Hackathon London

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages