Skip to content

feat: multi-agent demo behind one Sentinel gateway - #8

Merged
Chirudeva-Reddy merged 1 commit into
mainfrom
feat/multi-agent-demo
Sep 24, 2026
Merged

Chirudeva-Reddy merged 1 commit into
mainfrom
feat/multi-agent-demo

Conversation

@Chirudeva-Reddy

Copy link
Copy Markdown
Owner
  • sentinel demo [--offline] [--approve-escalations]: three agents on the Anthropic SDK
    (claude-opus-5, manual tool-use loop, strict tool schemas, server-side refusal fallbacks).
    Every tool call goes through execute_gated.
  • The page the agents read is poisoned. Researcher and mailer share one taint session, so the
    injection and the copied customer data follow the hand-off between agents.
  • The reviewer agent screens each held call with structured output and can only reject or
    escalate; a human decides escalations (TTY prompt; non-interactive denies unless
    --approve-escalations). It reads call arguments as fenced untrusted data.
  • Offline mode replays scripted model turns through the same loop, playing a researcher that
    falls for the injection, so the defences are exercised deterministically without a key.
  • Tests: offline scenario (exfiltration rejected, legitimate email sent only after a human OK,
    nothing sent when the human denies); real SDK request/response round-trip against a local fake
    Messages API; a live test that runs when ANTHROPIC_API_KEY is set.
  • [demo] extra (anthropic).

…s behind one gateway

- sentinel demo [--offline] [--approve-escalations]: three agents on the Anthropic SDK
  (claude-opus-5, manual tool-use loop, strict tool schemas, server-side refusal fallbacks).
  Every tool call goes through execute_gated.
- The page the agents read is poisoned. Researcher and mailer share one taint session, so the
  injection and the copied customer data follow the hand-off between agents.
- The reviewer agent screens each held call with structured output and can only reject or
  escalate; a human decides escalations (TTY prompt; non-interactive denies unless
  --approve-escalations). It reads call arguments as fenced untrusted data.
- Offline mode replays scripted model turns through the same loop, playing a researcher that
  falls for the injection, so the defences are exercised deterministically without a key.
- Tests: offline scenario (exfiltration rejected, legitimate email sent only after a human OK,
  nothing sent when the human denies); real SDK request/response round-trip against a local fake
  Messages API; a live test that runs when ANTHROPIC_API_KEY is set.
- [demo] extra (anthropic).
@Chirudeva-Reddy
Chirudeva-Reddy merged commit 90a5142 into main Sep 24, 2026
7 checks passed
@Chirudeva-Reddy
Chirudeva-Reddy deleted the feat/multi-agent-demo branch September 24, 2026 17:44
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant