feat: phase 5 real MCP proxy and adapter contract (finding 13) - #6
Merged
Merged
Conversation
Indirect prompt injection arrives in tool output, which the gateway never saw before.
- gateway.inspect_result(call, result): runs output-capable detectors (SCANS_OUTPUT) on the
result, wraps output from policy taint_sources in <<untrusted-data>> fences with a
do-not-follow note (the fence can't be closed from inside), logs a TOOL_RESULT ledger event,
and never raises on bytes, dicts, None or multi-MB output.
- sentinel.taint.TaintTracker: per-session 20-char shingle fingerprints of untrusted output,
plus a "compromised" flag once injected text is seen. In inspect, a call to a taint_sinks
"high" tool whose arguments copy untrusted text, or that runs in a compromised session, gets
a taint_tracker CRITICAL finding and so needs approval.
- execute_gated takes session_id, guards every result, returns sanitized_result/result_guard.
- Policy: taint_sources, taint_sinks {tool: high|low}, taint_min_match.
- tests/fakes.ScriptedAgent replays tool calls + canned results without an LLM. Scenarios:
poisoned page -> send_email to attacker (gated, did not run); read file -> http_post of
copied text (gated); fetch docs -> summarize -> email team (all ALLOW); taint isolated per
session.
Known ceiling: paraphrased or re-encoded exfiltration is not caught (substring matching).
- sentinel mcp-proxy -- <upstream cmd...> | <url>: an MCP server (official mcp SDK, [mcp] extra) that forwards tools/list and routes every tools/call through execute_gated, so MCP gets the same inspect -> approval -> redeem -> execute -> output-guard flow as everything else. Untrusted-source tools drop structured content and their output_schema (it would bypass the guard). Upstream processes get the proxy's env minus every SENTINEL_* secret. - SentinelOpenAIWrapper now waits for approval instead of hard-blocking, guards results, and returns `content` ready for the tool message; async `call`, sync `intercept_and_call`. - execute_gated awaits async tools; an approval that fails verification (e.g. approver on a different SENTINEL_APPROVAL_KEY) becomes a blocked response instead of an exception. - Adapter contract suite, run for both adapters: allowed runs, blocked never reaches the tool, approval pauses then resumes, rejection doesn't run, output is fenced as untrusted. - Integration test over real processes: MCP client -> proxy subprocess -> toy server subprocess, approval resolved from another process through the shared SQLite store. - Removed the SDK-less SentinelMCPMiddleware; examples/mcp_proxy_example.py uses the proxy. - mypy --strict now covers all of sentinel/ (dashboard excluded); coverage gate 90%. Not shipped: OpenAI Agents SDK and LangGraph adapters (heavy optional deps; the contract suite is where they'd plug in).
Chirudeva-Reddy
changed the base branch from
phase-4-output-guard-taint
to
main
September 24, 2026 17:22
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Phase 5 of the architecture improvement plan. Stacked on
phase-4-output-guard-taint: review only this PR's commits.feat: phase 5 real MCP proxy and adapter contract (finding 13)
extra) that forwards tools/list and routes every tools/call through execute_gated, so MCP
gets the same inspect -> approval -> redeem -> execute -> output-guard flow as everything else.
Untrusted-source tools drop structured content and their output_schema (it would bypass the
guard). Upstream processes get the proxy's env minus every SENTINEL_* secret.
returns
contentready for the tool message; asynccall, syncintercept_and_call.different SENTINEL_APPROVAL_KEY) becomes a blocked response instead of an exception.
approval pauses then resumes, rejection doesn't run, output is fenced as untrusted.
subprocess, approval resolved from another process through the shared SQLite store.
Not shipped: OpenAI Agents SDK and LangGraph adapters (heavy optional deps; the contract suite
is where they'd plug in).
Checks
ruff check,ruff format --check,mypy --strict,pytest --covpass locally on Python 3.10–3.13.