diff --git a/README.md b/README.md index 52e08a2..c63ca88 100644 --- a/README.md +++ b/README.md @@ -2,16 +2,20 @@ # 🛡️ SentinelAgent -**A deny-by-default security gateway, human-approval sandbox and MCP proxy for AI agents** +**Guard your agents' tool calls.** A deny-by-default security gateway, human-approval sandbox and MCP proxy for AI agents. [![PyPI](https://img.shields.io/pypi/v/sentinel-agent-gateway.svg)](https://pypi.org/project/sentinel-agent-gateway/) [![CI](https://github.com/Chirudeva-Reddy/sentinel-agent/actions/workflows/ci.yml/badge.svg)](https://github.com/Chirudeva-Reddy/sentinel-agent/actions) [![Python 3.10+](https://img.shields.io/badge/python-3.10%20%E2%80%93%203.14-blue.svg)](https://www.python.org/) -[![License](https://img.shields.io/badge/license-Apache%202.0-blue.svg)](LICENSE) +[![License](https://img.shields.io/badge/license-Apache%202.0-blue.svg)](https://github.com/Chirudeva-Reddy/sentinel-agent/blob/main/LICENSE) -**[Live demo: runs in your browser](https://chirudeva-reddy.github.io/sentinel-agent/)** · `pip install sentinel-agent-gateway` +**[Try it live](https://chirudeva-reddy.github.io/sentinel-agent/)**: the real gateway and the multi-agent demo, running in your browser. -[Why](#why) • [Quickstart](#quickstart) • [How it works](#how-it-works) • [Evaluation](#evaluation) • [Architecture](docs/ARCHITECTURE.md) +```bash +pip install sentinel-agent-gateway +``` + +[Architecture](https://github.com/Chirudeva-Reddy/sentinel-agent/blob/main/docs/ARCHITECTURE.md) @@ -135,11 +139,11 @@ flowchart LR O --> L ``` -Details, threat model and limits: [docs/ARCHITECTURE.md](docs/ARCHITECTURE.md). +Details, threat model and limits: [docs/ARCHITECTURE.md](https://github.com/Chirudeva-Reddy/sentinel-agent/blob/main/docs/ARCHITECTURE.md). ## Evaluation -All numbers come from [`docs/BENCHMARKS.md`](docs/BENCHMARKS.md), which `sentinel eval --markdown` generates from +All numbers come from [`docs/BENCHMARKS.md`](https://github.com/Chirudeva-Reddy/sentinel-agent/blob/main/docs/BENCHMARKS.md), which `sentinel eval --markdown` generates from the corpora in `sentinel/corpus/`. A test fails if the file and the code disagree. On the current 42 attack / 50 benign cases: every attack is flagged, 88% are stopped by detector evidence alone, and there are no false positives. Known misses are listed there. @@ -148,7 +152,7 @@ the CI `benchmark` job (`pytest -m benchmark`). ## Development -See [AGENTS.md](AGENTS.md) for the workflow (tests first, one regression test per bug, generated docs). +See [AGENTS.md](https://github.com/Chirudeva-Reddy/sentinel-agent/blob/main/AGENTS.md) for the workflow (tests first, one regression test per bug, generated docs). ```bash uv run ruff check . && uv run ruff format --check . && uv run mypy && uv run pytest --cov @@ -156,4 +160,4 @@ uv run ruff check . && uv run ruff format --check . && uv run mypy && uv run pyt ## License -[Apache 2.0](LICENSE) +[Apache 2.0](https://github.com/Chirudeva-Reddy/sentinel-agent/blob/main/LICENSE) diff --git a/site/index.html b/site/index.html index 3dd7931..0bd4452 100644 --- a/site/index.html +++ b/site/index.html @@ -3,121 +3,343 @@ -SentinelAgent - +SentinelAgent: a security gateway for AI agents + + + + + + + + -
-

🛡️ SentinelAgent

-

A deny-by-default security gateway for AI agents. It checks tool calls going out, guards tool - output coming back, and holds risky calls for a human. This page runs the real sentinel package - from main in your browser with Pyodide. No server, nothing leaves the page.

- -
Loading the Python runtime…
- -
-

Try the gateway

-

Pick an attack or benign call from the evaluation corpus, or write your own.

-
-
- - - - - - - - - + + +
+
+
+ Open-source agent security +

Guard your agents’ tool calls.

+

A deny-by-default gateway for AI agents. Risky calls wait for a human, and untrusted tool output is tracked.

+ -
The decision will appear here.
+
+ +
+
+
+ Inspect a tool call + Starting Python in your browser +
+ + +
+
+
+
+ + +
+
+
+
+
+
+
+ +
+
+ pip install sentinel-agent-gateway + Python 3.10+. Import name sentinel. + +
+
+ +
+

Checks both directions of every call.

+

Attacks on agents arrive in what tools return, not just in what the agent asks to do. Sentinel covers both, then asks a person.

+
+
+

Calls going out

+

Arguments are normalised, then scored by pluggable detectors. Shell is parsed into argv, so rm -r -f / is caught under any tool name. URL hosts are parsed as IP addresses, so decimal, hex and IPv6 forms can't slip past the SSRF check.

+

The policy can only make a decision stricter. Unknown tools wait for approval by default.

+
unknown_tool_action: "REQUIRE_APPROVAL" +aggregation: "noisy_or" +taint_sinks: + send_email: "high" + execute_bash: "high"
+
+
+

Output coming back

+

Web pages, emails and files are fenced as untrusted data and fingerprinted. If that text reaches a high-risk tool later in the session, a human decides.

+
+
+

One call, one approval

+

An approval is an HMAC token bound to a digest of the exact arguments. It expires, works once, and fails if anything changed while you were deciding.

+
+
+
+

A signed audit trail

+

Every decision is appended to an HMAC-chained ledger with sequence numbers and a signed head, so edits, deletions and truncation are all detected. Secrets are redacted before anything is written.

+
+
+
-
-

Multi-agent demo

-

A researcher agent and a mailer agent summarise a release-notes page that contains a hidden - instruction to email the customer list to an outside address. A reviewer agent screens every held call and - can only reject or escalate. The model turns are scripted here, and the researcher is written to fall for the - injection. Everything else (gateway, taint tracking, approvals, audit ledger) is the real code. Escalations are - auto-approved as the human.

- - +
+

Three agents, one gateway.

+

They summarise a release-notes page with a hidden instruction to email the customer list out. Model turns are scripted here; the gateway is the real code.

+
+
Researcherfetch_url, read_file
+
Mailersend_email
+
Reviewercan reject or escalate, never approve
+
+
+
+ Task: research the Q3 notes, email team@example.com a summary + +
+
Press run to watch the agents work. A person approves what the reviewer escalates; here that step is automatic.
+
+
-
-

Evaluation

-

Runs sentinel eval on the bundled corpora: attacks and benign calls under the default policy.

- -
+
+

Measured, not claimed.

+

Computed just now in this tab from the bundled corpus of attack and benign tool calls, under the default policy.

+
+
…
of attacks flagged
+
…
stopped on detector evidence alone
+
…
stopped under the default policy
+
…
false positives on benign calls
+
+

The corpus is small and hand-written, so treat these as regression numbers. Known misses and method.

- The approval API, the dashboard and the MCP proxy need a server. Run them with - docker compose -f docker/docker-compose.yml up or - pip install "sentinel-agent-gateway[server,dashboard,mcp]". + Everything above runs in your browser with Pyodide. The approval API, dashboard and MCP proxy need a server. + + PyPI + Architecture + GitHub +