A signed audit trail
+Every decision is appended to an HMAC-chained ledger with sequence numbers and a signed head, so edits, deletions and truncation are all detected. Secrets are redacted before anything is written.
+diff --git a/README.md b/README.md index 52e08a2..c63ca88 100644 --- a/README.md +++ b/README.md @@ -2,16 +2,20 @@ # 🛡️ SentinelAgent -**A deny-by-default security gateway, human-approval sandbox and MCP proxy for AI agents** +**Guard your agents' tool calls.** A deny-by-default security gateway, human-approval sandbox and MCP proxy for AI agents. [](https://pypi.org/project/sentinel-agent-gateway/) [](https://github.com/Chirudeva-Reddy/sentinel-agent/actions) [](https://www.python.org/) -[](LICENSE) +[](https://github.com/Chirudeva-Reddy/sentinel-agent/blob/main/LICENSE) -**[Live demo: runs in your browser](https://chirudeva-reddy.github.io/sentinel-agent/)** · `pip install sentinel-agent-gateway` +**[Try it live](https://chirudeva-reddy.github.io/sentinel-agent/)**: the real gateway and the multi-agent demo, running in your browser. -[Why](#why) • [Quickstart](#quickstart) • [How it works](#how-it-works) • [Evaluation](#evaluation) • [Architecture](docs/ARCHITECTURE.md) +```bash +pip install sentinel-agent-gateway +``` + +[Architecture](https://github.com/Chirudeva-Reddy/sentinel-agent/blob/main/docs/ARCHITECTURE.md) @@ -135,11 +139,11 @@ flowchart LR O --> L ``` -Details, threat model and limits: [docs/ARCHITECTURE.md](docs/ARCHITECTURE.md). +Details, threat model and limits: [docs/ARCHITECTURE.md](https://github.com/Chirudeva-Reddy/sentinel-agent/blob/main/docs/ARCHITECTURE.md). ## Evaluation -All numbers come from [`docs/BENCHMARKS.md`](docs/BENCHMARKS.md), which `sentinel eval --markdown` generates from +All numbers come from [`docs/BENCHMARKS.md`](https://github.com/Chirudeva-Reddy/sentinel-agent/blob/main/docs/BENCHMARKS.md), which `sentinel eval --markdown` generates from the corpora in `sentinel/corpus/`. A test fails if the file and the code disagree. On the current 42 attack / 50 benign cases: every attack is flagged, 88% are stopped by detector evidence alone, and there are no false positives. Known misses are listed there. @@ -148,7 +152,7 @@ the CI `benchmark` job (`pytest -m benchmark`). ## Development -See [AGENTS.md](AGENTS.md) for the workflow (tests first, one regression test per bug, generated docs). +See [AGENTS.md](https://github.com/Chirudeva-Reddy/sentinel-agent/blob/main/AGENTS.md) for the workflow (tests first, one regression test per bug, generated docs). ```bash uv run ruff check . && uv run ruff format --check . && uv run mypy && uv run pytest --cov @@ -156,4 +160,4 @@ uv run ruff check . && uv run ruff format --check . && uv run mypy && uv run pyt ## License -[Apache 2.0](LICENSE) +[Apache 2.0](https://github.com/Chirudeva-Reddy/sentinel-agent/blob/main/LICENSE) diff --git a/site/index.html b/site/index.html index 3dd7931..0bd4452 100644 --- a/site/index.html +++ b/site/index.html @@ -3,121 +3,343 @@
-A deny-by-default security gateway for AI agents. It checks tool calls going out, guards tool
- output coming back, and holds risky calls for a human. This page runs the real sentinel package
- from main in your browser with Pyodide. No server, nothing leaves the page.
- GitHub - PyPI - Architecture - Benchmarks -
-Pick an attack or benign call from the evaluation corpus, or write your own.
-A deny-by-default gateway for AI agents. Risky calls wait for a human, and untrusted tool output is tracked.
+pip install sentinel-agent-gateway
+ Python 3.10+. Import name sentinel.
+
+ Attacks on agents arrive in what tools return, not just in what the agent asks to do. Sentinel covers both, then asks a person.
+Arguments are normalised, then scored by pluggable detectors. Shell is parsed into argv, so rm -r -f / is caught under any tool name. URL hosts are parsed as IP addresses, so decimal, hex and IPv6 forms can't slip past the SSRF check.
The policy can only make a decision stricter. Unknown tools wait for approval by default.
+Web pages, emails and files are fenced as untrusted data and fingerprinted. If that text reaches a high-risk tool later in the session, a human decides.
+An approval is an HMAC token bound to a digest of the exact arguments. It expires, works once, and fails if anything changed while you were deciding.
+Every decision is appended to an HMAC-chained ledger with sequence numbers and a signed head, so edits, deletions and truncation are all detected. Secrets are redacted before anything is written.
+A researcher agent and a mailer agent summarise a release-notes page that contains a hidden - instruction to email the customer list to an outside address. A reviewer agent screens every held call and - can only reject or escalate. The model turns are scripted here, and the researcher is written to fall for the - injection. Everything else (gateway, taint tracking, approvals, audit ledger) is the real code. Escalations are - auto-approved as the human.
- - +They summarise a release-notes page with a hidden instruction to email the customer list out. Model turns are scripted here; the gateway is the real code.
+Runs sentinel eval on the bundled corpora: attacks and benign calls under the default policy.
Computed just now in this tab from the bundled corpus of attack and benign tool calls, under the default policy.
+The corpus is small and hand-written, so treat these as regression numbers. Known misses and method.