feat: phase 6 generated evaluation, metrics, honest docs (findings 12, 14) - #7
Merged
Merged
Conversation
…, 14) - sentinel eval [--json|--markdown]: detection and false-positive rates from the corpora, now shipped in the package (sentinel/corpus/). It reports three separate attack rates (flagged, stopped on detector evidence, stopped under default policy) so deny-by-default can't inflate the headline number. - docs/BENCHMARKS.md is generated; tests/eval fails if it drifts from the code, and gates flagged >= 95%, detector-stop >= 85%, FP <= 5%. - Attack corpus 15 -> 42 cases, covering every Phase 1 bypass row plus fullwidth, zero-width, hidden-markdown, IFS, gopher/IPv6/v4-mapped SSRF, revshell and DB privilege cases. Current: 100% flagged, 88% stopped by detectors, 0/50 false positives. Misses are listed. - Reading credential material (id_rsa, .env, .aws/credentials, /etc/shadow, .pem, kubeconfig) is now CRITICAL on its own (attacks/ssh_key_read: WARN -> REQUIRE_APPROVAL). - GET /metrics (Prometheus text, no new deps): decisions, output-guard hits, detector errors, inspect latency, pending approvals. - README and ARCHITECTURE rewritten to match the code: removes the LangChain/CrewAI, "AST", 0.06 ms and "100%" claims; documents the threat model, keys and known limits. Not done: AgentDojo run (needs paid LLM calls), ML detector plugin, OpenTelemetry spans.
…m to SENTINEL_HOME AuditLedger and ApprovalCoordinator accept an explicit key.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Phase 6 of the architecture improvement plan. Stacked on
phase-5-integrations: review only this PR's commits.feat: phase 6 generated evaluation, metrics, honest docs (findings 12, 14)
shipped in the package (sentinel/corpus/). It reports three separate attack rates (flagged,
stopped on detector evidence, stopped under default policy) so deny-by-default can't inflate
the headline number.
flagged >= 95%, detector-stop >= 85%, FP <= 5%.
hidden-markdown, IFS, gopher/IPv6/v4-mapped SSRF, revshell and DB privilege cases.
Current: 100% flagged, 88% stopped by detectors, 0/50 false positives. Misses are listed.
is now CRITICAL on its own (attacks/ssh_key_read: WARN -> REQUIRE_APPROVAL).
inspect latency, pending approvals.
0.06 ms and "100%" claims; documents the threat model, keys and known limits.
Not done: AgentDojo run (needs paid LLM calls), ML detector plugin, OpenTelemetry spans.
fix: sentinel eval uses throwaway HMAC keys instead of persisting them to SENTINEL_HOME
AuditLedger and ApprovalCoordinator accept an explicit key.
Checks
ruff check,ruff format --check,mypy --strict,pytest --covpass locally on Python 3.10–3.13.