Skip to content

feat: phase 6 generated evaluation, metrics, honest docs (findings 12, 14) - #7

Merged
Chirudeva-Reddy merged 2 commits into
mainfrom
phase-6-eval-observability
Sep 24, 2026
Merged

Chirudeva-Reddy merged 2 commits into
mainfrom
phase-6-eval-observability

Conversation

@Chirudeva-Reddy

Copy link
Copy Markdown
Owner

Phase 6 of the architecture improvement plan. Stacked on phase-5-integrations: review only this PR's commits.

feat: phase 6 generated evaluation, metrics, honest docs (findings 12, 14)

  • sentinel eval [--json|--markdown]: detection and false-positive rates from the corpora, now
    shipped in the package (sentinel/corpus/). It reports three separate attack rates (flagged,
    stopped on detector evidence, stopped under default policy) so deny-by-default can't inflate
    the headline number.
  • docs/BENCHMARKS.md is generated; tests/eval fails if it drifts from the code, and gates
    flagged >= 95%, detector-stop >= 85%, FP <= 5%.
  • Attack corpus 15 -> 42 cases, covering every Phase 1 bypass row plus fullwidth, zero-width,
    hidden-markdown, IFS, gopher/IPv6/v4-mapped SSRF, revshell and DB privilege cases.
    Current: 100% flagged, 88% stopped by detectors, 0/50 false positives. Misses are listed.
  • Reading credential material (id_rsa, .env, .aws/credentials, /etc/shadow, .pem, kubeconfig)
    is now CRITICAL on its own (attacks/ssh_key_read: WARN -> REQUIRE_APPROVAL).
  • GET /metrics (Prometheus text, no new deps): decisions, output-guard hits, detector errors,
    inspect latency, pending approvals.
  • README and ARCHITECTURE rewritten to match the code: removes the LangChain/CrewAI, "AST",
    0.06 ms and "100%" claims; documents the threat model, keys and known limits.

Not done: AgentDojo run (needs paid LLM calls), ML detector plugin, OpenTelemetry spans.

fix: sentinel eval uses throwaway HMAC keys instead of persisting them to SENTINEL_HOME

AuditLedger and ApprovalCoordinator accept an explicit key.

Checks

  • ruff check, ruff format --check, mypy --strict, pytest --cov pass locally on Python 3.10–3.13.

…, 14)

- sentinel eval [--json|--markdown]: detection and false-positive rates from the corpora, now
  shipped in the package (sentinel/corpus/). It reports three separate attack rates (flagged,
  stopped on detector evidence, stopped under default policy) so deny-by-default can't inflate
  the headline number.
- docs/BENCHMARKS.md is generated; tests/eval fails if it drifts from the code, and gates
  flagged >= 95%, detector-stop >= 85%, FP <= 5%.
- Attack corpus 15 -> 42 cases, covering every Phase 1 bypass row plus fullwidth, zero-width,
  hidden-markdown, IFS, gopher/IPv6/v4-mapped SSRF, revshell and DB privilege cases.
  Current: 100% flagged, 88% stopped by detectors, 0/50 false positives. Misses are listed.
- Reading credential material (id_rsa, .env, .aws/credentials, /etc/shadow, .pem, kubeconfig)
  is now CRITICAL on its own (attacks/ssh_key_read: WARN -> REQUIRE_APPROVAL).
- GET /metrics (Prometheus text, no new deps): decisions, output-guard hits, detector errors,
  inspect latency, pending approvals.
- README and ARCHITECTURE rewritten to match the code: removes the LangChain/CrewAI, "AST",
  0.06 ms and "100%" claims; documents the threat model, keys and known limits.

Not done: AgentDojo run (needs paid LLM calls), ML detector plugin, OpenTelemetry spans.
…m to SENTINEL_HOME

AuditLedger and ApprovalCoordinator accept an explicit key.
@Chirudeva-Reddy
Chirudeva-Reddy deleted the branch main September 24, 2026 17:22
@Chirudeva-Reddy
Chirudeva-Reddy changed the base branch from phase-5-integrations to main September 24, 2026 17:24
@Chirudeva-Reddy
Chirudeva-Reddy merged commit 1b373aa into main Sep 24, 2026
14 checks passed
@Chirudeva-Reddy
Chirudeva-Reddy deleted the phase-6-eval-observability branch September 24, 2026 17:25
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant