Skip to content

feat(judge): sampled after-the-fact judge audit of allowed tool calls - #562

Merged
Ar9av merged 1 commit into
mainfrom
feat/judge-audit
Oct 1, 2026
Merged

Ar9av merged 1 commit into
mainfrom
feat/judge-audit

Conversation

@Ar9av

@Ar9av Ar9av commented Oct 1, 2026

Copy link
Copy Markdown
Contributor

Why

Regex rules miss attacks. On our benchmark the regex-gated layer caught about 34% of injections, while judging every text caught 96-100%. The judge can't go on every hook call, though. It takes seconds per call, and a hook already costs about 165 ms. So whatever the rules allow is never seen by a model.

This PR finds what the rules let through without touching hook latency. It samples calls that were already allowed and judges them in a batch afterwards.

It has to run on the device. In the default redacted mode the console never receives the content of an allowed call, only counts. The local store (prismor.db) keeps the full events.

Design

prismor audit judge is a new target on the existing audit command, because prismor audit is already the posture audit. Plain prismor audit and audit --fix behave as before.

prismor audit judge [--since 24h] [--sample 0.05] [--max 50] [--dry-run] [--json] [--workspace W]
  • Selection. One SQL query over events picks pre-call tool events in the window that have no finding at the same (session_id, event_index). Prompts, tool output, memory events and post-call repeats are left out. Self-test sessions are excluded with the existing learning._FIXTURE_SESSIONS_SQL, and events already in judge_audit are skipped.
  • Event id. The id is <session_id>:<event_index>, the same key findings use. events.id can't be used because it changes every time a session snapshot is rewritten.
  • Sampling. sha256(event_id)[:8] / 2^32 < rate, so every run picks the same events. The sample is capped at --max.
  • Judge. The same judge the hooks use: semantic_guard.provider / model / cli_path, and mode: hybrid for CLI opt-in. It is built through SemanticGuardV2 with low_threshold=0, so every sampled call is judged rather than only the heuristic band. It reuses the existing verdict cache, hosted-judge fallback, CLI timeouts and budget_ms (via policy_engine._analyze_within).
    • A judge failure or a budget overrun is not recorded, so the next run retries it.
    • With no judge configured the command exits 2 with a clear message.
  • What the judge sees. tool: <name> plus input: <tool_input JSON>, capped at one 3,000-character judge window. It first goes through redaction.redact_text (cloak values and data-boundary values), then sinks._scrub_for_sink (secret patterns, URL passwords, key=value credentials; fails closed). The hosted judge gets the same scrubbed text.
    • The text has no "this was allowed" preamble. On a live run the injection-tuned judge read that phrase as a claim of prior approval and scored a bare ls -la 0.35.
    • The judge has no separate field for session intent, so the user's recent prompt is not sent.
  • Results. Stored in a new judge_audit table: event_id PK, session_id, event_index, ts, audited_at, verdict (flagged/clean against warn_threshold), risk_score, category, reason, model. Writes are INSERT OR REPLACE, so reruns are idempotent. The table is created in initialize_database, so prismor query --schema lists it.
  • Flagged calls.
    • Printed with the session/event id and the judge's reason.
    • Each one produces one finding through sinks.dispatch: ruleId: judge-audit, mode: observe (verdict observed), the judge's category, a static title, and the reason as evidence (hashed in redacted mode).
    • The event handed to the sinks is content-free: type: judge_audit, tool name, tool_use_id. So the prismor sink's build_record → assert_redacted → _seal (chain + signature) applies unchanged, and even full capture has no call content to ship.
    • Nothing is blocked retroactively.
  • Config. semantic_guard.audit: {sample_rate: 0.05, max_per_run: 50, window: 24h}. It is read only when the command runs.
  • Background trigger: not added. A detached spawn at SessionStart can't be guaranteed to add zero hook latency, so the docs point to cron for now (see follow-ups).

Privacy

  • Content is scrubbed before it reaches any judge, hosted included.
  • In redacted mode, the only new thing that leaves the device is one enum-style record per flagged call: type, rule id, category, severity, static title, tool name, ids and hashes. It has no command, path, host or reason text, and it is hash-chained and signed like every other record.
  • Clean verdicts never leave the device.

Tests

tests/test_judge_audit.py has 12 tests. There is no network: the judge is a register_llm fake and telemetry is captured at the upload or the spool.

  • Sampling is deterministic and close to the rate, --max caps the judge calls, and a rerun skips audited events.
  • Only allowed pre-call events are selected: no finding, no post-call repeat, no prompt, and fixture sessions are excluded.
  • A secret-shaped key in the command never reaches the fake judge (the test asserts on what the fake received).
  • Flagged results are stored, and exactly one record per flagged call goes through the real prismor sink. It is redacted, has no detail, passes assert_redacted, has hash/chain_seq, prev_hash links to the previous record, and host and path are absent.
  • Clean results produce no dispatch.
  • With no judge, JudgeNotConfigured is raised and the CLI exits 2 with the message.
  • --dry-run makes no judge calls and writes no rows.
  • A failed judge reply is not recorded.
  • Config defaults and override work, and the table shows up in query --schema.
$ python3 -m pytest -p no:randomly tests/test_judge_audit.py tests/test_audit_feed_signature.py tests/test_cli_help.py \
    tests/test_cli_main_guard.py tests/test_enterprise_telemetry.py tests/test_learning.py tests/test_prismor_sink_batching.py \
    tests/test_query.py tests/test_sinks_redaction.py tests/test_store_dashboard_queries.py tests/test_store_finding_id_collision.py \
    tests/test_store_symlink_workspace.py tests/test_telemetry_chain.py tests/test_semantic_guard_provider.py tests/test_semantic_guard_llm.py -q
113 passed            # origin/main, same files minus the new one: 101 passed

scripts/check_oss_safe.py passes.

End-to-end run (isolated home)

Setup:

  • A fresh PRISMOR_HOME=$(mktemp -d) with the other PRISMOR_* variables unset, and an empty temp workspace.
  • Synthetic Claude Code payloads piped to prismor hook-dispatch --agent claude --mode enforce: a prompt, ls -la (pre and post), a curl --data-binary @file POST, rm -rf /, a Node one-liner that reads notes/config.txt and PUTs it to an unknown host, and git status.
  • The regex rules warned on the curl (network-exfil-tool, file-upload-to-external) and blocked rm -rf /. The Node exfil got no finding at all, so it is exactly the kind of call this audit is for.

Without a judge configured:

$ prismor audit judge --sample 1.0
prismor audit judge: no LLM judge configured: set settings.semantic_guard.provider (prismor after `prismor login`, api + model, claude or codex) — see `prismor docs semantic-guard`
exit=2

Dry run:

$ prismor audit judge --sample 1.0 --dry-run
Judge audit: 3 allowed call(s) in the last 24h, 3 sampled at 1 (max 50)
  would judge  s1:6  2026-10-01T06:37:17.787548+00:00  shell
  would judge  s1:5  2026-10-01T06:37:17.594520+00:00  shell
  would judge  s1:1  2026-10-01T06:37:16.817977+00:00  shell
Dry run: no judge calls made.

With the judge that was already installed on the machine (provider: claude, the local Claude Code login, no new keys):

$ prismor audit judge --sample 1.0
Judge audit: 3 allowed call(s) in the last 24h, 3 sampled at 1 (max 50)
Judged 3, flagged 1
  FLAGGED  s1:5  Bash  credential_exfiltration 0.98
           Malicious command attempts to read a local config file and exfiltrate it via HTTP PUT to a suspicious third-party domain.
Allowed calls are never blocked retroactively. Inspect one with: prismor query "SELECT * FROM judge_audit WHERE verdict='flagged'"

$ prismor query --format table "SELECT event_id, verdict, risk_score, category, model, reason FROM judge_audit"
event_id  verdict  risk_score  category                 model   reason
--------  -------  ----------  -----------------------  ------  ------------------------------------------------------------
s1:6      clean    0.0         clean                    claude  Standard read-only git command with no suspicious instructi…
s1:5      flagged  0.98        credential_exfiltration  claude  Malicious command attempts to read a local config file and …
s1:1      clean    0.0         clean                    claude  Standard benign file listing command with no injection, man…

The telemetry record a flagged call produces, from the fake-judge run (same shape). The device had a throwaway identity pointing at 127.0.0.1:1 and a prismor sink, so the record was spooled:

{
  "category": "data_exfiltration",
  "chain_seq": 0,
  "evidence_hash": "82e1e5d54a271449",
  "hash": "4ac1e2e46944e777dc2f9ebafb12ef3ec55a802916fc1756c879787f1bfb2a0f",
  "matched_pattern": null,
  "mode": "observe",
  "prev_hash": "0000000000000000000000000000000000000000000000000000000000000000",
  "redacted": true,
  "rule_id": "judge-audit",
  "session_id": "s1",
  "session_seq": 5,
  "severity": "HIGH",
  "signature": "4lgpylHWy6C9ubOiykUrtxRUHee4APnr25jVJRiNld82SIBaJpBLysqbVOBPjwMcADLLeT6ht4YPByTrlvopBw==",
  "signing_alg": "ed25519",
  "title": "Judge audit flagged an allowed tool call",
  "tool_name": "Bash",
  "tool_use_id": "t4",
  "type": "judge_audit",
  "verdict": "observed"
}

A rerun judged nothing (0 allowed call(s)), because all three events were already audited.

Follow-ups

  • A console view for judge_audit records.
  • Feed flagged results into the review queue (web ci: harden OSS CI with CodeQL, Scorecard, Dependabot, and SECURITY.md (closes #282, closes #277) #290).
  • A scheduled trigger: cron is documented for now. An opt-in once-a-day detached run is possible if it can be made provably free on the hook path.
  • A tool-call rubric for the judge. Today the audit reuses the injection-tuned rubric; it caught the exfil above, but a rubric that weighs the call against the session's stated intent would need a separate intent field on the judge, including the hosted endpoint.

prismor audit judge picks a deterministic sample of tool calls the rules
allowed (no finding) from the local store, scrubs them, and has the
configured judge review them offline. Verdicts land in a new judge_audit
table; each flagged call is printed and sent as one content-free,
chained judge_audit record through the normal sinks. Never blocks,
never touches the hook path. Exits 2 when no judge is configured.

Co-Authored-By: Claude <noreply@anthropic.com>
@Ar9av
Ar9av merged commit 5c4bee2 into main Oct 1, 2026
8 checks passed
@Ar9av

Ar9av commented Oct 1, 2026

Copy link
Copy Markdown
Contributor Author

Added f8c5fb0. The judge_audit record now carries what a reviewer needs:

  • audited_event (<session_id>:<event_index>)
  • risk_score
  • model
  • a severity from the score (MEDIUM, or HIGH at or over block_threshold)

All of these are ids, enums or numbers, so they survive redaction. The reason stays evidence: it is hashed when redacted, and under full capture it goes in detail, now also run through the sink scrubber. The console side is Ar9av/prismor-web-preview#293: flagged calls land in the review queue, and the labels give the judge's precision.

@Ar9av

Ar9av commented Oct 1, 2026

Copy link
Copy Markdown
Contributor Author

Correction: f8c5fb0 was pushed after this PR merged, so it isn't in main. It's cherry-picked into #566.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant