Repository navigation
feat(judge): sampled after-the-fact judge audit of allowed tool calls - #562
Merged
Merged
Conversation
prismor audit judge picks a deterministic sample of tool calls the rules allowed (no finding) from the local store, scrubs them, and has the configured judge review them offline. Verdicts land in a new judge_audit table; each flagged call is printed and sent as one content-free, chained judge_audit record through the normal sinks. Never blocks, never touches the hook path. Exits 2 when no judge is configured. Co-Authored-By: Claude <noreply@anthropic.com>
Contributor
Author
|
Added f8c5fb0. The judge_audit record now carries what a reviewer needs:
All of these are ids, enums or numbers, so they survive redaction. The reason stays evidence: it is hashed when redacted, and under full capture it goes in |
Contributor
Author
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Why
Regex rules miss attacks. On our benchmark the regex-gated layer caught about 34% of injections, while judging every text caught 96-100%. The judge can't go on every hook call, though. It takes seconds per call, and a hook already costs about 165 ms. So whatever the rules allow is never seen by a model.
This PR finds what the rules let through without touching hook latency. It samples calls that were already allowed and judges them in a batch afterwards.
It has to run on the device. In the default redacted mode the console never receives the content of an allowed call, only counts. The local store (
prismor.db) keeps the full events.Design
prismor audit judgeis a new target on the existingauditcommand, becauseprismor auditis already the posture audit. Plainprismor auditandaudit --fixbehave as before.eventspicks pre-call tool events in the window that have no finding at the same(session_id, event_index). Prompts, tool output, memory events and post-call repeats are left out. Self-test sessions are excluded with the existinglearning._FIXTURE_SESSIONS_SQL, and events already injudge_auditare skipped.<session_id>:<event_index>, the same key findings use.events.idcan't be used because it changes every time a session snapshot is rewritten.sha256(event_id)[:8] / 2^32 < rate, so every run picks the same events. The sample is capped at--max.semantic_guard.provider/model/cli_path, andmode: hybridfor CLI opt-in. It is built throughSemanticGuardV2withlow_threshold=0, so every sampled call is judged rather than only the heuristic band. It reuses the existing verdict cache, hosted-judge fallback, CLI timeouts andbudget_ms(viapolicy_engine._analyze_within).tool: <name>plusinput: <tool_input JSON>, capped at one 3,000-character judge window. It first goes throughredaction.redact_text(cloak values and data-boundary values), thensinks._scrub_for_sink(secret patterns, URL passwords,key=valuecredentials; fails closed). The hosted judge gets the same scrubbed text.ls -la0.35.judge_audittable:event_idPK,session_id,event_index,ts,audited_at,verdict(flagged/clean againstwarn_threshold),risk_score,category,reason,model. Writes areINSERT OR REPLACE, so reruns are idempotent. The table is created ininitialize_database, soprismor query --schemalists it.sinks.dispatch:ruleId: judge-audit,mode: observe(verdictobserved), the judge's category, a static title, and the reason as evidence (hashed in redacted mode).type: judge_audit, tool name,tool_use_id. So the prismor sink'sbuild_record→assert_redacted→_seal(chain + signature) applies unchanged, and even full capture has no call content to ship.semantic_guard.audit: {sample_rate: 0.05, max_per_run: 50, window: 24h}. It is read only when the command runs.Privacy
Tests
tests/test_judge_audit.pyhas 12 tests. There is no network: the judge is aregister_llmfake and telemetry is captured at the upload or the spool.--maxcaps the judge calls, and a rerun skips audited events.detail, passesassert_redacted, hashash/chain_seq,prev_hashlinks to the previous record, and host and path are absent.JudgeNotConfiguredis raised and the CLI exits 2 with the message.--dry-runmakes no judge calls and writes no rows.query --schema.scripts/check_oss_safe.pypasses.End-to-end run (isolated home)
Setup:
PRISMOR_HOME=$(mktemp -d)with the otherPRISMOR_*variables unset, and an empty temp workspace.prismor hook-dispatch --agent claude --mode enforce: a prompt,ls -la(pre and post), acurl --data-binary @filePOST,rm -rf /, a Node one-liner that readsnotes/config.txtand PUTs it to an unknown host, andgit status.network-exfil-tool,file-upload-to-external) and blockedrm -rf /. The Node exfil got no finding at all, so it is exactly the kind of call this audit is for.Without a judge configured:
Dry run:
With the judge that was already installed on the machine (
provider: claude, the local Claude Code login, no new keys):The telemetry record a flagged call produces, from the fake-judge run (same shape). The device had a throwaway identity pointing at
127.0.0.1:1and a prismor sink, so the record was spooled:{ "category": "data_exfiltration", "chain_seq": 0, "evidence_hash": "82e1e5d54a271449", "hash": "4ac1e2e46944e777dc2f9ebafb12ef3ec55a802916fc1756c879787f1bfb2a0f", "matched_pattern": null, "mode": "observe", "prev_hash": "0000000000000000000000000000000000000000000000000000000000000000", "redacted": true, "rule_id": "judge-audit", "session_id": "s1", "session_seq": 5, "severity": "HIGH", "signature": "4lgpylHWy6C9ubOiykUrtxRUHee4APnr25jVJRiNld82SIBaJpBLysqbVOBPjwMcADLLeT6ht4YPByTrlvopBw==", "signing_alg": "ed25519", "title": "Judge audit flagged an allowed tool call", "tool_name": "Bash", "tool_use_id": "t4", "type": "judge_audit", "verdict": "observed" }A rerun judged nothing (
0 allowed call(s)), because all three events were already audited.Follow-ups
judge_auditrecords.