fix: close the ten code-review findings (v0.2.1) - #11
Merged
Merged
Conversation
Every fix has a regression test in tests/regressions/test_issue_<n>.py that failed first (27 red + 1 collection error before this change). Detection - 17/18: shell evasion. `sh -c '...'`, `bash -lc`, `eval` scripts are re-parsed (3 levels), and wrapper programs are skipped together with their own flags (sudo -u root, env -i, nice -n 10, timeout 5, stdbuf, ionice, ...). Benign cases (git rm -r --cached, sudo apt list) stay clean. - 19: scheme-less SSRF. An argument that is only an address (169.254.169.254/latest, localhost:6379, //10.0.0.1/x, [::1]:8080) is treated as a network target. Prose mentioning an IP and small numbers (42/7) stay clean. - 20: tool output is scanned in overlapping 64 KB chunks up to 1 MB (1 MB takes about 127 ms); untrusted output beyond that fails closed; taint fingerprints cover the head and tail. - 21: base64 decoding has no candidate cap; the decoded text is joined and scanned once. Integrity - 22: verify_integrity catches up with other writers and reads under the writers' lock; the reader never consumes a partially written line. Fixes the false "truncated" health alarm. - 23: keys are written to a temp file and link()ed into place (atomic, never overwritten); an empty or short key file is refused instead of silently used. Coverage and structure - 24: POST /api/v1/results (agent key) runs tool output through the guard, so HTTP agents get spotlighting and taint tracking too. - 25: one fold_tool_name() used by normalize, the policy, approval digests, the gateway and the MCP proxy (the policy previously skipped NFKC). - 16: moved the duplicate-entry-point regression test to tests/regressions as AGENTS.md requires. Corpus: +6 attacks (nested shell, eval, sudo flags, timeout, 2 scheme-less SSRF), +5 benign near-misses. 48 attacks: 100% flagged, 90% stopped on detector evidence; 0/55 false positives. No existing snapshot row changed.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Every fix has a regression test in tests/regressions/test_issue_.py that failed first
(27 red + 1 collection error before this change).
Detection
sh -c '...',bash -lc,evalscripts are re-parsed (3 levels), andwrapper programs are skipped together with their own flags (sudo -u root, env -i, nice -n 10,
timeout 5, stdbuf, ionice, ...). Benign cases (git rm -r --cached, sudo apt list) stay clean.
localhost:6379, //10.0.0.1/x, [::1]:8080) is treated as a network target. Prose mentioning an IP
and small numbers (42/7) stay clean.
untrusted output beyond that fails closed; taint fingerprints cover the head and tail.
Integrity
reader never consumes a partially written line. Fixes the false "truncated" health alarm.
empty or short key file is refused instead of silently used.
Coverage and structure
spotlighting and taint tracking too.
the MCP proxy (the policy previously skipped NFKC).
Corpus: +6 attacks (nested shell, eval, sudo flags, timeout, 2 scheme-less SSRF), +5 benign
near-misses. 48 attacks: 100% flagged, 90% stopped on detector evidence; 0/55 false positives.
No existing snapshot row changed.