feat(research): scan the files a skill ships, not just the file it documents - #58
Merged
Merged
Conversation
…cuments Step one measured the surface: 47.2% of a 290-skill sample name files no scan has ever read, 182 of them executable code. This reads them. The output that matters is the divergence set -- skills whose SKILL.md is clean but whose referenced code is not. That is hidden behaviour by construction: the documentation a reviewer or user sees says one thing, the code that runs says another. It is the only shape of finding worth calling malicious, and nothing in the pipeline could surface it before now. Two constraints, both learned the expensive way this fortnight: It does not convict. The rule engine is calibrated for SKILL.md prose; its false-positive profile on JavaScript and Python is unmeasured, and this project has already published one inflated number by assuming a rule meant what it appeared to mean. The report labels its own output as leads to read by hand. A 404 is not absence. A referenced path that fails to fetch is counted separately, because "names a script we could not retrieve" and "ships nothing" are different facts, and collapsing them is how a blind spot gets reported as larger than it is. Extraction moves into malwar.research.references so both scripts share one implementation with real test coverage. The tests are weighted toward negative cases -- every rule defect shipped here was a false positive found by hand afterwards, so the must-not-match set is deliberately larger than the must-match set and covers the shapes that fooled the detection rules: remote URLs, the user's own dotfiles, paths escaping the package, and prose that merely resembles a path. 1752 passing, 34 new. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DNoTXU8k3pfSBzR7aJubqL
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What this does
Step one measured the surface: 47.2% of a 290-skill sample name files no scan has ever read, 182 of them executable. This fetches and scans them.
The output that matters is the divergence set — skills whose
SKILL.mdis clean but whose referenced code is not. That is hidden behaviour by construction: the documentation a reviewer or user reads says one thing, the code that runs says another. It is the only shape of finding worth calling malicious, and nothing in the pipeline could surface it before now.Two constraints, both learned the expensive way
It does not convict. The rule engine is calibrated for SKILL.md prose. Its false-positive profile on JavaScript and Python is unmeasured, and this project has already published one inflated number by assuming a rule meant what it appeared to mean. The report labels its own output as leads to read by hand, in the output itself, not just in a docstring.
A 404 is not absence. A referenced path that fails to fetch is counted separately. "Names a script we could not retrieve" and "ships nothing" are different facts, and collapsing them is how a blind spot gets reported as larger than it is.
Structure
Extraction moves to
malwar.research.referencesso both scripts share one implementation with real coverage, rather than the copy-paste that existed after step one.On the tests
Deliberately weighted toward negative cases. Every rule defect shipped here was a false positive found by hand afterwards — MULTI-001 reading consent as evasion, PERSIST-001 convicting a crontab, PERM-001 firing on
--yoloinside a sentence warning against--yolo. So the must-not-match set is larger than the must-match set and covers the shapes that fooled those rules:curl … | sh,irm … | iex) — not a file the registry serves~/.bashrc,/etc/init.d/…)../../../etc/passwd.js)SKILL.md,README.md,package.json,node_modules/…)e.g. tax and billing,version 2.0.1)Plus a whole-file case from
buddy-card(both scripts captured) and the common case that must yield nothing: a skill that is purely instructions.What it will not tell us
Whether anything is malicious. It narrows 74,000 skills to a set worth reading. If the divergence set comes back empty, that is a real result — the unread surface is large and, in that sample, not hiding anything — and the report says so rather than implying a threat.
🤖 Generated with Claude Code
https://claude.ai/code/session_01DNoTXU8k3pfSBzR7aJubqL
Generated by Claude Code