Skip to content

feat(research): scan the files a skill ships, not just the file it documents - #58

Merged
Ap6pack merged 1 commit into
mainfrom
claude/scan-referenced-payloads
Sep 11, 2026
Merged

Ap6pack merged 1 commit into
mainfrom
claude/scan-referenced-payloads

Conversation

@Ap6pack

@Ap6pack Ap6pack commented Sep 10, 2026

Copy link
Copy Markdown
Owner

What this does

Step one measured the surface: 47.2% of a 290-skill sample name files no scan has ever read, 182 of them executable. This fetches and scans them.

The output that matters is the divergence set — skills whose SKILL.md is clean but whose referenced code is not. That is hidden behaviour by construction: the documentation a reviewer or user reads says one thing, the code that runs says another. It is the only shape of finding worth calling malicious, and nothing in the pipeline could surface it before now.

Two constraints, both learned the expensive way

It does not convict. The rule engine is calibrated for SKILL.md prose. Its false-positive profile on JavaScript and Python is unmeasured, and this project has already published one inflated number by assuming a rule meant what it appeared to mean. The report labels its own output as leads to read by hand, in the output itself, not just in a docstring.

A 404 is not absence. A referenced path that fails to fetch is counted separately. "Names a script we could not retrieve" and "ships nothing" are different facts, and collapsing them is how a blind spot gets reported as larger than it is.

Structure

Extraction moves to malwar.research.references so both scripts share one implementation with real coverage, rather than the copy-paste that existed after step one.

On the tests

Deliberately weighted toward negative cases. Every rule defect shipped here was a false positive found by hand afterwards — MULTI-001 reading consent as evasion, PERSIST-001 convicting a crontab, PERM-001 firing on --yolo inside a sentence warning against --yolo. So the must-not-match set is larger than the must-match set and covers the shapes that fooled those rules:

  • remote code (curl … | sh, irm … | iex) — not a file the registry serves
  • the user's own machine (~/.bashrc, /etc/init.d/…)
  • paths escaping the package (../../../etc/passwd.js)
  • already-scanned or non-package files (SKILL.md, README.md, package.json, node_modules/…)
  • prose that merely resembles a path (e.g. tax and billing, version 2.0.1)

Plus a whole-file case from buddy-card (both scripts captured) and the common case that must yield nothing: a skill that is purely instructions.

1752 passed, 16 skipped   (34 new)

What it will not tell us

Whether anything is malicious. It narrows 74,000 skills to a set worth reading. If the divergence set comes back empty, that is a real result — the unread surface is large and, in that sample, not hiding anything — and the report says so rather than implying a threat.

🤖 Generated with Claude Code

https://claude.ai/code/session_01DNoTXU8k3pfSBzR7aJubqL


Generated by Claude Code

…cuments

Step one measured the surface: 47.2% of a 290-skill sample name files no scan
has ever read, 182 of them executable code. This reads them.

The output that matters is the divergence set -- skills whose SKILL.md is
clean but whose referenced code is not. That is hidden behaviour by
construction: the documentation a reviewer or user sees says one thing, the
code that runs says another. It is the only shape of finding worth calling
malicious, and nothing in the pipeline could surface it before now.

Two constraints, both learned the expensive way this fortnight:

It does not convict. The rule engine is calibrated for SKILL.md prose; its
false-positive profile on JavaScript and Python is unmeasured, and this
project has already published one inflated number by assuming a rule meant
what it appeared to mean. The report labels its own output as leads to read
by hand.

A 404 is not absence. A referenced path that fails to fetch is counted
separately, because "names a script we could not retrieve" and "ships nothing"
are different facts, and collapsing them is how a blind spot gets reported as
larger than it is.

Extraction moves into malwar.research.references so both scripts share one
implementation with real test coverage. The tests are weighted toward negative
cases -- every rule defect shipped here was a false positive found by hand
afterwards, so the must-not-match set is deliberately larger than the
must-match set and covers the shapes that fooled the detection rules: remote
URLs, the user's own dotfiles, paths escaping the package, and prose that
merely resembles a path.

1752 passing, 34 new.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DNoTXU8k3pfSBzR7aJubqL
@Ap6pack
Ap6pack merged commit f8b65e3 into main Sep 11, 2026
5 checks passed
@Ap6pack
Ap6pack deleted the claude/scan-referenced-payloads branch September 11, 2026 16:38
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants