Show built-in Claude Code skill activations without scoring them - #223
Merged
Merged
Conversation
…g them Skills the CLI ships inside its binary (claude-api, debug, ...) stay discoverable in every run, isolated or not. The detector dropped any name outside the declared and user sets, so a bundled skill winning a prompt read as nothing firing. claude-code now reads the skills it exposed from the init event; any that are neither declared nor the user's own are auto-bundled. Their activations land in AttemptRecord.bundled_activated and an 'Auto-bundled' line under the activation table, and never count against activates:. Closes #220
A plugin skill whose SKILL.md sits under a bundled skill's basename was also credited to the bundled skill. detect_bundled now strips plugin paths first, as detect does.
…ndled 'Auto-bundled' read as something Caliper bundled. The report now says 'Built-in skills claude-api 2/4 (ship with claude-code; not scored)', and the record field is builtin_activated.
This was referenced Oct 1, 2026
Merged
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes #220.
Skills the agent CLI ships itself (Claude Code's
claude-api,debug, …) stay discoverable in every run, isolated or not. Until now, one firing read as "nothing fired": the detector dropped any name outside the declared, user and plugin sets.Change
claude-codereads theskillsitsinitevent lists. Any that are neither declared nor the user's own are recorded as built-in (AttemptResult.builtin_skill_names). Other backends returnNone, meaning not observed.ActivationDetector.detect_builtinapplies the same matching rule to those names and records the result in a newAttemptRecord.builtin_activated. It is never scored, soactivates: []still passes.Built-in skills claude-api 2/4 (ship with claude-code; not scored)line under the activation table, plus abuilt-inrow in each attempt's details.Verification
test_attempt.py,test_claude_harness.pyandtest_reporter.py.activated=[],activation_passed=true,builtin_activated=["claude-api"], and the report showsBuilt-in skills claude-api 2/4while activation stays at 100%.What it looks like
Taken from the #217 probe run (
caliper report --verbose).A new line under the activation table. It only appears when a built-in skill fired, and it also shows on specs that have no
activates::A
built-inrow in the attempt details. These panels show for tasks that didn't fully pass, or for every task with--verbose:And in the results JSON, per attempt: