A read-only static audit skill for Python projects: one command, one Markdown report, one A–F grade.
python-health-audit is a Kilo / Claude Code skill that any developer (lead, solo, acquirer) can invoke to get an instant, opinionated health check of a Python codebase. It runs four industry-standard linters through uvx — no virtualenv, no install — and writes a single python_health_report.md at the root of the audited project.
| Without the skill | With the skill |
|---|---|
| Remember 4 CLI invocations + flags + exclusions | One natural-language request |
| Stitch outputs from Ruff, Vulture, Radon, Pylint by hand | One Markdown report with deterministic A–F grade |
Risk of ruff --fix / pylint --fix mutating source |
Strict read-only guarantee |
| Ad-hoc structure every time | Stable 5-section template + 3-item action plan |
The skill works in any agent runtime that supports the skills format — Kilo, Claude Code, Roo Code, Cline, Cursor, Continue, etc.
Recommended (one-liner):
npx skills add laurentvv/python-health-auditThe CLI accepts the GitHub shorthand owner/repo and auto-discovers SKILL.md at the repo root.
Per-runtime:
# Kilo
kilo skills add laurentvv/python-health-audit
# Claude Code
claude skills add laurentvv/python-health-audit
# Roo Code (installs into .roo/skills/ of the current project)
npx skills add laurentvv/python-health-audit
# Global install (shared across all projects on the machine)
npx skills add -g laurentvv/python-health-auditGlobal installs use symlinks so subsequent skills update calls keep every project in sync.
Open your agent in the Python project you want to audit and ask, in plain English (or French):
I'm taking over a 3-year-old Flask project. Can you give me a full health check
before I start refactoring? The code lives in ./backend. Tell me where the
technical debt is.
The agent will:
cdinto the target project (or use theworkdirparameter)- Run sequentially, capturing all stdout + stderr in memory:
uvx ruff check .— local dead code (unused imports, unused vars)uvx vulture . --min-confidence 80— global dead code (unused symbols)uvx radon cc . -a -nc— cyclomatic complexity hotspotsuvx radon mi .— Maintainability Index per fileuvx pylint --disable=all --enable=duplicate-code (...)— copy-paste detection
- Compute an A–F grade from the metrics (deterministic heuristic, see below)
- Write
python_health_report.mdat the root of the audited project
That's it. No source files are touched.
# Python Health Report — backend
Generated on 2026-06-04 11:09 by python-health-audit.
## 1. Executive Summary
- Global grade: C
- Reason: Grade C assigned — Ruff reports 21 findings (just above the 20
threshold), but only 1 C/D hotspot (`login` in `auth.py`) and average MI
at 71.64 keep the project out of D/F territory.
## 2. Dead Code
### 2.1 Local — Ruff
21 findings (20 unused imports F401 + 1 undefined name F821). All fixable
mechanically; no auto-fix applied per audit policy.
### 2.2 Global — Vulture
13 entries at confidence ≥ 80%.
> ⚠️ Vulture produces false positives by construction (global static
> detection). Verify each entry before removal.
## 3. Complexity Hotspots (Radon)
| Rank | File:Line | Block |
|------|-----------------|--------|
| D | backend/auth.py | `login`|
## 4. Code Duplication (Pylint)
No duplication detected — Pylint `duplicate-code` reports 10.00/10.
## 5. Recommended Action Plan
1. **Fix the F821 undefined name in `backend/app.py:18`** — latent runtime crash.
2. **Reduce complexity of `login` (rank D, 22)** — extract validation chain.
3. **Purge the 20 unused imports flagged by Ruff F401** — mechanical cleanup.The grade is computed deterministically. First matching rule wins (A → F):
| Grade | Criteria |
|---|---|
| A | Ruff = 0 AND 0 C/D/E/F hotspot AND average MI ≥ 80 AND 0 Pylint duplication |
| B | Ruff ≤ 5 AND 0 E/F hotspot AND average MI ≥ 65 |
| C | (Ruff ≤ 20) OR (≤ 3 C/D hotspots); average MI ≥ 50 |
| D | ≥ 1 E hotspot OR average MI ∈ [30, 49] |
| F | ≥ 1 F hotspot OR average MI < 30 OR Vulture > 20 entries |
- Average MI = arithmetic mean of
radon miscores per file. - Hotspot = function/class graded C, D, E or F by Radon cc (A and B hidden).
- Duplication = at least one pair returned by Pylint
duplicate-code.
- Strict read-only — no
Edit/Write/Set-Contenton.pyfiles. The only artifact written is the report itself. - No auto-fix — never invokes
ruff --fix,autoflake,radon rawwith rewrite, orpylint --fix. - Silent execution — no progress messages in chat. Only the path of the generated report is returned.
Tested against 2 realistic prompts × 2 configurations (with vs without the skill), on a synthetic Python codebase with intentional debt (dead code, complexity hotspot, copy-paste).
| Configuration | Pass rate | Detail |
|---|---|---|
| With skill | 100% (13/13) | eval-1: 8/8, eval-2: 5/5 |
| Without skill (baseline) | 29% (4/13) | eval-1: 3/8, eval-2: 1/5 |
| Delta | +71 pp | — |
The baseline (no skill) consistently fails to:
- Produce the expected
python_health_report.mdfilename (agents inventAUDIT.md,AUDIT_REPORT.md, etc.) - Follow the 5-section template and the 3-item action plan
- Display the A–F grade
- Include the Vulture false-positive warning
Reproduce: see evals/evals.json and the iteration-1 artifacts archived in the v1.0 release.
Single-skill repo, flat layout (auto-discovered by npx skills add):
python-health-audit/
├── README.md
├── LICENSE
├── .gitignore
├── SKILL.md # frontmatter + role / objective / steps / grading / format / constraints
└── evals/
└── evals.json # 2 benchmark test cases
python-health-audit/ ├── README.md ├── LICENSE ├── .gitignore └── python-health-audit/ ├── SKILL.md # frontmatter + role / objective / steps / grading / format / constraints └── evals/ └── evals.json # 2 benchmark test cases
## Compatibility
| Runtime | Status |
|---|---|
| Kilo | ✅ supported |
| Claude Code | ✅ supported |
| Roo Code (`.roo/skills/`) | ✅ supported |
| Cline, Cursor, Continue, … | ✅ any agent that reads the `SKILL.md` frontmatter format |
Required on the host machine:
- [`uv`](https://docs.astral.sh/uv/) (provides the `uvx` command) — **or** `pipx` as fallback
- PowerShell 7+ on Windows, bash on Linux/macOS
## Contributing
This is a single-skill repo (flat layout). To iterate:
1. Edit `SKILL.md` at the repo root.
2. Add / modify test prompts in `evals/evals.json`.
3. Re-run benchmarks with the [`skill-creator`](https://github.com/Kilo-Org/skills/tree/main/skill-creator) workflow.
4. Open a PR — describe the change in the skill's behavior, not just the file diff.
## License
MIT — see [LICENSE](LICENSE).