Plan uncertain work · preserve system knowledge · recover divergent branches
trace data · attack correctness · pressure-test assumptions · research decisions · prune context
Most agent instructions focus on producing code. This collection focuses on producing justified engineering decisions: grounded in repository evidence, constrained by explicit authority, and closed with verification.
Each skill is deliberately opinionated. Together they cover eleven recurring failure zones in long-lived codebases.
| When you need to… | Use | What it changes |
|---|---|---|
| Turn a fuzzy request into work another agent can execute safely | $deep-plan |
Writes a risk-ordered plan, not implementation code |
| Build durable understanding of a large or long-lived codebase | $steward-brownfield |
Maintains an evidence-backed project world model |
| Recover useful work from a stale, divergent, or oversized branch | $harvest-agent-branches |
Ports coherent slices without overwriting newer work |
| Explain where a value came from—or where its meaning broke | $trace-data-provenance |
Traces one specimen across every semantic boundary |
| Try to falsify a change that appears correct | $adversarial-review |
Runs bounded attacks and reports reproducible findings |
| Pressure-test a system from independent reasoning perspectives | $reasoning-codebase-review |
Coordinates investigators, red/blue challenge, and evidence-based judgment |
| Choose between consequential engineering options | $decision-recon |
Separates requirements from current evidence and preserves reversal conditions |
| Learn from a completed body of engineering work | $evidence-retrospective |
Reconstructs goals, diffs, verification, systemic patterns, and follow-through |
| Understand a change before making a human review call | $checkpoint-walkthrough |
Builds a concern-grouped trail, risk prompts, observations, and decision packet |
| Pressure-test an idea while changing direction is cheap | $idea-forge |
Attacks assumptions, defends the strongest version, and renders an honest outcome |
| Keep repository agent instructions useful without letting context sprawl | $context-pruner |
Verifies loading, accounts for every rule, and prunes only with evidence |
Start with the narrowest workflow that owns the decision. See ROUTING.md.
idea-forge -> decision-recon -> deep-plan
deep-plan -> implementation -> adversarial-review -> checkpoint-walkthrough
completed work -> evidence-retrospective
long-running system -> steward-brownfield throughout
Turns a rough engineering request into a plan that can survive fresh sessions and shallow review. It investigates the real repository first, sizes the work, exposes ambiguity, orders phases by risk, and produces self-contained sub-prompts with exact verification and stop conditions.
Use it when: the request spans several files, hides product decisions, or could easily become an unreviewable diff.
Use $deep-plan to investigate this repository and turn the request into a risk-ordered execution plan.
Treats project knowledge as durable infrastructure. It initializes or resumes a versioned world model, refreshes only affected knowledge, coordinates bounded specialist work, and checkpoints evidence so the next agent does not have to rediscover the system from scratch.
Use it when: continuity across sessions matters more than a one-off answer.
Use $steward-brownfield to resume this project safely and recommend the next highest-value step.
Recovers intent from abandoned or divergent agent work without assuming the whole branch deserves to land. It pins source and target evidence, builds a path matrix, decomposes work by behavior, chooses the least risky transfer method, and preserves recovery before cleanup.
Use it when: an old branch contains valuable work, but
mergeis too blunt an instrument.
Use $harvest-agent-branches to salvage coherent work from this branch onto current main.
Follows a concrete datum from authoritative input to delivered output. It checks identity, time semantics, units, missingness, fallbacks, lineage, and historical behavior—then identifies the earliest boundary where meaning becomes unproven or incorrect.
Use it when: a metric is wrong, stale, missing, duplicated, inconsistent, or simply impossible to explain.
Use $trace-data-provenance to trace this value end to end and identify the first unsafe boundary.
Attempts to falsify behavioral and security correctness rather than merely confirming the happy path. It derives an attack ledger from actual contracts and code, selects adaptive techniques, minimizes reproductions, and reports bounded residual risk instead of vague reassurance.
Use it when: a parser, workflow, API, or bug fix needs more than an ordinary code review.
Use $adversarial-review to attack this change with adaptive edge-case tests.
Orchestrates independent reviewers through pre-mortem, first-principles, inversion, Socratic, constraint, stakeholder, and analogical lenses. It then subjects normalized claims to separate red and blue challengers before the coordinator accepts, qualifies, or rejects them.
Use it when: an architecture or codebase needs more than one mental model—and the disagreements matter as much as the findings.
Use $reasoning-codebase-review to pressure-test this codebase through independent reasoning methods.
Turns technology, vendor, architecture, migration, and build-versus-buy choices into evidence-backed decisions. It frames hard gates before researching candidates, tracks claim freshness, exposes lock-in and lifecycle cost, tests sensitivity, and challenges the provisional leader before recommending it.
Use it when: the expensive part is choosing—and being able to explain when that choice should change.
Use $decision-recon to compare these options and produce a reversible, evidence-backed recommendation.
Reconstructs a release, milestone, sprint, migration, or multi-session feature from the evidence it left behind. It analyzes the whole change for architecture drift, integration seams, duplicated patterns, verification gaps, and follow-through—while keeping missing evidence distinct from missing work.
Use it when: the lessons live across several tickets or commits, and memory is too easy to rewrite after the fact.
Use $evidence-retrospective to review this milestone and propose sourced, owned follow-ups.
Guides a human from intent and surface-area orientation through a concern-grouped code trail, high-blast-radius questions, and safe ways to observe behavior. It keeps risk prompts separate from verified defects and ends with an explicitly bounded Approve, Rework, or Discuss decision.
Use it when: a diff needs to become understandable before it can become acceptable.
Use $checkpoint-walkthrough to show me what changed, where to look, and what remains a judgment call.
Pressure-tests a product, feature, architecture, workflow, or business concept before planning momentum makes it expensive to question. It separates actors and incentives, alternates attack with strongest-case defense, identifies discriminating experiments, and closes as Hardened, Clarified, Parked, or Killed.
Use it when: the cheapest useful deliverable may be a better idea—or permission to drop it.
Use $idea-forge to attack this idea, defend what survives, and render an honest outcome.
Treats always-loaded agent instructions as scarce infrastructure. It maps what each harness actually loads, verifies rules against the repository, tracks every instruction through a change ledger, prefers mechanical enforcement, and removes human-authored guidance only when evidence or explicit approval supports it.
Use it when:
AGENTS.md,CLAUDE.md, Copilot, Cursor, or linked rule files have become stale, contradictory, duplicated, or simply too large to trust.
Use $context-pruner to audit this repository’s agent instructions and remove context debt safely.
Install from a versioned release into the personal skills directory.
macOS or Linux
git clone --branch v1.0.1 --depth 1 https://github.com/chrisduvillard/codex-engineering-skills.git
mkdir -p "$HOME/.agents/skills"
cp -R codex-engineering-skills/skills/* "$HOME/.agents/skills/"
python -m pip install -r "$HOME/.agents/skills/steward-brownfield/requirements.txt"PowerShell
git clone --branch v1.0.1 --depth 1 https://github.com/chrisduvillard/codex-engineering-skills.git
New-Item -ItemType Directory -Force "$HOME/.agents/skills" | Out-Null
Copy-Item -Recurse -Force "codex-engineering-skills/skills/*" "$HOME/.agents/skills/"
python -m pip install -r "$HOME/.agents/skills/steward-brownfield/requirements.txt"For repository-scoped installation, copy only the required skill directories into
.agents/skills/ in that repository. Restart Codex after installation and invoke explicit-only
workflows with their $name.
Upgrade by replacing a skill directory from a newer release. Uninstall by removing only that skill's directory.
The eleven skills are different tools, but they enforce the same engineering instincts:
- Evidence before confidence. Read the repository, contracts, history, and runtime artifacts.
- Risk before tidiness. Detect breakage and secure boundaries before restructuring code.
- Narrow authority. Inspection does not imply permission to edit, publish, deploy, or delete.
- Concrete specimens. Trace one value, replay one failure, or minimize one counterexample.
- Mechanical closure. Finish with exact checks, preserved recovery, and explicit residual risk.
assets/
└── skills/
└── eleven README workflow illustrations
skills/
├── deep-plan/
├── steward-brownfield/
│ ├── assets/
│ ├── references/
│ ├── scripts/
│ └── tests/
├── harvest-agent-branches/
├── trace-data-provenance/
├── adversarial-review/
│ └── references/
├── reasoning-codebase-review/
│ └── references/
├── decision-recon/
│ └── references/
├── evidence-retrospective/
│ └── references/
├── checkpoint-walkthrough/
│ └── references/
├── idea-forge/
│ └── references/
└── context-pruner/
└── references/
Every skill is rooted at a SKILL.md. Supporting metadata lives in agents/; larger skills may also carry progressive references, scripts, tests, schemas, and templates.
The repository checks skill names, parsed frontmatter, local support links, agent metadata, Python syntax, validator regressions, and the Brownfield Steward test suite on pull requests and pushes to main.
Run the same checks locally:
python3 -m pip install -r requirements-dev.txt
python3 scripts/validate_skills.py
python3 scripts/validate_catalog.py
python3 -m unittest discover -s tests -p 'test_*.py'
python3 -m unittest discover -s skills/steward-brownfield/tests -p 'test_*.py'Released under the MIT License.
Changes should make a skill safer, clearer, or more falsifiable—not merely longer. A strong proposal includes the failure mode it addresses, the evidence behind it, the smallest instruction change that fixes it, and a way to verify the new behavior.










