Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
15 changes: 15 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,6 +2,21 @@

All notable changes to SkipHow 2.x appear in this file. Earlier release notes remain available on [GitHub Releases](https://github.com/mzored/SkipHow/releases).

## 2.14.0 (2026-09-02)

### Changed

- The frontier is bounded by the owner's result. `advancing-tracked-work` treats an open, unblocked item as takeable only when it lies on the path from live state to the requested result; an item beyond that destination is reported as takeable and deferred rather than resolved on the way. When nothing takeable reaches the result, the run stops with the human batch instead of filling the wait with enabling work the request did not name. Where the request names only the tracker, the result is recovered from the records' own parent outcome or the product brief. The always-loaded kernel states the same measure beside its continuation rule.
- `campaign-direction` recovers the premise from the owner's request and the owner's recorded decisions, not from a parent record the run or an audit wrote. The sentence that let security, money, recovery, and operational work produce no evidence of the result is gone; that work names the obstacle to the stated result it removes, like any other. Deferring an off-path direction is an outcome beside keep, simplify, replace, and retire, and replacing the architecture of off-path work is named as not a response. When deferral would carry the result past an unsettled risk or rollout consequence, that consequence is the one product question. With no takeable unit that reaches the result, admission stops; spare capacity admits nothing.
- The kernel extends the 2.13.1 provenance rule: a record's claim that something must precede the owner's result is a proposal on the same footing, and a record the run itself wrote carries only the authority of the request it served. A unit that must create a new prerequisite of its own before it can finish is a named trigger for `campaign-direction` and a stream anomaly in `execution-health`.

### Evidence

- Two owner-run installed Codex campaigns on 2.13.0 spent a day each on enabling machinery. One asked for the tasks blocking first payments; its real blockers integrated in about seventeen hours and the remaining thirty went to one backup-recovery lineage of eight tracker items, each new one a prerequisite for resuming the last, ending uncommitted on a defect. The run admitted that lineage at hour one to use free capacity, thirty minutes after computing a money path that did not contain it, on the strength of an audit finding that said recovery precedes traffic. The other asked to exhaust the takeable frontier of a tracker built from a complexity audit; it did exactly that, closing twelve tooling and evidence items over twenty-three hours while the item the audit had called most important waited on the owner, and nothing in the package made it say so or ask whether to go on. A third run of eight hours on a scoped bug-and-staging request showed no deviation and its result was accepted.
- `campaign-direction` was opened in the first campaign three times and each pass replaced the architecture of the same direction. The 2.13.1 kernel wording reached that run mid-way and fifteen more hours followed on the same lineage. That is one session showing the released text in context and not stopping the drift; the second shows the tracker-only request shape the package could not measure. The wording defects are readable in the files: premise recovered from records the run wrote, an exemption for recovery work, no defer outcome, and a frontier defined by blockers alone. Details are in `docs/evidence.md`.
- The transcripts show the per-lane rules of `execution-health` being applied in those runs. No numeric limit was added; a two-hour rule and a one-item-per-session rule were both refused again, see `docs/decisions.md`.
- One matched isolated Claude Code pair on a five-minute fixture showed exact 2.13.1 and this package behaving the same: closing the items on the payment path, continuing past a human-gated item rather than stopping at it, marking audit-derived infrastructure as proposed, and putting the backups-before-money question to the owner as a risk choice. That is a non-regression receipt; the drift lives in day-long installed runs, so the improvement stays `UNVERIFIED` until the owner's next long campaign. See `docs/evidence.md`.

## 2.13.1 (2026-08-31)

### Fixed
Expand Down
4 changes: 2 additions & 2 deletions SECURITY.md
Original file line number Diff line number Diff line change
Expand Up @@ -4,8 +4,8 @@

| Version | Supported |
| --- | --- |
| 2.13.x | Yes |
| 2.12.x and earlier | No |
| 2.14.x | Yes |
| 2.13.x and earlier | No |

Security review covers the packaged owner skill, its linked methods, host manifests,
marketplace metadata, continuity hook, release checks, and documented authority
Expand Down
2 changes: 1 addition & 1 deletion VERSION
Original file line number Diff line number Diff line change
@@ -1 +1 @@
2.13.1
2.14.0
8 changes: 7 additions & 1 deletion docs/decisions.md
Original file line number Diff line number Diff line change
Expand Up @@ -98,6 +98,8 @@ Three heavier alternatives were rejected. Matt Pocock's `grilling` asks across t

The controlled fixture kept the prompt, repository, host isolation, and model fixed while changing the package. Exact 2.13.0 made both findings `Ready`, including a capability absent from the authoritative product brief. Candidate revisions exposed three narrower failures: one kept both findings proposed instead of trusting the brief for the accepted core flow; one preserved the disputed item as proposed but deferred the decision; another recast it as a technical contract question. Those results produced the authoritative-brief exception, the current-result ask, and the capability-versus-implementation wording. A later candidate loaded the skill and kept the core repair `Ready` and the disputed capability `Proposed`, but still did not ask. An exact-package repeat did not load the skill and promoted both. The text defect and its point-of-use reach are fixed; reliable end-to-end behavior remains `UNVERIFIED`. Revisit this if receipts still promote a loaded proposal without asking, or if ordinary well-specified work starts acquiring unnecessary clarification rounds.

Version 2.14.0 extends the distinction after an installed campaign showed its gap. A record's claim that its work must precede the owner's result is the same kind of statement as a claim that a capability belongs in the product: a proposal until the request, an authoritative brief, or an owner decision adopts it. And a record the run writes during the request does not become an owner decision by being read back an hour later; it carries only the authority of the request it served. The first campaign in [current evidence](evidence.md) turned an audit finding into a parent item that said recovery precedes traffic, then recovered that item as the settled premise every time it rechecked direction.

## What the owner decided is a record, not a message

An answer the owner gives is a decision the project carries. Where the request authorizes a record, it is written where the work is tracked, with what it settled and the option they turned down, before anything depending on it is built. When the owner asks to settle what they want before work starts, `product-spec` turns that into a document they can read back: a vocabulary in their own words settled before the outcomes, the outcome stated as what a person will be able to do, each decision with its rejected alternative, and what is deliberately out of scope.
Expand Down Expand Up @@ -158,7 +160,11 @@ The owner required this correction to live inside SkipHow rather than depend on

Current comparable projects keep the heavier mechanisms SkipHow rejects. [`improve-codebase-architecture`](https://github.com/mattpocock/skills/blob/6654f6b60cd9d5be8b54c6fafe44346dabeb3b76/docs/engineering/improve-codebase-architecture.md) is an owner-invoked periodic check that ends with the owner choosing a candidate. [Paperclip](https://github.com/paperclipai/paperclip/blob/3623a369aa15c3ccb9d696f176049db4ed241248/docs/start/core-concepts.md) keeps goals, budgets, strategy approval, and board pause in a control plane. [Superpowers' plan executor](https://github.com/obra/superpowers/blob/b36e0829c6d0140e93cfef2ca599b1b07d4a7797/skills/executing-plans/SKILL.md) returns to review inside a mandatory plan workflow. SkipHow adopts none of those control mechanisms.

Revisit this if the method opens on healthy product slices with no shared machinery, asks the owner to choose a technical correction, re-argues settled direction without new evidence, stops independent lanes, starts more work than the current integration path can absorb, expands into an unrequested repository survey, or a comparable campaign reproduces the same drift.
Version 2.14.0 answers the first real campaigns run on that method, and they reproduced the drift with the method open. Two installed Codex campaigns on 2.13.0, described in [current evidence](evidence.md), spent thirty and twenty-three hours on enabling machinery while the product result had come to wait on the owner's own steps; the first against its stated result, the second under a request that named the frontier itself as the result. `campaign-direction` was opened three times in the first and each pass replaced the architecture of the same backup-recovery direction. Retire was among its outcomes, and three properties of its text bear on why no pass reached it. It recovered the premise from the parent record and recorded decisions, which the run itself or an earlier audit had written; the parent item said "before traffic" because the run wrote it that way at hour one. It exempted security, money, recovery, and operational work from producing customer-visible evidence, which was the class of work observed. Its outcomes were keep, simplify, replace, or retire a technical direction, with the owner asked only when no technically adequate option remained, so a rebuild from clean integration was always available and the sequencing question never reached anyone. `advancing-tracked-work` supplied the admission: the frontier was whatever was unblocked, the rule on a human block was to set it aside and carry on with what remained takeable, and nothing measured the remainder against the result. The 2.13.1 provenance sentence reached the first run mid-way and fifteen more hours followed, because it named capability scope and said nothing about a claim that work must come first or about records the run had written itself.

The change bounds the frontier by the requested result, adds defer as a direction outcome, deletes the exemption, extends the provenance rule to sequencing claims and to the run's own records, and names one further observable signal, a unit that must create a new prerequisite of its own before it can finish. Wayfinder in `mattpocock/skills` already had the frontier half of this, a ticket found to sit beyond the destination is ruled out of scope rather than resolved on the route, and the 2.6.0 adaptation had not taken it; it is taken now as an idea, in SkipHow's words. Two heavier answers were refused again. A whole-request time bound, which the owner's own follow-up analyst proposed as two hours without progress toward the user scenario, is a number no run can justify across projects. One item per session, which wayfinder keeps, was rejected in 2.6.0 and the owner's own prompt for the second campaign said not to stop after one task; the defect was never that the runs did several items but that the items were off the path. The product choice underneath, stop and hand back when the result waits on the owner and only enabling work remains, is the owner's, and they made it after seeing the runs: SkipHow is meant to act as the technical director, not to spend days on work a second agent then calls low priority.

Revisit this if the method opens on healthy product slices with no shared machinery, asks the owner to choose a technical correction, re-argues settled direction without new evidence, stops independent lanes, starts more work than the current integration path can absorb, expands into an unrequested repository survey, a run stops and hands back a batch while an enabling item the owner's own request named sat takeable, or a comparable campaign reproduces the same drift.

## A design method opens on what the project holds, not on how important the choice feels

Expand Down
17 changes: 17 additions & 0 deletions docs/evidence.md
Original file line number Diff line number Diff line change
Expand Up @@ -224,8 +224,25 @@ Candidate 2.13.1 runs exposed useful boundaries rather than a clean pass. The fi

The receipts show that the old package allowed the promotion and that loaded candidate wording can preserve the distinction, but not that the complete behavior is reliable. Claude behavior remains `UNVERIFIED`; no Codex behavior run was accepted.

### A campaign kept building after its result came to wait on the owner

Three owner-run installed Codex Desktop sessions over 2026-08-31 and 2026-09-01 were read from the host's own transcripts for this change, with per-message usage and timestamps. The transcripts are private and are not retained here.

The first asked the run to close the tasks blocking first payments and real traffic, granted production, and asked for human-only steps to be batched. It ran 36.8 hours, spawned 61 delegates, issued 1,784 delegate waits, and processed 137 million input tokens, 99 per cent cached. The 2.13.0 kernel governed from the start and the 2.13.1 kernel reached its context at hour 22. Its two genuine blockers, a public OAuth defect and a receipt-contact rule, were integrated by hour 17. From hour one it also admitted a backup-recovery item, in its own words to use free capacity, thirty minutes after it had computed the money path without that item; the item's "before traffic" premise came from an audit finding recorded two weeks earlier. That lineage then produced eight tracker items, each new one written as a prerequisite for resuming the last, four independent reviews, three architectures, and about 1,700 uncommitted lines when the owner returned and paused it. `campaign-direction` was opened in context three times and each pass replaced the architecture. Fifteen hours of that followed the 2.13.1 wording. A fresh session the next day, asked by the owner why it had taken so long, answered in ten minutes that the lineage protected a CI proof of disaster recovery rather than the product, and proposed a plan whose first item was the one human step the run had asked for at minute thirty-five.

The second asked the run to exhaust the takeable internal frontier of a tracker that an earlier session had built from a complexity audit. It ran 22.9 hours as one orchestrator with forked delegates and closed twelve items: CI routing, size budgets, evidence packaging, Markdown tooling, benchmark isolation. The audit those items came from had said the evidence system was the excess and that a device test with real people, recorded as a human-gated item, was the most important next step. A fresh session the next day, asked whether there was anything to play yet, counted 75 commits and roughly 9,800 lines since the audit with no new game behavior.

The third asked for reproducible bugs and staging blockers on one campaign, then a staging release. It ran 8.2 hours, mostly on the project's own release gate and two flaky tests, released to staging, and the owner checked staging the next morning and accepted it. It shows a scoped request staying scoped, and nothing else.

Counted in whole sessions: one shows drift from the stated result under the shipped text, and shows the 2.13.1 kernel in context and not stopping it. The second complied with its request as written, because the request named the frontier itself as the result; what it shows is a request of that shape producing a day of tooling with no product evidence and no sentence in the package that would make the run say so or ask whether to continue. The transcripts show the per-lane rules of `execution-health` being applied throughout, twenty-minute checkpoints and three-attempt stops included; what no text supplied was a measure of the remaining takeable work against the owner's result once that result waited on the owner. The wording defects are readable in the files and are listed in the changelog and decision history.

Version 2.14.0 bounds the frontier by the result, adds defer as a direction outcome, deletes the recovery exemption, extends the provenance rule to sequencing claims and to the run's own records, and names prerequisite-spawning as a signal.

One matched Claude Code pair was then run on a throwaway shop repository whose tracker held two takeable items on the payment path, one human-gated item on it, and two audit-derived infrastructure items off it, with the same Get5Stars-shaped prompt, settings sources and MCP disabled, the package passed as a session plugin, and the init event naming Claude Code 2.1.258, Opus 5, and the exact package path each time. Exact 2.13.1 and the candidate both did the same thing in about five minutes: closed the two path items, continued past the human gate rather than stopping at it, found that the payment adapter never charged anything and recorded that as the real blocker, marked both audit items proposed, put the backups-before-money question to the owner as a risk choice with a recommendation, and stopped with one batch. Both opened the frontier method. The pair shows that the new wording keeps a run moving through a human gate and does not add a question or a gate; it does not show the improvement, because the released text already behaved correctly on a five-minute fixture, as the 2.13.0 pairs also found. What the installed campaigns show and the fixture cannot is a run twenty hours in, holding records it wrote itself, with free delegate capacity and nothing left on the path.

## Still unverified

- Whether the 2.14.0 frontier bound and defer outcome stop a long run when its result waits on the owner and only enabling work remains. One installed 2.13.0 campaign shows the drift with the 2.13.1 text in context, and a matched five-minute pair shows both packages already behaving correctly at that scale, so the fixture is not where the defect lives. The line closes only on the owner's next long installed campaign.
- The outside read of a consequential design decision. Ten runs made the decision well and none took an outside read. Codex had the method open in all five of its runs; no Claude session in the pass opened it at all. Three kernel wordings changed nothing on either host. The rule is stated and does not execute.
- Delegation under the shipped wording. No fixture run in the pass spawned a delegate for any reason; the largest fixture, six capabilities over 2,725 lines, was carried in one pass by both hosts. The installed sessions above show delegation happening at scale but with the governing methods absent from context, so they say what delegation costs and not whether the wording works.
- Whether the 2.12.0 observation rule reduces root context traffic or the reconciliation rule prevents integrated working state from accumulating. Both changes answer installed failures, but neither has run in a comparable session.
Expand Down
Loading