Skip to content

v0.1.0 ship-readiness: acceptance criteria, remediation cycles, and the security half that needs no breaking change - #252

Open
letuhao wants to merge 148 commits into
release/v0.1.0from
fix/v0.1.0-release-gaps
Open

letuhao wants to merge 148 commits into
release/v0.1.0from
fix/v0.1.0-release-gaps

Conversation

@letuhao

@letuhao letuhao commented Sep 13, 2026 •

Copy link
Copy Markdown
Owner

Closes the v0.1.0 ship-readiness gaps that can be closed by engineering, and records what cannot.

Driven by docs/plans/2026-09-13-v0.1.0-ship-acceptance.md — the PO-approved ship bar — through the new Remediation Cycle. This does not ship anything. The release fires on a pushed vX.Y.Z tag, and no tag is pushed here.

Two new standards, both enforced

  • Plan Acceptance Criteria (AC-1..7) — a plan states what must be TRUE above its board of what someone intends to DO. Written because a release candidate reached main with 23 rows ticked, every gate green, and "can this ship?" had no answer.
  • Remediation Cycle (RC-1..6) — a criterion reaches met only through a recorded cycle: investigate, post issues, fix, PROVE, update the criterion.

Both gated, both self-tested (18/18 and 13/13), both bitten on live content.

Security — the serious part

Measured across nine lockfiles rather than trusting CI's summary: 2 critical, 52 high, 68 moderate, 22 low. The split matters more than the total. The frontend's critical and high are dev tooling. The 43 highs across four gateways are server runtime on the request path — CRLF injection in http-proxy-middleware, lodash code injection, ws memory disclosure.

This PR takes the non-breaking half: high severity 52 → 29, cms-frontend cleared, datasets 2.21.0 → 5.0.1 with pip-audit clean. Proved by effect, not by lockfile diff: npm ci plus a real build and suite for both gateways, 202 and 286 tests green.

Issues #247, #249, #250 carry the semver-major work and the rsa advisory that has no fix.

CI legs that had stopped checking

  • govulncheck had not run at all. The job pins Go 1.25, so setup-go pins the toolchain, and govulncheck@latest now needs 1.26 — the install failed and the scan loop never executed, while the check looked exactly as it looks when it finds something. Raised, plus a guard so scanning zero modules can never read as finding nothing.
  • MinIO moved off Docker Hub. DB live-smoke was exiting 125 — docker refusing to pull, not a test failing. Repointed to quay.io and pinned, at all six call sites.

Tests

Two chat-service tests were red because the guard asked the wrong question: the corpus directory is tracked for two notes while the recorded runs are not, so CORPUS.exists() said present and the loader found nothing. Absent is not truncated. Now 3966 passed, 9 skipped, zero failed, up from 3960 with 2 failures.

Verified in a real browser

Persona journeys pass against a stack rebuilt from this commit — image verified, not the build log — and the Campaigns rename renders correctly in English and Vietnamese with heading and nav agreeing.

The stack rebuilds again, and two gates that could not report their own failure

Asked for a full infra rebuild before final testing; it failed, and the background task reported exit code 0 while doing it. docker compose build does not do that — reproduced on a throwaway project, a failing target exits 1.

  • pgvector (infra: pgvector image fails to build — Alpine 3.24 dropped clang19/llvm19, and the pin was the wrong major anyway #255). postgres:18-alpine is a floating tag and moved to Alpine 3.24, which ships no clang19/llvm19. The pin was also already wrong: the image's Postgres is configured with LLVM_CONFIG=/usr/lib/llvm21/bin/llvm-config, so pgvector's JIT bitcode was being built against a different LLVM major than the server that loads it. Now derived from pg_config. Proven live — CREATE EXTENSION vector → 0.8.1 in the running container, not from a build log. Correction posted on infra: pgvector image fails to build — Alpine 3.24 dropped clang19/llvm19, and the pin was the wrong major anyway #255: I first called this release-blocking. It is not — release-targets.py returns 33 services and postgres is not among them. It blocks local and CI stack bring-up, which blocks the journeys and the authoring run, but not the publish.
  • The other 33 failed on truncated downloads, not bad packages: RemoteDisconnected, IncompleteRead, a half-read JSON index, and a wheel failing its hash at 32.8/818.2 kB — the hash check working. 34 concurrent builds saturated the link and the in-Dockerfile 5-attempt retry burned all five while it stayed saturated. Rebuilt in batches of four: all 33 release targets built, zero failures, 42 containers up and healthy.
  • all-gates could not say why a gate failed. It reported one by quoting [emit-0013] SELFTEST PASS ... (non-vacuous) — a success banner offered as the cause of a failure. gate-wiring-gate printed the first non-empty line of failing output, and this repo's gates print a self-test banner before their real work, so the quoted line was guaranteed to be the wrong one. It now quotes the tail.
  • Underneath it, a guard that was dead code. emit-migration-0013-lint.sh did text=$(cat "$f"); rc=$? under set -euo pipefail — errexit kills the script at the assignment, so rc was never read and its FAIL message could never print. Bitten both ways with a cat shim: the committed version prints the banner and nothing else; the fixed one names the file, and skips a file that genuinely vanished mid-sweep rather than going red on another gate's normal operation.

The changelog — the release gate now passes

[0.1.0] is written from the 39 user-facing commits in v0.1.0-rc.1..HEAD, verified in the exact mode the workflow invokes (VERSION="${TAG#v}" → changelog-gate.py --release 0.1.0):

changelog-gate: structure + release 0.1.0 OK (3 section(s)).   EXIT=0

Bitten both directions against copies via the gate's own --path, so CHANGELOG.md was never mutated to test it: section removed → exit 1; section present with only its five headings → exit 1 ("an empty section is the same failure as a missing one"); as written → exit 0. The hollow case is the one worth checking, because a bare heading is what a hurried release adds.

The section records what is knowingly not fixed beside what is — rsa with no published fix (#250), four gateways staying on NestJS 10 (#247), and 17 of 18 locales machine-translated and unread by a native speaker. A changelog listing only wins is the same class of artifact as a README claiming unbuilt features.

(A premise in #251 did not hold and the measured numbers are used instead: the issue records 100 commits / 36 feat-fix on main; today main has 65 (26), and the range that will actually be tagged has 89 (39).)

What is NOT done — and what changed about it

Nothing mechanical is left in the way of shipping, and that is a different statement from "ship it".

Every rehearsable step of the release act has now been rehearsed: the tag shape, the version derivation, the changelog gate, the 33-target matrix, and the 33 images building from source. Pushing the tag IS the ship act, and the PO holds it. No tag is pushed here.

Four criteria still need a person, and no amount of engineering moves them:

needs
T7 a named person accepting rsa RUSTSEC-2023-0071 in writing — there is no version to move to (#250)
T8 a native read of the ten new Vietnamese strings; placeholder-perfect and wrong passes every gate
T9 a decision on 战役 vs 活動 — zh-CN renders Campaigns as a military campaign and already disagrees with zh-TW
T11 the live authoring run: whether an author can plan and draft a book with canon intact (AC-10)

🤖 Generated with Claude Code

letuhao and others added 26 commits September 13, 2026 05:34
The remediation branch was verified almost entirely at the unit level, which is
the wrong shape of proof for what it changed. The defects it fixed were found by
a browser driving a running stack, and its largest single change -- ten strings
across eighteen locales -- is 197 pieces of user-facing text nobody has read.

Eight rows. Four produce evidence rather than code: CI with its legs named, the
full suites re-run on the merged state, the persona journeys against a rebuilt
container, and the rename verified in a real browser. One adds a gate for
placeholder loss in translations, which is a real bug class rather than a style
issue -- a dropped {{target}} renders a sentence with a hole in it. One settles
the two red chat-service tests, since "pre-existing" is a cause and not a
disposition. Two need a person: a native read of the new Vietnamese, and the
Chinese label, which currently means a military campaign and disagrees with the
Traditional Chinese locale.

PR #246 is not to be merged until this board is closed. It carries the closing
keywords for fourteen issues, so merging it is the release act.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…t cut it mid-sentence

goal-prompt.py emits only the RESUME's FIRST line. A wrapped one was truncated
at 'CI was', which is worse than no RESUME: it reads as a complete instruction
and is not one.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
dep-vuln.yml pinned go-version 1.25, so setup-go pinned GOTOOLCHAIN=local, and
govulncheck@latest is now golang.org/x/vuln v1.8.0 which requires go >= 1.26.0.
The install step failed and the scan loop under it never ran. The Go ecosystem
has been unscanned, while the check looked exactly like it looks when it finds
something -- advisory red, exit 1.

That indistinguishability is the real defect, so the fix has two halves. The
toolchain is raised to 1.26, which is safe here because the highest `go`
directive across the repo's go.mod files is 1.25.0. And the step now counts the
modules it scanned and exits 2 with an explicit error when that count is zero,
because scanning nothing is not the same as finding nothing.

Bitten by extracting the step body from the workflow itself and driving it with a
stubbed govulncheck: zero modules exits 2, two modules clean exits 0, two modules
with a finding exits 1. Before this change the first and third were both 1.

Not verified here: that Go 1.26 resolves in setup-go is taken from the tool's own
requirement message. The next run of this workflow settles it; if it does not
resolve, pin govulncheck to its last 1.25-compatible release instead.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
CI on PR #246 is not one red thing. Three legs fail for unrelated reasons and
none of them is a test disagreeing with the code:

- DB live-smoke exits 125 because docker cannot pull minio/minio at all.
- classify deploy refuses an EMPTY changed-file list, which is correct: the two
  branches differ by 160 files, so the diff did not resolve.
- chat-service unit is red for exactly the two tests T7 already covers, which
  makes that row blocking rather than cosmetic.

And the advisory legs split in two: datasets 2.21.0 has a published fix and is a
bump, while the rsa Marvin Attack advisory says outright that no fixed upgrade
exists. The second kind is a decision, not work, so it is written as one.

None of it was introduced by this release, verified by diffing main against the
release branch for every lockfile and manifest pattern: zero matches.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…broken

Two corrections, both found by checking rather than assuming.

This plan opened by saying PR #246 must not be merged until the board closes. It
was already merged when that was written -- 22:17:45Z, merge commit 0d56840 --
and all fourteen issues are closed, leaving the repository with no open issues.
So every gap on this board shipped. That makes the work more urgent, not less:
the unreviewed strings are in front of users and the two red tests are red on
main. Recorded rather than edited away, because this plan's subject is the
difference between what was verified and what was assumed, and its own first line
was assumed.

And classify deploy is not a workflow bug. The theory in its row -- a shallow
fetch with no merge-base -- did not survive the log, which shows the full
refs/heads fetch and no git error. Reproduced against the exact commit CI used:
the three-dot range gives 0 files and the two-dot gives 160, because
merge-base(origin/main, MERGE) resolves to the merge's own second parent. The PR
was merged into main while its own checks were still running, so origin/main
moved to contain the change under test.

The script was correct throughout. BDR-82's reach floor refused to call an
unresolved diff a patch, which is exactly what it is for.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…1..7)

On 2026-09-13 a release candidate reached main with 23 remediation rows ticked,
every one carrying pasted bite evidence, every gate green -- and "can this ship?"
had no answer. The PO: "this work is ad-hoc ... what have we done and not,
without ACs we cannot decide this repo can ship v0.1.0 or not".

Nothing had failed. A board records intentions carried out; it says nothing about
whether the result is acceptable, because acceptable was never written down.

Every plan now declares acceptance criteria above its board: falsifiable
statements about the product, each naming a verification method that is an
artifact someone can re-run, each tied to the rows that serve it, each carrying a
status. A met tick carries its evidence and a waiver names its author.

"Reviewed" and "tested" are rejected as verification methods. They name no
artifact anyone can re-run or read, which is the entire point of the column.

Scoped to plans dated 2026-09-13 or later, read from the filename, and the gate
PRINTS how many it skipped -- 511 today. Retrofitting 511 closed plans would
paint a wall of red and teach everyone to scroll past the gate, and a gate people
ignore enforces nothing.

Composes with Non-Vacuity rather than duplicating it: AC says which checks must
exist, NV says those checks must be able to fail.

Self-test 16/16. Bitten twice on the live plan: claiming a criterion met with
"reviewed" as its method reddens AC-3 and AC-5 together, and deleting the
criterion that covers three rows reddens AC-7 by name. Restored byte-exact.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…aft needed

The ship bar did not exist, which is why the ship question had no answer. Twelve
criteria, each stated so it can be false, each naming what settles it. DRAFT: the
PO writes the bar, and nothing is built or published until they say so.

The honest assessment is three met, five not met, four unknown. The three met are
real work: the claims are accurate, the naming is accurate, and the human-sim
findings are closed with evidence. The four unknown are the ones that matter --
whether the product does its core job, whether it builds and publishes at all,
and whether it can be rolled back. Those are not failing, they are unasked, and
the release act itself has never been rehearsed.

Writing it immediately found a hole in the standard written an hour earlier. The
gate rejected the draft for using an "unknown" status that was not in its token
list. The rejection was right about the list and wrong about the world: "we
checked and it is false" and "nobody has ever checked" are different states, and
collapsing them hides the more dangerous one. A failure is a known quantity; an
unasked question is not.

So the token was added rather than the document softened -- the same distinction
this repo already enforces for a skipped CI leg versus a passed one, and for a
scanner that scanned zero modules. Self-test 16/16 to 18/18.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…ROVED (RC-1..6)

PO approved all twelve v0.1.0 ship criteria. The standing verdict is recorded in
the document itself: not shippable, five criteria unmet and four never assessed.

A criterion now reaches "met" only through a recorded cycle with all five stages
-- investigate, post issues, fix, prove the fix, update the criterion. Each stage
is there because skipping it has failed in this repo: seven premises in the
2026-09-12 plan were wrong and were corrected by looking rather than recalling; a
fix nobody can find is a fix nobody reviews; "I fixed it" is a claim where a
watched red-then-green is evidence; and a fix that moves no criterion did not
move the bar, which is the exact failure that created the acceptance-criteria
standard a few hours ago.

The loop is meant to run many times. A cycle whose impact is "AC-9 still fails,
here is why" is a complete cycle -- the gate asks for five stages, not good news.

Its own self-test caught it rejecting every correctly-formed cycle on the first
run: Proof's content is the fenced block below the label, so its inline value is
empty by design, and the emptiness check did not know that. Fixed in the check
rather than in the fixtures. 13/13.

The three already-met criteria are marked pre-cycle rather than back-filled with
invented cycles.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…breaking changes

The PO: "this repo never pass security check, that is serious problem, and we
need serious upgrade libraries". Measured rather than assumed: 2 critical, 52
high, 68 moderate, 22 low across nine npm lockfiles.

The split matters more than the total. The frontend's critical (vitest) and high
(vite) are dev tooling, reachable only when a developer's test UI or dev server
is listening. The 43 highs across the four gateways are server runtime on the
request path -- CRLF injection in http-proxy-middleware, which is the proxying
layer itself, plus lodash code injection and ws memory disclosure. That is the
serious half.

This cycle takes the non-breaking half: npm audit fix with --package-lock-only
across six packages, so lockfiles moved and no manifest did. High severity 52 to
29, cms-frontend cleared entirely, and datasets 2.21.0 to 5.0.1 with pip-audit
confirming no known vulnerabilities.

Proved rather than assumed, because a lockfile change nobody installs is not a
verified fix: npm ci and a real build for both major gateways, 202 tests green on
api-gateway-bff and 286 on ai-gateway.

Issues #247, #248, #249, #250 record what remains. AC-6 moves from not met to
partial -- the semver-major upgrades are untouched and rsa still needs a named
person to accept it, because no fixed version exists.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…quay.io (T10)

DB live-smoke was failing with exit 125: docker refusing to pull minio/minio at
all, not a test failing. docker.io/minio/minio is no longer anonymously pullable;
MinIO publishes to quay.io.

Verified both ways rather than guessed: manifest inspect is DENIED on Docker Hub
and OK on quay.io. Then proved by effect rather than by metadata -- ran the CI
step's exact command, the image pulled, the container started, and the
workflow's own health loop reported live after two attempts.

Pinned rather than :latest, which is the real lesson here. An unpinned tag from a
registry that can revoke or move it is how a green pipeline turns red with no
commit to blame, and that is precisely what happened. Pinned to the release
infra/docker-compose.meta-ha.yml already used, so CI and that compose agree on
the version under test.

Fixed at all six call sites, not only the one that failed: the CI workflow, three
infra compose files, the RAID test-infra template and the Antithesis compose.
Every one would have failed identically the moment it ran. One defect with six
call sites is one commit.

Not established here: that the archive-worker smoke passes against this image.
That needs the full CI environment. The image starting and answering health is
what can be shown locally.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Yesterday's fix added ten keys across seventeen languages -- 197 strings produced
by a local model and read by nobody. Several carry runtime interpolation, and a
dropped placeholder is not a quality opinion, it is a broken string: i18next
renders what is left, so a Vietnamese author would see "Asked for ~ words,
delivered (%)" while every test stays green, because tests assert on keys and the
English bundle is fine.

i18n_translate.py already checks this at generation time. Nothing checked it at
rest, which is where a hand-edit, a bad merge, or a filtered regeneration lands --
and all three happened to these files in one day.

The audit came back clean, which is the useful part: 150,365 translated strings
across 35 namespaces and 17 locales keep every placeholder. The machine
translations did not drop an interpolation, and that is now established rather
than hoped.

Bitten on a real shipped string: removing two placeholders from the Vietnamese
length report reddens it by name and prints both sides. Restored byte-exact.
Self-test 12/12, including the cases that stop it over-firing -- reordered
placeholders pass, a format spec is the same placeholder, and a key merely
missing from a locale is left to i18n-completeness-gate, because two gates
reporting one defect teach people to read neither.

What it does not establish: whether the strings mean the right thing. A
placeholder-perfect mistranslation passes cleanly. That half of AC-8 needs a
person.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
I ran doc-language-gate --staged, captured exit 1, and committed anyway because
the commands were chained with ';' rather than '&&'. Rule 5 says check the exit
code; checking it and then ignoring it is worse than not checking, because the
log now shows a gate was consulted.

Two real causes behind the red, both fixed here rather than suppressed:

The gate's self-test fixtures used real Vietnamese where the subject under test
is placeholder SETS, not language. Synthetic markers test exactly the same logic
and keep the source file English, so no pragma is needed at all -- the better fix
than asserting an exemption.

The plan's evidence block quotes the gate's own output, which necessarily
contains the Vietnamese string it rejected. That IS the subject matter, so it
takes the documented region pragma. Mine was placed after the block instead of
around it, which exempts nothing.

Self-test still 12/12; the real bundles still clean at 150,365 strings.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…lan closed

The gap-closure plan's goal required PR #246 to be ready to merge, and #246 had
already merged when that plan was written. A condition that can never be
satisfied keeps a board alive forever while the bar it serves sits elsewhere.

So the PO-approved ship criteria become the single driving document: twelve
criteria, an eleven-row board where every row names the criterion it moves, and a
cycle log. Ticking a row does not move the bar; only a cycle's AC impact does.

Cycle 2 recorded. Neither of its two defects was what its message said. The
chat-service failures reported "corpus looks truncated" when the corpus was
ABSENT -- the directory is tracked for two notes while the recorded runs are not,
so a directory-existence guard said present and the loader found nothing. And DB
live-smoke was never a failing test: exit 125 is docker refusing to pull an image
that is no longer anonymously pullable. Both are NV-3 in shape, a check whose
scope never reaches its subject.

chat-service is now 3966 passed, 9 skipped, zero failed, up from 3960 passed with
2 failures. AC-5 moves to partial rather than met, because "no test is red on
main" is a claim about CI and CI has not run these fixes. AC-4 and AC-7 stay put
for the same reason: a fix nobody has watched run is not a green leg.

Eight rows from the closed plan finished with evidence; the rest are re-stated
against the criteria they actually serve.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…e one (T1)

AC-11 was unknown because nobody had ever attempted the release. Now attempted,
and the answer is concrete.

oss-release.yml fires on a pushed vX.Y.Z tag and does four things in order:
validate the changelog, compute a matrix, build 33 images and push them to ghcr,
publish a GitHub Release. Pushing the tag IS the ship act -- there is no separate
confirmation step -- so the rehearsal stopped short of tagging and every image
built locally was deleted afterwards.

It fails at the first gate in about ten seconds: no [0.1.0] section in
CHANGELOG.md. The gate is correct and the workflow is deliberately ordered to
fail before building anything. The harder half is that there is nothing to move
into that section: [Unreleased] has its five headings and zero entries, while 100
commits have landed on main since the release candidate, 36 of them feat or fix.
So the file this repo calls the source of truth for what shipped records none of
the last week. Issue #251.

The build half does work. 33 targets resolve, and the two gateways whose
lockfiles cycle 1 rewrote both build -- chosen deliberately, because a lockfile
diff and a passing unit suite do not prove an image. One boots far enough to load
every module and then fail fast on an absent environment variable, which is
correct behaviour rather than a broken build.

AC-11 moves from unknown to partial. The publish half is now blocked rather than
unknown, and unblocking it is the PO's writing, not an agent's.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The row said run against lw-iso:25174 and that was wrong twice, which is what
rule 6 is for. lw-iso is not running -- nothing answers on that port -- and the
skill's own TOOLS.md says the documented account does not log in there at all,
because LOCAL_TEST_ENV describes infra on 5174. Pointing a journey at the wrong
stack returns AUTH_INVALID_CREDENTIALS, which reads like a broken login.

The running infra stack was two days old, so it predated every change in this
release and would have tested the previous commit while looking exactly like a
passing regression. Rebuilt it, and verified the IMAGE rather than the build log:
the string only this change contains is present, and the old name is absent. The
rebuilt frontend source is byte-identical to the release commit.

Both persona journeys pass. The frequent one is not vacuous: ensureScale fails
closed with an explicit assertion, because below 21 books a client-side filter
over the first page is indistinguishable from a working search, so the defect the
journey exists to catch cannot appear.

And the rename renders. In English the heading and nav both read Campaigns; in
Vietnamese both read the same translated term, and the old name appears nowhere
in the rendered page in either language. Screenshots kept; the verification spec
was temporary and is deleted, since campaign-naming-gate is the standing guard.

AC-9 moves to partial rather than met. The suite is two journeys -- find a book
by name, and an empty library offering a next action. That is not the core author
journeys; planning and drafting a book is AC-10 and remains unrun.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The frontend suite was 5 failed / 6508 passed on main, and it also reported an
unhandled rejection that vitest warns "might cause false positive tests". Both
are the same defect: a vi.mock factory replaces its module wholesale, so a key
the factory omits is not "unused", it is undefined.

DefaultModelsCard: the component renders a composer row and imports
COMPOSER_CAPABILITY, which the mock never returned. That took the whole file
down, which is why all five tests failed rather than one. Underneath it sat a
second, quieter problem -- the component renders four rows and the test still
asserted three. Completed the constant from the real value rather than inventing
one, and corrected the count. Composer appends, so the two tests that address
rows by index still point where they did; checked against the component rather
than assumed, because an inserted row would have moved them silently.

KgNoProjectState: the sonner mock returned success but not error, while
ProjectFormModal's failure path calls toast.error. So the suite was green while
that component could not report a save failure at all -- the error path was not
merely untested, it was broken, and only an unhandled rejection revealed it.

Fourth occurrence of this shape here: react-i18next missing initReactI18next,
sonner missing error twice, and a settings mock missing a capability constant.

Frontend now 6513 passed, zero failed. Both fixes bitten: removing the constant
reproduces the import error, and reverting the count reproduces the length
assertion. Restored byte-exact between bites. Issue #253.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…t to demand

CI's all-gates leg was red on PR 252 because phase0-reconcile-gate refused six
entries. Three of them were mine: the two ship-readiness plans and the
remediation plan carried no Reconciles line at all, so nothing said which
existing rules they apply rather than replace.

That matters more for these plans than most. They are sweeps: the ship criteria
apply the acceptance-criteria and remediation-cycle shapes to one release, the
remediation board applies OUT-5, IN-4, SET-1..8 and DOCK-7 to defects a human run
found, and none of them invents a rule. Saying so is the whole point of the
question.

The other three were pre-existing and red on main too. The human-sim report had a
Reconciles line that described the REQUEST rather than naming index rows, which
the gate reads as phantom registration -- and it was right. Its original prose is
kept verbatim, moved under a heading that says what it actually is. The run log
and the go-live breakdown had no line at all.

Gate now green: 68 specs checked against 138 standards rows, each naming its
prior art. Its own self-test confirms it is non-vacuous in both directions --
it flags a missing line, a bare "none" and a phantom row, and does not flag a
real citation or a reasoned one.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…to users (T6)

Of the frontend's outstanding vulnerabilities, react-router was the only one a
user can reach: an open redirect via backslash in Link and useNavigate, a bypass
of CVE-2025-68470, and arbitrary constructor injection in deserializeErrors
during SSR hydration. The critical and high in that package set are vitest and
vite -- developer tooling, exploitable against a workstation when a test UI or
dev server listens, not against anyone using the app. So this one goes first and
alone, as the row asked.

Checked the API surface before assuming a major was safe: the app uses the
classic v6 component API throughout -- Routes, Route, Link, NavLink, Outlet,
Navigate, useNavigate, useParams, useLocation, useSearchParams -- and no data
router. That is the compatible path to v7.

Proved three ways rather than one, because a green install proves nothing about a
major: tsc clean, the full frontend suite at 867 files and 6513 tests with zero
failures, and both persona journeys passing in a real browser against a
rebuilt image. The journeys are the ones that matter here, since they navigate.

Frontend moderates drop 5 to 3. The remaining critical and high are the vite and
vitest chain, untouched here on purpose: they need their own major and they do
not ship. Issue #249.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
… npm ci (#254)

Found while upgrading react-router: frontend/package-lock.json was gitignored,
and the Dockerfile copied package.json alone and ran npm install. All 74 frontend
dependencies use a floating range, so the image resolved its entire tree fresh at
build time.

That does not merely make builds untidy, it undoes the fix. Two builds of the
same commit could produce different trees, so "we upgraded react-router" was a
statement about one machine at one moment. A vulnerable transitive dependency
could re-enter on the next build with nothing in the diff -- the same shape as
minio/minio:latest in this very release, where an unpinned reference turned a
green pipeline red with no commit to blame. And --no-audit disabled the audit at
exactly the moment the tree was decided.

The lock was suppressed in THREE places: .gitignore, frontend/.gitignore, and
frontend/.dockerignore. The third only surfaced when the build failed after the
first two were fixed, because COPY could not find a file the context excluded.

Every other lockfile here is tracked, and .gitignore already carried a
commented-out line for the api-gateway-bff lock reading "needed for Docker
build". The frontend was simply left behind.

Proved both directions. It builds: npm ci added 827 packages. And it fails closed
on drift, which is the property worth having -- making package.json demand
react-router 6 while the lock holds 7 gives EUSAGE, "lock file's
react-router-dom@7.18.3 does not satisfy react-router-dom@6.30.6". Restored and
rebuilt green.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
… T12)

Pushed the branch and opened PR #252 so the earlier fixes would be RUN rather
than assumed, then read every leg. Two things the PR could not answer, and saying
so is the point: DB live-smoke is SKIPPED on it, because the heavy gates are
scope-guarded to main-targeted PRs and this one targets the release branch; and
the dep-vuln workflow does not trigger on this branch at all, so govulncheck had
to be dispatched manually to be seen.

Dispatched, it printed "govulncheck: scanned 76 module(s)" and found nothing --
the guard from cycle 2 doing its job, on a scanner that had not run for months.
pip-audit turned green too, which is the datasets bump landing in CI. AC-7 is
therefore met: all four scanners run and say what they covered. Whether they find
things is AC-6's question.

The suites found two more partial mocks. The frontend was 5 failed plus an
unhandled rejection that vitest warns can cause false positives; one of those
mocks meant a component could not report a save failure at all. Now 6513 passed,
zero failed, zero errors, alongside chat-service 3966, composition 4194 and five
Go packages.

react-router 6 to 7 is done and proved three ways, because a green install proves
nothing about a major: tsc, the full suite, and both persona journeys in a real
browser against a rebuilt image.

That upgrade then exposed something larger. The frontend lockfile was gitignored
and the image ran npm install on 74 floating ranges, so nothing was ever actually
pinned. It was suppressed in three places, the third only surfacing when the
build failed after the first two were fixed. Now tracked and built with npm ci,
which fails closed on drift.

AC-4 moves to partial, not met: the MinIO fix remains unobserved because its leg
only runs on main-targeted PRs. A skipped leg is not a passed leg.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
AC-12 was unknown because nobody had asked. No rollback path existed anywhere --
not in versioning-and-releases, not in the runbooks, not in the go-live
checklist.

The image mechanism turns out to be the easy half. A release pushes 33 images to
ghcr and deletes nothing, so rolling back is redeploying the previous tag.

The hard half is the schema, and it is worse than expected. The novel-platform
services do not use versioned migrations: six of them run an idempotent CREATE
TABLE IF NOT EXISTS blob at startup, so there is no down path at all and the
schema only moves forward. That blob is cumulative and contains real non-additive
statements -- columns renamed from usd to tokens, language renamed to
original_language, tables dropped. An older image meeting a renamed column fails,
and no image rollback fixes it. The foundation track is the opposite and better
off, with 73 up and 73 down migrations in exact parity.

So the runbook makes rollback safety a command rather than a judgement: diff the
schema-touching paths between the two tags. For v0.1.0 that returns nothing --
zero schema files changed since the release candidate -- so this release is
rollback-safe by test.

Exercised rather than described, on two genuinely different images confirmed by
bundle digest rather than by tag: deploy, roll back, roll forward, all three
serving. One gotcha found by hitting it -- outside its compose network the
container dies instantly on an unresolvable upstream, which looks exactly like a
broken image and is not one.

AC-12 moves to partial, not met: one service of 33, no data written across the
boundary, and no production target to rehearse against.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The frontend's last vulnerabilities were the vite and vitest chain -- dev
tooling, reachable against a developer's machine when a test UI or dev server
listens, not against a user. Taken after react-router because that one shipped
and these do not.

vite 5 to 8, vitest 2 to 4.1.11, and @vitejs/plugin-react 4 to 6.1.1, which the
first two require: without it npm reports a valid install and an invalid tree,
since the plugin accepts vite 4 through 7.

Frontend npm audit is now 0 critical, 0 high, 0 moderate, 0 low. It was 1
critical, 1 high, 5 moderate this morning.

vite 8 is not a drop-in. It bundles with rolldown, which accepts manualChunks
only as a function -- the object form fails the build with "manualChunks is not a
function", after a softer "Expected Function but received Object" warning that is
easy to scroll past. Converted to the function form, matching on the package
boundary rather than a bare substring so react cannot claim react-router-dom, and
letting the longest match win. The grouping is unchanged and the build proves it:
vendor-react and vendor-tiptap still emit under their own names.

Proved the way a major has to be: tsc clean, 867 files and 6513 tests passing, a
real image build, and both persona journeys green in a browser against it.

One thing worth knowing for CI: vitest 4 removed the "basic" reporter. Nothing in
this repo used it -- checked across json, yml, sh and ts -- but a stale flag would
fail at startup rather than in a test.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
… gateways name their blocker

Two major chains, taken separately because they fail differently.

The frontend is now clean at every severity: 0 critical, 0 high, 0 moderate, 0
low, from 1 critical and 1 high this morning. vite 5 to 8, vitest 2 to 4, and the
react plugin 4 to 6, which the other two require.

vite 8 is not a drop-in. It bundles with rolldown, which accepts manualChunks
only as a function; the object form dies outright, after a softer warning that is
easy to scroll past. Converted, matching on the package boundary so react cannot
claim react-router-dom. The chunking is unchanged and the build says so rather
than me -- the vendor chunks still emit under their own names.

The gateways stop somewhere specific, and that is recorded on #247 so the next
attempt does not rediscover it. npm audit fix --force bumps four @nestjs packages
to 12 and leaves common at 10, so the tree will not install. Aligning common
fixes resolution and the build passes, and then seven of fourteen test suites
fail on ESM syntax: NestJS 12 pulls ESM-only dependencies into a service whose
Jest setup is CommonJS. That is a Jest migration across four gateways, not a
version bump.

Reverted byte-exact -- 202 tests passing again -- because a half-migrated gateway
is worse than the advisories it fixes.

One thing worth carrying forward: npm's own advice was wrong. After the bump it
offered the remaining highs a fix at 7.5.5, a downgrade from 12. Following the
tool blindly takes the framework backwards.

AC-6 stays partial. Every user-facing advisory in this release is now closed, but
rsa still has no fix and needs a named person.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Re-read my own T5 evidence instead of trusting the tick. It claimed the full
backend and frontend suites and had run four: chat-service, composition-service,
book-service and the frontend. The repo has ten Python services with tests,
including knowledge-service, the largest at 459 test files, which had never been
run at all.

So AC-5 -- no test is red -- was resting on less than half its subject. That is
the same shape as every other finding in this plan: a claim measured against a
fraction of what it covers, reported as though it covered all of it. The
difference is that this time the fraction was mine.

Ran the missing seven. Every one is clean: knowledge 5004, lore-enrichment 1275,
translation 1209, worker-ai 511, learning 201, campaign 184, jobs 142, video-gen
64. With the four from earlier that is roughly 23,000 tests across eleven suites
and zero failures.

Re-verified the five gates behind the met criteria rather than assuming they
still held, including their self-tests. PR 252 now has zero failures after the
all-gates cause was fixed.

Named rather than implied: the Go modules beyond book-service are not covered
here. This branch changes zero Go files and CI's Go leg is scope-guarded to
main-targeted PRs, so it is skipped. An unobserved leg is not a passing one.

AC-5 stays partial, but the word means something different now. It was green on
four suites; it is green on eleven. It is unmet for one reason only: the
criterion is about main, and these fixes are not merged.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…t say why they failed

Cycle 9. The PO asked for a full infra rebuild before their final test. It
failed -- zero images built -- while the background task reported exit code 0.
`docker compose build` does not do that: a failing target exits 1, reproduced
on a throwaway project both with and without bake.

pgvector (#255, T13): `infra/postgres-pgvector.Dockerfile` pinned clang19 and
llvm19 by hand. `postgres:18-alpine` is a FLOATING tag and moved to Alpine 3.24,
which ships neither. The pin was also already wrong -- the image's Postgres is
configured with LLVM_CONFIG=/usr/lib/llvm21/bin/llvm-config, so pgvector's JIT
bitcode was being built against a different major than the server loading it.
Now derived from pg_config, failing loudly if it cannot be read. Proven live:
CREATE EXTENSION vector -> 0.8.1 in the running container.

My own premise, corrected on #255 rather than edited away: this is NOT
release-blocking. release-targets.py returns 33 services and postgres is not
among them -- the 33 published images are exactly the compose services minus
postgres. It blocks local and CI stack bring-up, which blocks the journeys and
the authoring run, but not the publish.

The other 33 failed on TRUNCATED downloads, not bad packages: RemoteDisconnected,
IncompleteRead, a half-read JSON index, and a wheel failing its hash at
32.8/818.2 kB -- the hash check working. 34 concurrent builds saturated the link
and the 5-attempt retry burned all five while it stayed saturated. Rebuilt in
batches of four: all 33 built, zero failures, 42 containers up and healthy.

Two gates (T14). all-gates reported a failing gate by quoting the line
"[emit-0013] SELFTEST PASS ... (non-vacuous)" -- a success banner offered as the
cause of a failure. gate-wiring-gate printed the FIRST non-empty line of a failed
gate's output, and this repo's gates print a self-test banner BEFORE their real
work, so the quoted line was guaranteed to be the wrong one. It now quotes the
tail. Underneath it, emit-migration-0013-lint.sh had a read guard that was dead
code: `text=$(cat "$f"); rc=$?` under `set -euo pipefail` dies at the assignment,
so rc was never read and its FAIL message could never print. Bitten both ways --
the committed version prints the banner and nothing else; the fixed one names the
file, and skips a file that genuinely vanished mid-sweep instead of going red on
another gate's normal operation.

AC-11 and AC-4 stay partial; AC-11's partial now rests on all 33 targets building
in one pass rather than on representative images. Publishing is still blocked by
the changelog (#251).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…e now passes (T15, #251)

Cycle 10. #251 was the only thing the release act had ever reached, and it
failed there in about ten seconds. The section is written from the 39
user-facing commits in v0.1.0-rc.1..HEAD rather than summarised from memory.

Verified in the exact mode oss-release.yml invokes -- the workflow derives
VERSION="${TAG#v}", so a v0.1.0 tag calls `changelog-gate.py --release 0.1.0`:

  changelog-gate: structure + release 0.1.0 OK (3 section(s)).   EXIT=0

Bitten both directions against copies via the gate's own --path, so CHANGELOG.md
was never mutated to test it: with the section removed the gate reports "no
`## [0.1.0] - YYYY-MM-DD` section found" (exit 1); with the section present but
carrying only its five Keep-a-Changelog headings it reports "exists but has no
entries under it" (exit 1). The second is the case worth checking, because a
bare heading is what a hurried release adds.

A premise in #251 did not hold and the section uses the measured numbers: the
issue records 100 commits / 36 feat-fix on main; today main has 65 (26), and the
range that will actually be tagged -- main plus this branch -- has 89 (39).

The section records what is knowingly NOT fixed next to what is: rsa
RUSTSEC-2023-0071 has no published fix (#250), four gateways stay on NestJS 10
because 12 needs a Jest/ESM migration (#247), and 17 of 18 locales are
machine-translated and unread by a native speaker. A changelog listing only wins
is the same class of artifact as a README claiming unbuilt features.

Also corrects a stale fact in the file being validated: oss-release.yml's header
said it builds 41 images. release-targets.py returns 33; 41 is the compose
service count, which includes seven pulled images and postgres.

AC-11 stays partial, but the remaining distance is now a decision rather than a
task: every rehearsable step of the release act has been rehearsed, and pushing
the tag IS the ship act, which is the PO's.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
letuhao and others added 3 commits September 13, 2026 09:48
…a skip tallied as green

The local --run-all Cycle 9 left unfinished has finished. It proved the fix in
situ: a synthetic gate shaped like the failure mode (self-test banner first,
real failure after) is now reported by the line that explains it rather than by
its banner, and the gate Cycle 9 actually fixed came back GREEN under the
concurrency that broke it in CI.

Two of the three other reds were collateral from my own synthetic file -- it has
no self-test, so gate-teeth-gate failed listing it, and gate-number-visibility
then failed because a failing gate-teeth-gate does not print CI_GATE_FLOOR.
Removing it returns both to green, verified rather than assumed.

The third is real and is filed as #256: a gate that reads its universe from a
live Postgres prints "this is a skip, not a green" and then returns 0, so
all-gates recorded it GREEN (0.2s) in CI while it could see nothing. With a
stack up it is RED, carrying 2 ACCOUNT-scoped findings -- one on `motif`, the
exact table its own docstring was written about.

T16 stays OPEN. Both obvious repairs were tried and rejected rather than
shipped: NEEDS_STACK would convert a true red on a stacked box into a permanent
skip (live_plan skips every row not in LIVE_BARE, and the anchor probes lw-iso
while this gate needs infra's Postgres), and exit 2 would make CI permanently
red. The repair that fits is a SKIP status in the shared runner contract.

A premise of mine, wrong, and caught by testing rather than reasoning: I first
diagnosed the gate as having no guard at all and wrote one. Running the
committed version against a dead docker daemon showed the guard is already
there -- I had stopped reading main() two lines short. Reverted byte-exact; the
tree carries none of it. Fourth premise this plan has corrected that was its own.

AC-4 stays partial, and now for a named reason rather than an absence: a GREEN
leg in this repo can mean "skipped" and be indistinguishable from "checked",
which weakens what a green all-gates is worth until the runner can say SKIP.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…ading as a pass (T16, #256)

Cycle 12. Cycle 11 parked this as "a protocol change the PO should see coming".
Re-read against the STOP list that was over-caution, not a blocker: a local gate
runner publishes nothing, decides nothing, and touches no shared target. Parking
work nobody has to decide on is how a board grows rows that are nobody's.

The contract was binary -- 0 passes, anything else fails -- so a gate whose
subject is absent had no way to say so except `return 0` and explain in prose.
test_seed_assert_applies_the_lifecycle_predicate_gate.py did exactly that, and
CI recorded it GREEN (0.2s) while it could see nothing. The prose had been right
since it was written; only the exit code was lying, and the exit code is the
half a 150-row sweep is read by.

GATE_SKIP_RC = 3. _run returns the code as well as the boolean, the verdict
moves into a pure classify(), and the gate returns 3 where it returned 0.

Proved by exit code, with Postgres made unreachable by pointing DOCKER_HOST at a
dead daemon rather than by stopping the running stack: same message, OLD_EXIT=0
-> NEW_EXIT=3. The working path is unchanged and checked, not assumed: with the
stack up the gate is still RED with 42 findings.

Bitten by disabling classify()'s SKIP arm -- precisely "the runner was never
taught exit 3". The self-test goes red naming both consequences, and restoring
is byte-exact (cmp clean). Six assertions, each a way the fix could be undone
silently: exit 3 is SKIP; exit 3 on a KNOWN_RED row is STILL SKIP, or a deferral
gets renewed by a gate that never ran; 0 is still GREEN; 1 is still RED, so the
SKIP arm cannot swallow real failures; a TIMEOUT is RED and never SKIP, because
a killed gate has not told us its subject was absent; and 3 must not collide
with pass/finding/misuse.

classify() is pure and split from the printing for the same reason live_plan is,
and that reason IS this bug: a branch that only runs during a fifteen-minute
--run-all is a branch nobody checks.

T17 opened for what the false green was hiding and this does NOT fix: two
ACCOUNT-scoped seed assertions blind to their lifecycle column, one of them on
`motif` -- the exact table the gate's own docstring was written about.

AC-4 stays partial for the reason it has since Cycle 5 (the MinIO leg is
scope-guarded and unobserved), but a green leg is worth what it says again.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…swers (T17, #256)

Cycle 13. What T16's exit code made visible. The gate dedupes by (table, tool),
so "2 ACCOUNT-scoped" was 2 pairs across 3 instances; they were enumerated by
re-applying the gate's own predicate rather than inferred from the summary.

The gate's docstring says it cannot tell you whether the tool's read predicate
really is status='active', so that was read from the service:

  motif_repo.get_by_codes:  WHERE code = ANY($2) AND status = 'active' AND <visible>
  arc_template_repo:        list_for_caller(..., status: str | None = "active", ...)

And the hazard is present rather than theoretical -- this database holds 27
archived arc_template rows and 162 archived motif rows.

Narrower than the generic case, and worth saying rather than overstating: every
code carries {run_id}, so a stale row from a PREVIOUS run cannot collide. The
2026-08-23 defect this family is named after had no such discriminator.

Fix: AND status='active' on all three, with the justification written into each
scenario's own `why` -- where the gate's docstring says per-repository read
predicates belong. Narrowed to what the tool sees, NOT relaxed. Edited through a
JSON round-trip proved byte-identical on unmodified input first, so the diff is
exactly six lines and nothing is reformatted.

  BEFORE  ACCOUNT-scoped: 2   EXIT=1
  AFTER   ACCOUNT-scoped: 0   EXIT=0

Bitten: removing the predicate again flags motif/composition_motif_bind_edit by
name with the right count and exits 1; restoring is byte-exact (cmp clean) and
returns to 0.

T17 stays PARTIAL, deliberately. The scenarios were NOT re-run -- they are live
MCP probes needing a stack, a book, a Work and a minted token, which is why
gate-wiring-gate classifies that family NEEDS_STACK. The assertions now ASK the
right question; that they still PASS when asked is owed. Narrowing can only make
them stricter, and a fixture failing on an archived row would be correct -- but
that is reasoning, not a run.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
letuhao and others added 30 commits September 14, 2026 07:55
Once `continue` returned real prose, studio-inline-correction failed on Accept:
"Element is outside of the viewport". InlineGhost is position-fixed at the caret
with only a width limit, so a ~300-word suggestion ran off the bottom of the
screen. Fixed content does not scroll with the page, scrolling re-anchors it to
the caret, Esc discards and nothing accepts from the keyboard -- a full-length
suggestion could not be accepted at all. Every earlier suggestion was a short
request for context, so it always fit.

The card is bounded to the viewport; the prose scrolls inside it and the actions
stay on screen. Width and anchoring unchanged.

Bitten through a rebuilt frontend: HEAD InlineGhost reproduces the viewport
error. Restored, green.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The recurring "page.goto: Timeout 15000ms exceeded" -- three full runs, blamed on
host starvation -- was a render-blocking Google Fonts stylesheet.

A host sampler caught the latest occurrence at 26-45% CPU, 17 GB free, no model
running, which refutes the starvation hypothesis. The trace network named the
request that never completed: fonts.googleapis.com/css2, status -1, with every
same-origin resource done in under 200ms. index.html loaded it as a normal
<link rel="stylesheet">, and a stylesheet blocks the load event.

That is a product defect: LoreWeave is self-hostable and ships zh-CN, and on a
network that cannot reach Google every page would hang behind a cosmetic font.

The link moves out of index.html; loadWebFonts inserts the same stylesheet after
`load`, so it can no longer hold it. display=swap and the existing fallback stacks
mean text renders at once and swaps when the font arrives.

offline-font-cdn.spec.ts routes the CDN to never answer, turning the intermittent
timeout into a deterministic one. Bitten: HEAD index.html fails it with the same
timeout; restored, it passes, and fonts still load when the CDN is reachable.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…0 skipped (Z2)

Final run on rebuilt images: 201 passed, 0 failed, 0 skipped, read from the Allure
summary by content; the evidence gate passes on all 202 test directories.

The handover says plainly what one green run does not prove: plan-forge analyze
still truncates on 6.5% of real runs (a different mode, not yet measured for a
fix), the SPEC string bounds ship as a guardrail at p~0.27, and the inline
critique timeout fix could not be re-broken without controlling LM Studio by hand.

Owed to the PO: AC-7, H1 (OQ-1), and whether one screen should show the critic
twice. Nothing pushed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Each was partial only because its last test moved to another row: F1's fourth test to F6,
F3's second to F7, and J3's blank distiller to F9. All three of those rows are fixed and re-broken,
and every test is green in the final run. J3 also carried a wrong verdict ('H2's constraint, not a
defect'); the correction is written into the row rather than dropped.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
/ideate brainstorms, researches prior art and scores ideas into docs/ideas/, and promotes a chosen
idea into a CLARIFY spec draft. AGENTS.md gains a Pipeline table mapping every phase to the command
that drives it; CLAUDE.md points at it. IDEA-001 and its spec draft (close the v0.1.0 leftovers)
are the first use.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…-T3)

A plan-forge loop that runs to the token cap comes back with finish_reason=length. chat() raised
before the regeneration ladder existed for it, so the retry built for loops was unreachable for the
failure that remained (2 of 31 real runs). Truncation now has its own error type carrying no text,
the first call lives inside the ladder, and truncation regenerates - never repairs.

Also: the schema-rejected fallback dropped the escalated frequency_penalty; and a running worker job
now heartbeats updated_at through cancel_check, so the 900s sweeper stops starting a slow job twice.

Each fix re-broken and restored; evidence in docs/plans/2026-09-18-close-v010-leftovers.md Cycle 1.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…backfill on sign-in (T6-T10)

A book created through REST now asks composition for its Work right after commit, off the request
path, with the author's own bearer - the same call the Studio makes on open, made earlier. OQ-1
holds: nothing is minted and no other identity is sent. MCP book_create is unchanged.

POST /v1/books/provision-missing provisions the caller's own active books (never a diary, never
anyone else's) that lack a ready Work: a GET per book, a POST only when needed, concurrency 4,
90s deadline, detached from the client that fires it.

Also fixes an order-dependent migrate test that counted every scene in the shared test DB.
Each change re-broken and restored; live race and create checks in the plan's Cycle 2.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…T11, T13)

setTokens runs on sign-in only (login, register), never on a silent refresh or a reload, so it now
fires POST /v1/books/provision-missing once with the new bearer: keepalive, fire-and-forget, and
opted out of the global operation tracker so background housekeeping never lights the progress bar.

When the critic panel is on screen (floated, popped out, or the active docked tab), Compose no
longer repeats the verdict inline. It keeps the C26 override gate, whose Regenerate action exists
only there.

Each change re-broken and restored; evidence in the plan's Cycle 3.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…y chapter editor (T5, T12, T16)

composition POST /work ran the pending-Work backfill only when it had just created a knowledge
project. A book with an existing unmarked project AND a pending Work fell through to a second Work
insert, hit the one-Work-per-book index and answered 409 WORK_CREATE_CONFLICT on every attempt,
the Studio's own open included. Found by the live sign-in backfill: 9 of one owner's 105 books.

The legacy chapter editor is retired (PO): its route now only redirects the same chapter into the
Writing Studio, every in-app link goes there through studioChapterPath, and the page is deleted.

Also records the live measurement (30/30 plan runs, one truncation regenerated) and the critic
ceiling re-broken through a slow stand-in provider. Evidence in the plan's Cycle 4.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…T14)

[0.1.0] now says what the gap-closure branch changed (provisioning at creation and on sign-in,
the Studio as the only writing surface), what it fixed, that the legacy chapter editor was
removed, and what is still open: MCP-created books provision later, a plan run can fail if the
model loops three times, and the critic waits at most 240s. Evidence in the plan's Cycle 5.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
remediation-cycle-gate (RC-1/RC-5) reads **Investigated:**, **Issues:**, **Fix:**, **Proof:** and
**AC impact:** labels; the cycles used a period and had no Issues line, so the gate went red and
could not trace any met criterion to its cycle. Labels normalised, content unchanged.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…T17)

Eight specs reached the legacy editor through ChapterComposePanel; it now drives the Studio (deep
link, palette-opened scene-compose, editor tab) with the same property names. Where the Studio
changed a precondition the claim was re-expressed, never loosened: the co-writer is ready without a
setup step (U1), one scene not two (U2), the what-if goes through the canon picker now that both
books have a Work. B7.3 (no Work, so Publish ungated) is unreachable in the Studio and stays pinned
by usePublishGate.test.tsx; flagged for the PO. Decisions and runs in the plan's Cycle 6.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…plained (T15)

Run 3: 201 passed, 0 failed, 0 skipped. Runs 1-2 had 5 failures: one real (T7 made both world
books canon, fixed in Cycle 6) and four that pass alone and in run 3 with no confirmed cause.
AC-10 stays partial: green once is not reliably green. Evidence in Cycle 7.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…lock

Go's platformjwt.Verify never validates iat; PyJWT does by default with zero leeway, so a few
seconds of clock difference between auth-service and a Python service 401'd a live, unexpired
token that Go services accepted. The dev Docker VM's wall clock steps back 1.4s every ~30s,
which is how studio-publish failed with a token its own setup had just used. verify_iat off,
exp and signature unchanged. Live probe: 14 x 401 before, 0 after, over the same 70s window.

Also records the revision-order-by-wall-clock fragility as DEFERRED #164. Evidence in Cycle 8.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…e worker takes over

sw.js claims clients on activate, so a visitor's FIRST install fires controllerchange on a page
that had no controller; registerSW reloaded on every controllerchange, wiping what the person had
started typing about a second after load. It now reloads only when an already-controlled page
switches to an accepted update, as its own comment said. Live: 5/5 first visits reloaded before,
0/5 after. Found as an E2E sign-in that lost both fields (Cycle 9).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Run 6 on images carrying every fix: 201 passed, 0 failed, 0 skipped. Every red across six runs is
fixed, explained as an environment condition, or tracked: two single occurrences with no proven
cause are DEFERRED #165, not called flaky. The handover report gains the update for the PO; the
ship decision (AC-11) is still theirs.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
- #289: the outline-node test posted kind 'arc', which a996750 removed from NodeKind, so
  validation answered 422 before either claim was tested; it posts 'chapter' now.
- #290: the undirected-yield test ran git grep, absent from the service image; a Python walk
  over services/ makes the same assertion.
- #291: CI now runs book-service internal/migrate DB tests too, with -p 1 (shared DB).
- #292: frontend/tests/e2e gets its own tsconfig and typecheck:e2e (pre-commit wired); the ten
  type errors it found are fixed with real types.

Each fix re-broken and restored; IDEA-002, its spec and the plan are included.
Evidence: docs/plans/2026-09-19-evidence-runner-and-model-lease.md Cycle 1.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…er (T5-T7)

run-evidence-suite.py warns and records (never refuses): wall-clock steps, schedulers due in
the run, models LM Studio has loaded, and provenance of the image each container actually runs
(sha vs HEAD, uncommitted changes in its build scope, replaced by a rebuild). The run writes to
runs/<id>/ so a later re-run cannot erase a red's trace, and appends to runs/LEDGER.jsonl.
why-red.py assembles one red's trace, stack logs in its window, llm_jobs and clock steps.
iso.sh now stamps git sha, build time and a new git_dirty_scope label on every image it builds.

Bitten: evidence survives a plain re-run; labels go 'unknown' without the export.
Evidence: docs/plans/2026-09-19-evidence-runner-and-model-lease.md Cycle 2.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…ycle format

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
… T9-T10)

A per-credential, opt-in flag for a local server that can hold one model. It is the user's
statement about their own hardware, so it defaults to false and a patch that does not mention it
keeps what the user chose. Nothing uses it yet (T11-T12).

T9: run 4's circuit breaker was not opened by the load abort (a permanent 400, which Guard does not
count) but by at least five transient attempt failures inside jobs that nothing logs; T13 will
log each failed attempt and classify from the live replay.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…oint take turns (#286, T11)

For a local server that holds one model, two callers asking for different models at once abort
both loads. The lease (Redis + one atomic Lua script, in the governor's style) lets calls for the
held model share the endpoint and makes a different model wait; FIFO with an aging bound so no
side starves; crashed holders and silent waiters are pruned; fails open on a Redis error. It only
orders requests and never loads or unloads a model. Not wired yet (T12).

Tested through the real script with miniredis; three rules re-broken and restored.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
… job and stream paths (#286, T12)

For a credential whose owner turned on "serve one model at a time", each provider call takes the
endpoint's model lease: inside Guard per attempt on the job path (a retry re-acquires it), and
around adapter.Stream on /v1/llm/stream, which bypasses Guard. Without the opt-in both paths are
exactly as before. A lease timeout is its own code, LLM_MODEL_BUSY, kept out of the SDK retry list.
Tunables MODEL_LEASE_TTL_S / _WAIT_S / _AGING_S. Vision is not wired (recorded).

Both wirings re-broken and restored; evidence in the plan's Cycle 5.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…ever a breaker failure (#286, T13)

LM Studio answers two concurrent loads of different models with a 400 "Failed to load model ...
Engine protocol startup was aborted". It is now ErrUpstreamModelContention: retried (the same call
succeeds once the other model has loaded) and never counted toward the circuit breaker, whatever the
credential's setting. Every failed attempt is logged at WARN with its class.

Also: the evidence runner's preflight flags a container whose log is not being captured - the
reason run 4's failed attempts left no trace: provider-registry's stdout had not reached docker
logs since its restart.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
… the live collision replayed (#286, T14-T15)

T14: a checkbox in the add and edit provider dialogs. The edit dialog shows
the stored value and sends the field only when the user changed it; create
sends it only when ticked, so the UI never turns the setting on by itself.
English keys plus the 17 other locales filled by i18n_translate.py.
Vitest: 3 tests, bitten three ways.

T15: run 4's collision replayed live on lw-iso, 3 x gemma 12B + 3 x gemma
26B at once on one LM Studio. Setting off: 1/6 completed, 5 x
LLM_CIRCUIT_OPEN. Setting on: 6/6 completed, wall 25.2 s. The replay found
that LM Studio also answers a model swap with HTTP 500, which counts toward
the breaker even with the lease on (4 of 5); filed as #295.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…closed with evidence (#286, T16)

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…fuses an account that cannot log in (#288, T17 run 1)

Run 1 through the evidence runner: 194 passed, 7 red. Every red has a cause
(plan Cycle 10). Three were test defects:
- demo-pipeline-3b/3c read the book row from a list React had not yet
  replaced; create lands in the Studio since 88d3e97, and e86508e
  repaired only 3a.
- inline-correction looked for inline-discard while the ghost was still
  streaming (Discard is inline-stop then); real prose since afe2454.
- quality-conformance's trace.or(empty) matched the loading wrapper, so
  a blank panel passed, or matched two elements.
Each fix is bitten on the product side or by timing. The other four reds
were environmental (slow LM Studio, Chromium context-close hangs).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…unt; the runner waits for a restarting stack (#288, T17 run 2)

The library's first page is 84 system motifs plus the user's by name. 47
leaked test motifs pushed a freshly seeded one to row 101 and turned three
motif tests red. The specs now archive what they create (by id, or by code
for motifs made through the UI); bitten by removing one cleanup and watching
the active count grow. The enrichment red was LM Studio throughput plus a
cross-test backlog; its bound is not raised.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant