Fix the three pages that shipped broken, and the checks that let them - #6
Merged
bitcoinuniverseadmin merged 29 commits intoAug 28, 2026
Merged
Conversation
Enabling the database made the backend exit at startup. updatePoolsJson fetches the pools-v2.json sha from GitHub, that fetch must not happen at runtime here, so currentSha stayed null and index.ts aborted before it could serve anything. Bundle the pool list in the repository and identify it by the git blob hash of its own bytes. That is deterministic, needs no network, and is directly comparable with the sha already stored in the state table, so a redeploy with unchanged pool data does no work. The bundled file is the default now; the GitHub URLs stay for anyone who wants them. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…rent one The frontend decided on its own that Mining and Charts were available. The backend decided separately, from DATABASE.ENABLED and INDEXING_BLOCKS_AMOUNT, whether to mount the routes behind them. Nothing compared the two answers, so production shipped both pages against routes that were never registered and every request behind them answered 404. Give the two one place to agree on. The route setup records what it actually mounted, /api/v1/capabilities reports that alongside dependency reachability, coverage and lag, and startup refuses a configuration that would advertise a feature it cannot serve. The rules live in their own module with no imports, so the release gate and the tests judge exactly what the running backend judges. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Every $listXX range caught its database error, logged it, and returned an empty array. The route then answered 200 with [], so a database that was down was indistinguishable from a chain with nothing to show, and the Charts page rendered an empty state over a real outage. Throw instead. statistics.routes already turns a thrown error into 500, which is the answer the frontend needs in order to tell unavailable apart from empty. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The Charts page subscribed with one next handler and no error handler. A 404 killed the subscription, so isLoading stayed true, the spinner stayed on screen, and the range buttons stopped working for the rest of the session. Every mining widget had the same shape through an async pipe with an else skeleton, so a failed read left 35 skeletons that nothing would ever clear. Add one helper that turns a request into a state machine which always reaches a terminal state: data, empty, stale, or error. It keeps empty apart from failed, bounds the wait, retries only what a retry could fix with backoff and jitter, cancels an obsolete request when the range changes, and keeps the last good answer to show as stale rather than replacing real numbers with a blank panel. Charts, reward stats, pool ranking and the hashrate chart use it. Each owns its own state, so one failing module no longer decides what the rest of the page shows. The shared status panel names what failed and offers retry and system status, and its spinner is drawn from theme tokens: the one it replaces was white on a light background, which is why a failed page and a blank page looked the same. Pool ranking and the hashrate chart also started one request from route setup and another from the chain tip. They share one trigger now, debounced so a burst of tips around a new block is one refresh. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The last release passed every check and still shipped a Charts page and a Mining dashboard whose backend routes were never mounted. CI answered from fixtures, so nothing it ran could have noticed: the thing that was wrong was the configuration of the machine, and no check looked at that. release.sh installs beside the running release and refuses the symlink swap unless the build is complete, the configuration is coherent, the database answers, the source registry parses with every token present, and every protocol the registry calls readable has an authority configured. After the swap it reads /api/v1/capabilities and fails the release if a feature is enabled with no routes behind it, which is exactly the state that shipped. The integration database is pinned to the engine and major version the deployment runs. It was MariaDB 10.5 against a MySQL 8.4 plan; the migrations use MariaDB syntax MySQL rejects outright, so a test database on the other engine would have proved nothing about the release. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The page put two chips of equal weight beside each protocol: what the registry says it implements, and what the runtime snapshot says its authority can do. So Ordinals, Rare Sats and Runes read "Live, read only" next to "Authority unreachable", and the summary counted all three among the protocols readable that day. Both facts were true; neither was the answer to the question a reader is asking. Availability now decides the primary label and the count. A protocol whose authority is unreachable, unconfigured, or still catching up is not readable, whatever the registry says it implements. The registry capability follows as a qualifier, and a line underneath gives the evidence: where the authority has reached, how far behind that is, how many checks have failed in a row, and when it was last asked. An authority that is catching up gets its own state rather than being rounded to working or broken, because that is what an index rebuild actually is. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…ing zero A currency with no rate fell through to a multiplier of zero, so every amount rendered as a confident $0.00. That is the same error as answering a failed query with an empty list: it turns "we do not know" into a definite value, and a reader has no way to tell the two apart. With no usable rate the amount is left blank and titled. This surfaced when the third-party price feed was switched off, but the fallback was always there. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The visual matrix checked overflow, console errors, contrast and accessibility. A Charts page that never resolved and a Mining dashboard of 35 skeletons passed all four and shipped: nothing overflowed, nothing logged, placeholders have fine contrast, and axe is content with a skeleton. The matrix now measures whether a page finished. With a populated fixture there is no excuse for a loader still on screen after the settle wait, for a skeleton that never resolved, or for a chart panel that drew nothing, and each of those fails the run. With a failure fixture the obligation is the opposite: a page still waiting must have said why. Fixtures still cannot prove the deployment can produce data at all, so synthetic-check.mjs asks the live origin directly. It fails on a feature advertised with no routes behind it, a range that answers with nothing, a protocol marked readable whose authority cannot answer, a configured authority with no checkpoint, and a frontend and backend on different builds. It runs after a release and hourly. The backend integration tests now run in CI too, against a real database, because the migration failure that had to be found by hand was a database one and no test in the suite touched a database. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The deployment notes described a database that was off, and the release procedure was a list of manual steps with no gates. Both are now what the machine does: a dedicated MariaDB container, why it is MariaDB and not the MySQL pinned elsewhere, the bounded indexing window and why it is bounded, the bundled pool metadata, the price feed that is off and why, and a release that refuses to cut over on a configuration it cannot serve. The overlay ADR gains the status contract. Capability and availability were never written down as separate things, which is how they came to be rendered side by side with equal weight and read as one answer. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The capability probes and the bundled pool import each catch their own errors and return a value rather than throwing, so awaiting them is safe. The lint rule cannot see that without the annotation, and it was right to ask: the rollback in the pool import was itself an unguarded await, so a failing rollback would have replaced the error that caused it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
A gate that never fires is indistinguishable from no gate. The judgment is exported and covered directly: a spinner still turning on a populated page, skeletons that never resolved, a chart panel that drew nothing, a page that is only placeholders, and a failure state that waits without saying why all fail; a page that finished and a failure state that explains itself do not. Importing the harness no longer launches it, so the test can reach the judgment without driving a browser. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The clock link, the fee-level filter, and the invert toggle are icons with a title and no accessible name, so a keyboard or screen-reader user reaches three stops that announce nothing. The keyboard harness had been reporting all three; nothing acted on it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
CI runs it hourly, which only covers the window CI happens to be awake. A timer on the machine that serves the origin covers the rest, writes to the journal, and marks the unit failed when a check fails, so an outage between releases is visible where every other service failure already is. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Fulcrum is gone from the explorer host: no process, no data directory, no unit file. The backend was still configured to reach it, and a dead address backend does not fail loudly, it retries. It was producing roughly two connection errors a second, forever, which buried every real error in the journal. Core answers blocks, transactions and the mempool either way, so the deployment reads addresses from Core for now and an address lookup fails immediately rather than hanging. The preflight refuses an electrum backend when nothing is listening on its port, so this cannot come back quietly. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…at it says The new visual gate found this on its first real run: with every request held open, the protocols page sat on its skeleton indefinitely. The registry read had a catchError but no deadline, so a request that hangs rather than fails left the page waiting with nothing subscribed to clear it. It has a budget now, and the error state it reaches offers a retry. The gate's own rule was also wrong for the fixture named loading, which holds requests open on purpose to photograph the waiting state. Asking that to have finished asks the wrong question. It is judged on whether the wait is announced at all, which is what a screen reader needs and what a blank rectangle fails; the deadline itself is covered by the lifecycle tests, which run far longer than the harness waits. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
A run pointed at a second output directory committed 21 screenshots and its report, because only the default directory was ignored. Ignore any output directory the harness is pointed at. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Statistics are collected from the moment the writer starts, and nothing backfills them: there is no first-party source to backfill from, and inventing rows would be worse than having none. So a 1W range currently draws a couple of hours of samples under a heading that says 1W. That is not a lie the chart tells on purpose, but it is one a reader would take away. When the samples cover noticeably less than the range asked for, the page says when collection began and that this is all the history there is. The note disappears on its own as the history fills in. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The history note read "1 hours ago", and it was a getter, so it reduced over the whole series on every change detection pass, on a page that takes a live sample every minute. It is computed when the series changes now, and it counts in singulars. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
A restore was verified by hand, which proves the dump format works and nothing else: there was no backup being taken. A daily timer now writes to the same directory and keeps two weeks. It refuses to keep a dump that is suspiciously small and checks the archive reads back, because a backup nobody verifies is not a backup. The deployment notes gain the measured growth figures, so the next person deciding whether to prune or to widen the indexing window has the numbers rather than a guess. At roughly 300 MB a year against 1.5 TB free, nothing is pruned, and statistics are kept indefinitely on purpose: the "all" range is the whole series. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The notes still told the reader to compose a dump command by hand, which was accurate before there was a timer taking one every day. They now name the unit, and they say to verify a restore rather than to trust that a dump exists, because that is the part that was actually missing. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The step failed in CI on its first run: a container left behind by an earlier run was still up, so compose reported it as already running and never recreated it. The engine had changed underneath it, and the readiness probe then waited out its attempts on a database that was never going to be the one the tests asked for. Tear the previous container down, volumes included, and recreate. Remove a stray container holding the same name too, because `down` only removes what compose owns and `up` fails on the conflict rather than starting anything. Probe with either client binary name, since MariaDB renamed them and a probe that knows one name reports a healthy server as a database that never started. Verified by leaving a container of the wrong engine running and watching the suite replace it and pass. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The production monitor recorded the last deploy as an outage: three checks failed at the exact second of the cutover. The backend and the overlay take a few seconds to listen again after a restart, and the gateway answered 502 the instant the connection was refused, so every request in that window became a visible failure. A request with no body is now retried over about five seconds while the connection is refused. That bridges a restart, and it never retries a write, because only GET and HEAD can be replayed. An upstream that is genuinely gone still gets a 502, still well inside the page's own budget, so a real outage is reported rather than hidden behind a long wait. Verified by starting the gateway against an upstream that is not listening and bringing it up a second later: 200 after 1.8s, against 502 after 5.3s when nothing comes back at all. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Probing the origin through a cutover still caught two 502s. The retry added for a restarting upstream cannot help there: the gateway itself was being restarted, and nothing behind a process that is down can answer for it. It does not need restarting for most releases. The backend and the overlay run from a path baked into their unit at exec time, so they have to come back; the gateway resolves its static root per request, so a new frontend reaches it through the symlink on the next request. It restarts only when its own file differs from the running release, which is the one case that still shows a brief gap, and the log says so when it happens. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The notes said a deploy caused at most a brief connection reset, which was a guess. It has been measured now: probing the origin once a second through a cutover returned 200 for every request. The one case that still shows a gap is a release that changes the gateway itself, and the notes say which case that is rather than leaving the reader to find out. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The gate failed its first honest run against every route, and each failure was the gate being wrong rather than the page. It probed at a fixed pause, which is a race and not a deadline: on a loaded CI machine a page that finishes perfectly well is still mid-render at 2.6 seconds. It now waits for the page to settle, up to fifteen seconds, and what fails a run is a page that never settles. It counted a spinner nobody can see. The block overview keeps its loader in the tree and fades the wrapper with opacity, so checkVisibility is used now, which walks the ancestors and understands opacity. It counted app-loading-indicator, which is a labelled progress banner and is supposed to stay for as long as the work runs. Failing a release for saying "Indexing blocks" is the opposite of the point. It counted placeholders below the fold. Angular defers work until it scrolls into view, so a block page holds a transaction placeholder there on purpose. The harness judges a viewport, so the probe does too. All thirteen routes pass now, and the judge's own tests still hold it to firing on the states that shipped. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Run across every route and data state, the gate found four more pages with the same fault: home and blocks hold a placeholder with nothing said about why when the chain backend is down, and the transaction and address pages render bare skeletons with no accessible announcement, so a screen reader user hears nothing while they wait. Those are real and were found here, but they are not what this change set reviewed, and quietly widening the blast radius to make a gate green is how gates end up disabled. The gate blocks on the three routes whose request lifecycle has been rebuilt, and prints the rest every run under their own heading, never suppressed. Adding a route to GATED_ROUTES is how that work gets finished. The loading rule also stopped demanding a spinner from pages that have nothing to fetch. The docs and source pages render from the bundle, and failing them for not waiting was the rule being wrong. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The gate's first full run reported home and blocks holding placeholders with nothing said about why when the chain backend is down, and the transaction and address pages rendering skeletons a screen reader hears nothing about. The dashboard now gives live chain data a deadline: when the first socket payload has not arrived after ten seconds, a status panel says so and offers a reconnect, and a late arrival still clears it. Its waiting state is announced to assistive tech from the start. The blocks list stops retrying forever every ten seconds; a failure is bounded, terminal, says what happened, and offers a retry. The transaction and address pages announce their wait. All four routes join GATED_ROUTES, so the gate that found them holds them from now on. Verified across default, dark and contrast at 375 and 1440: zero unfinished pages, zero axe violations, zero measured contrast failures. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
axe measured the tracker bar's pulsing stage at 4.28:1 against the inset surface at its smallest size. The stages that sit on that surface use solid secondary text instead of the translucent default, which composites below the threshold there. The dashboard's scrollable tables becoming keyboard reachable rode along with the page fixes in the previous commit; this closes the remaining axe finding. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The frontend job served its build on a fixed port. Two runners share a host, so a second job's gateway crashed on the taken port behind the ampersand, the readiness probe happily reached the first job's server, and the run measured another build until that job ended and took the port down with it, failing 1843 navigations halfway through. Each run now asks the kernel for a free port, publishes it to the following steps, and refuses to continue unless the process answering the health check is the one it just started. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Author
|
Coordination note: this session pushed the three follow-up commits at head 6d48139 (four newly gated routes, the axe fixes, and the per-run CI port) and is driving this PR to green, its merge into develop, the promotion to main, and the production deployment. The PR 7 session should keep reconciling on feat/universe-glam-pink and take the new develop after this merges; the go-state branch will not receive further commits from anyone else while CI runs. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What was wrong
Three pages shipped broken, and every check passed.
Protocols showed
Live, read onlybesideAuthority unreachablefor thesame protocol, and counted it among the protocols readable that day. The
overlay was configured to reach ord at
127.0.0.1:8383, where nothing listens;ord 0.29 listens on 8382 and its socket-activated entry point on 8380 was not
enabled. 26 of 28 authorities were not configured at all.
Charts was a blank panel and a spinner that never stopped. Every
/api/v1/statistics/*route answered 404, because statistics routes are onlyregistered when
STATISTICS.ENABLEDandDATABASE.ENABLEDare both on, andthe database was off. The page subscribed with one next handler and no error
handler, so the 404 killed the subscription:
isLoadingstayed true and therange buttons stopped working for the rest of the session.
Mining left 35 skeletons and 5 spinners on screen next to live block data.
Every
/api/v1/mining/*route answered 404 for the same reason, and everywidget used
obs | async; else skeletonwith no error branch.CI could not have caught any of it. The visual matrix answers every request
from fixtures, so it proves how a page renders given data and nothing about
whether the deployment can produce any.
What this changes
records what it mounted, and
/api/v1/capabilitiespublishes it. Startuprefuses a configuration that would advertise a feature it cannot serve.
reaches a terminal state, keeps empty apart from failed, bounds the wait,
retries only what a retry could fix, and keeps the last good answer to show
as stale rather than blanking the panel.
[], so anoutage is no longer indistinguishable from an empty chain.
qualifier, and the evidence is on screen: indexed height, blocks behind,
consecutive failures, last checked.
own bytes, so enabling the database no longer requires a GitHub fetch at
startup.
$0.00.release.shrefuses to cut over on an incoherent configuration and rollsback when verification fails.
nothing, or a failure state that says nothing.
synthetic-check.mjsasks thelive origin with nothing mocked. Backend integration tests now run in CI
against a real database.
Verification
Deployed and verified on the live origin. Every route that answered 404 now
answers 200 with real data; Charts renders 91 chart paths with 0 spinners;
Mining renders real reward and pool figures with 0 skeletons; Protocols reports
0 of 31 can be read right nowwith Ordinals shown asCatching up, 604570 blocks behind, which is true while the ord index rebuilds.174 frontend unit tests, 257 overlay tests, backend unit and integration tests,
and all repository gates pass.
Pairs with bitcoinuniverseio/backend-apis#58 for the overlay side, merged as
f908097band already answering on the live origin with the versionedavailability contract.
Follow-up in this PR
The gate's first full run found the same fault on four more routes: home and
blocks held placeholders with nothing said about why when the chain backend is
down, and the transaction and address pages rendered skeletons a screen reader
hears nothing about. All four are fixed and now gated. Two axe findings from
the same run, keyboard access to the dashboard's scrollable tables and a
tracker stage at 4.28:1, are fixed too.
The frontend job also stopped serving its build on a fixed port: two runners
share a host, and a run once measured another job's gateway until that job
ended and took the port down mid-run. Each run now serves on its own free port
and proves the process answering is the one it started.
🤖 Generated with Claude Code