fix(backend): split liveness from readiness so an ES outage is not a green probe - #3385
Merged
Merged
Conversation
…green probe
/healthz answered {"ok": true, "es": <bool>} with `ok` hardcoded to true, so
apiary-backend reported healthy to its own compose healthcheck and to
anything else that probed it even when it could not reach Elasticsearch at
all. The `es` field was the only real signal in the response and it could
not change the status code. A backend with a dead dependency and a backend
with a working one were the same 200 -- the #3283 ingest outage would not
have shown up there.
One question, two endpoints:
- /livez is "the process is up". It never touches Elasticsearch, and that
is the load-bearing part rather than an omission: a probe that can block
on a dependency converts that dependency's outage into a restart loop of
a container that was never the thing that broke. The image HEALTHCHECK
moves to it, with a comment saying why it must not be /readyz.
- /readyz is "this can do its job": ES reachable, not red, and no
index.blocks.write on any index family this tier writes to. 503 with a
`reason` that names the blocked indices rather than counting them.
/healthz stays as an alias for /livez. The port-test harness, the ops
scripts and the image healthcheck have all used that name for years, and
breaking a health-probe name is how a stack ends up reporting permanently
unhealthy with nothing wrong.
Decisions worth stating, because each is a way to be wrong quietly:
- Yellow stays ready. It means unassigned replicas, which is the ordinary
shape of a replicated cluster mid-rolling-restart; gating on it would
take the backend out of service on every deploy.
- Red is not ready, and a *missing* index is not a block. Every
dashboard-owned index is created lazily on its first write, so a fresh
cluster correctly reads ready -- verified, not assumed.
- Unreachable outranks a block report. A write-block reading taken against
a cluster we could not reach is the absence of a fact, and reporting it
as the cause would send an operator after the wrong thing.
- Both probes run concurrently under a 5s READINESS_TIMEOUT rather than
inheriting the shared client's 30s, which is sized for real multi-second
queries. A probe that blocks for 30s is indistinguishable from the
outage it exists to report.
es::WRITE_TARGET_FAMILIES is the explicit inventory of what this tier
writes, deliberately the complement of EVENT_INDICES rather than a copy:
a read-only index somebody else owns going write-blocked is that tier's
outage, and reporting not-ready for it would train operators to ignore the
endpoint. `dashboard-*` is one wildcard rather than twenty concrete names
so it cannot fall out of date the way a hand-copied list can.
Verified against a real Elasticsearch 9.5.3 (the deployed major), not just
unit tests: green and unblocked -> 200; green with dashboard-config-v1
write-blocked -> 503 naming it while /livez stays 200; ES unreachable ->
503 carrying the failing URL; a write-blocked read-only index does not
gate readiness; multiple blocks all named, sorted.
Es::ping() is now dead and is deleted rather than left as a tempting way to
reintroduce this, and the two comments that described the old probe as a
ping caller are corrected -- es.rs's own transport-timeout rationale
mentioned /healthz as the reason a 30s bound was needed, which is no longer
true of the probe path. /livez and /readyz get their own bounded metrics
families so an operator watching request rates can tell a readiness
probe's 503s from an unrelated misshapen path. Contract documented in
docs/OPERATIONS.md, including what the endpoint deliberately cannot
distinguish (flood-stage watermark vs. an operator's _block/write).
Refs #3317
Dependency Review✅ No vulnerabilities or license issues or OpenSSF Scorecard issues found.Scanned FilesNone |
Xore
enabled auto-merge (squash)
September 27, 2026 02:11
Xore
pushed a commit
that referenced
this pull request
Sep 27, 2026
Both files conflicted because #3385 (liveness/readiness split) and #3315 (revision stamp) landed on the same two places. backend-service/src/main.rs — #3385 replaced the single `Health { ok, es }` struct with `Liveness`/`Readiness` and three handlers, so #3315's `revision` field had nowhere to land. Rather than resurrect the old struct, the field moves onto `Liveness`: /healthz stays an alias of /livez, so the deploy verifier and diagnostics both still read `revision` off the same body, and #3317's property (liveness never touches Elasticsearch) is preserved because git_revision() reads a compile-time env, not the cluster. `built` and `revision` are both kept -- they are not redundant: a timestamp can only be compared, an object name can be looked up, and an image rebuilt from an old commit passes the timestamp check while running month-old code. containers.yml — #3316 added `boot_smoke` to the same two matrix rows #3315 added `stamp` to, and both needed the `labels:` line. Resolution is the union: both rows carry both flags, `labels` keeps #3316's per-row build-row handle, and `build-args` is added alongside it. Local: cargo test 543 passed / 0 failed / 1 pre-existing ignored (544 total, unchanged from main); cargo clippy --all-targets -- -D warnings clean; cargo build --all-targets clean; also green with GIT_SHA set. test_3315_image_revision.py 40 passed; test_3316_image_boot_smoke.py 44 passed. zizmor --offline --min-severity medium reports the same single allowlisted dangerous-triggers advisory as main, and no new findings.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
/healthzreturned{"ok": true, "es": <bool>}withokhardcoded totrue, soapiary-backendreported healthy to its own compose healthcheck and to anything else probing it even when it could not reach Elasticsearch at all. Theesfield was the only real signal in the response and it could not change the status code — a backend with a dead dependency and a backend with a working one were the same200. The #3283 ingest outage would not have shown up there.What changes
/livezis "the process is up". It never touches Elasticsearch, and that is the load-bearing part rather than an omission: a probe that can block on a dependency converts that dependency's outage into a restart loop of a container that was never the thing that broke. The imageHEALTHCHECKmoves to it, with a comment saying why it must not be/readyz./readyzis "this can do its job": ES reachable, not red, and noindex.blocks.writeon any index family this tier writes to. 503 with areasonthat names the blocked indices rather than counting them./healthzstays as an alias for/livez. The port-test harness, the ops scripts and the image healthcheck have all used that name for years, and breaking a health-probe name is how a stack ends up reporting permanently unhealthy with nothing wrong.Decisions worth reviewing, because each is a way to be wrong quietly
ready— verified, not assumed.READINESS_TIMEOUT, not the client's 30s. The shared transport budget is sized for real multi-second queries. A probe that blocks for 30s is indistinguishable from the outage it exists to report. Both probes run concurrently.es::WRITE_TARGET_FAMILIESis the complement ofEVENT_INDICES, not a copy. A read-only index somebody else owns going write-blocked is that tier's outage, and reporting not-ready for it would train operators to ignore the endpoint.dashboard-*is one wildcard rather than twenty concrete names so it cannot fall out of date the way a hand-copied list can.Verification
Unit tests cover the decision truth table (the function is pure, so the interesting cases — two probes disagreeing — are reachable without a live cluster). Beyond that, run against a real Elasticsearch 9.5.3 (the deployed major):
/livez/readyzwrite_blocked: []dashboard-config-v1write-blockedml-worker-metrics) blockedcargo test540 passed / 0 failed;cargo clippy --all-targets -- -D warningsclean.Also in here
Es::ping()is now dead and is deleted rather than left as a tempting way to reintroduce this.ping()caller are corrected.es.rs's own transport-timeout rationale cited/healthzas the reason a 30s bound was needed, which is no longer true of the probe path — the bound now stands only for the real queries and worker loops that can saturate the search queue./livezand/readyzget their own bounded metrics families, so an operator watching request rates can tell a readiness probe's 503s from an unrelated misshapen path.docs/OPERATIONS.md, including what the endpoint deliberately cannot distinguish (flood-stage watermark vs. an operator's_block/write).Docs linters (
check-doc-paths-exist,check-docs-reachable,check-doc-stale-paths,check-backend-boot-contract,check-compose-env-docs) all pass. The test container and process used for verification were removed.Refs #3317