Skip to content

fix(backend): split liveness from readiness so an ES outage is not a green probe - #3385

Merged
Xore merged 1 commit into
mainfrom
oc/3317-backend-liveness-split
Sep 27, 2026
Merged

Xore merged 1 commit into
mainfrom
oc/3317-backend-liveness-split

Conversation

@Xore

@Xore Xore commented Sep 27, 2026

Copy link
Copy Markdown
Owner

/healthz returned {"ok": true, "es": <bool>} with ok hardcoded to true, so apiary-backend reported healthy to its own compose healthcheck and to anything else probing it even when it could not reach Elasticsearch at all. The es field was the only real signal in the response and it could not change the status code — a backend with a dead dependency and a backend with a working one were the same 200. The #3283 ingest outage would not have shown up there.

What changes

/livez is "the process is up". It never touches Elasticsearch, and that is the load-bearing part rather than an omission: a probe that can block on a dependency converts that dependency's outage into a restart loop of a container that was never the thing that broke. The image HEALTHCHECK moves to it, with a comment saying why it must not be /readyz.

/readyz is "this can do its job": ES reachable, not red, and no index.blocks.write on any index family this tier writes to. 503 with a reason that names the blocked indices rather than counting them.

/healthz stays as an alias for /livez. The port-test harness, the ops scripts and the image healthcheck have all used that name for years, and breaking a health-probe name is how a stack ends up reporting permanently unhealthy with nothing wrong.

$ curl -s -o /dev/null -w '%{http_code}\n' http://backend-service:8081/livez     # ES write-blocked
200        # container stays up, no restart loop
$ curl -s http://backend-service:8081/readyz
{"ready":false,"reason":"elasticsearch has index.blocks.write set on: dashboard-config-v1. The
flood-stage disk watermark sets this on every index at once, and so does an operator's
`PUT /<index>/_block/write`; this endpoint cannot tell those apart, so check _cat/allocation free
space before concluding which one it is.","cluster":"green","write_blocked":["dashboard-config-v1"]}

Decisions worth reviewing, because each is a way to be wrong quietly

  • Yellow stays ready. It means unassigned replicas, the ordinary shape of a replicated cluster mid-rolling-restart. Gating on it would take the backend out of service on every deploy. Red is not ready.
  • A missing index is not a block. Every dashboard-owned index is created lazily on its first write, so a fresh cluster correctly reads ready — verified, not assumed.
  • Unreachable outranks a block report. A write-block reading taken against a cluster we could not reach is the absence of a fact; reporting it as the cause would send an operator after the wrong thing.
  • 5s READINESS_TIMEOUT, not the client's 30s. The shared transport budget is sized for real multi-second queries. A probe that blocks for 30s is indistinguishable from the outage it exists to report. Both probes run concurrently.
  • es::WRITE_TARGET_FAMILIES is the complement of EVENT_INDICES, not a copy. A read-only index somebody else owns going write-blocked is that tier's outage, and reporting not-ready for it would train operators to ignore the endpoint. dashboard-* is one wildcard rather than twenty concrete names so it cannot fall out of date the way a hand-copied list can.

Verification

Unit tests cover the decision truth table (the function is pure, so the interesting cases — two probes disagreeing — are reachable without a live cluster). Beyond that, run against a real Elasticsearch 9.5.3 (the deployed major):

Scenario /livez /readyz
green, nothing blocked 200 200 write_blocked: []
green, dashboard-config-v1 write-blocked 200 503, names the index
three indices blocked 200 503, all three, sorted
a read-only index (ml-worker-metrics) blocked 200 200 — correctly not gated
Elasticsearch unreachable 200 503, carrying the failing URL
fresh cluster, no dashboard indices yet 200 200 — lazy creation, not a block

cargo test 540 passed / 0 failed; cargo clippy --all-targets -- -D warnings clean.

Also in here

  • Es::ping() is now dead and is deleted rather than left as a tempting way to reintroduce this.
  • Two comments that described the old probe as a ping() caller are corrected. es.rs's own transport-timeout rationale cited /healthz as the reason a 30s bound was needed, which is no longer true of the probe path — the bound now stands only for the real queries and worker loops that can saturate the search queue.
  • /livez and /readyz get their own bounded metrics families, so an operator watching request rates can tell a readiness probe's 503s from an unrelated misshapen path.
  • Contract documented in docs/OPERATIONS.md, including what the endpoint deliberately cannot distinguish (flood-stage watermark vs. an operator's _block/write).

Docs linters (check-doc-paths-exist, check-docs-reachable, check-doc-stale-paths, check-backend-boot-contract, check-compose-env-docs) all pass. The test container and process used for verification were removed.

Refs #3317

…green probe

/healthz answered {"ok": true, "es": <bool>} with `ok` hardcoded to true, so
apiary-backend reported healthy to its own compose healthcheck and to
anything else that probed it even when it could not reach Elasticsearch at
all. The `es` field was the only real signal in the response and it could
not change the status code. A backend with a dead dependency and a backend
with a working one were the same 200 -- the #3283 ingest outage would not
have shown up there.

One question, two endpoints:

- /livez is "the process is up". It never touches Elasticsearch, and that
  is the load-bearing part rather than an omission: a probe that can block
  on a dependency converts that dependency's outage into a restart loop of
  a container that was never the thing that broke. The image HEALTHCHECK
  moves to it, with a comment saying why it must not be /readyz.
- /readyz is "this can do its job": ES reachable, not red, and no
  index.blocks.write on any index family this tier writes to. 503 with a
  `reason` that names the blocked indices rather than counting them.

/healthz stays as an alias for /livez. The port-test harness, the ops
scripts and the image healthcheck have all used that name for years, and
breaking a health-probe name is how a stack ends up reporting permanently
unhealthy with nothing wrong.

Decisions worth stating, because each is a way to be wrong quietly:

- Yellow stays ready. It means unassigned replicas, which is the ordinary
  shape of a replicated cluster mid-rolling-restart; gating on it would
  take the backend out of service on every deploy.
- Red is not ready, and a *missing* index is not a block. Every
  dashboard-owned index is created lazily on its first write, so a fresh
  cluster correctly reads ready -- verified, not assumed.
- Unreachable outranks a block report. A write-block reading taken against
  a cluster we could not reach is the absence of a fact, and reporting it
  as the cause would send an operator after the wrong thing.
- Both probes run concurrently under a 5s READINESS_TIMEOUT rather than
  inheriting the shared client's 30s, which is sized for real multi-second
  queries. A probe that blocks for 30s is indistinguishable from the
  outage it exists to report.

es::WRITE_TARGET_FAMILIES is the explicit inventory of what this tier
writes, deliberately the complement of EVENT_INDICES rather than a copy:
a read-only index somebody else owns going write-blocked is that tier's
outage, and reporting not-ready for it would train operators to ignore the
endpoint. `dashboard-*` is one wildcard rather than twenty concrete names
so it cannot fall out of date the way a hand-copied list can.

Verified against a real Elasticsearch 9.5.3 (the deployed major), not just
unit tests: green and unblocked -> 200; green with dashboard-config-v1
write-blocked -> 503 naming it while /livez stays 200; ES unreachable ->
503 carrying the failing URL; a write-blocked read-only index does not
gate readiness; multiple blocks all named, sorted.

Es::ping() is now dead and is deleted rather than left as a tempting way to
reintroduce this, and the two comments that described the old probe as a
ping caller are corrected -- es.rs's own transport-timeout rationale
mentioned /healthz as the reason a 30s bound was needed, which is no longer
true of the probe path. /livez and /readyz get their own bounded metrics
families so an operator watching request rates can tell a readiness
probe's 503s from an unrelated misshapen path. Contract documented in
docs/OPERATIONS.md, including what the endpoint deliberately cannot
distinguish (flood-stage watermark vs. an operator's _block/write).

Refs #3317
@github-actions

Copy link
Copy Markdown

Dependency Review

✅ No vulnerabilities or license issues or OpenSSF Scorecard issues found.

Scanned Files

None

@Xore
Xore enabled auto-merge (squash) September 27, 2026 02:11
@Xore
Xore merged commit bf54a34 into main Sep 27, 2026
120 of 121 checks passed
@Xore
Xore deleted the oc/3317-backend-liveness-split branch September 27, 2026 02:24
Xore pushed a commit that referenced this pull request Sep 27, 2026
Both files conflicted because #3385 (liveness/readiness split) and #3315
(revision stamp) landed on the same two places.

backend-service/src/main.rs — #3385 replaced the single `Health { ok, es }`
struct with `Liveness`/`Readiness` and three handlers, so #3315's `revision`
field had nowhere to land. Rather than resurrect the old struct, the field
moves onto `Liveness`: /healthz stays an alias of /livez, so the deploy
verifier and diagnostics both still read `revision` off the same body, and
#3317's property (liveness never touches Elasticsearch) is preserved because
git_revision() reads a compile-time env, not the cluster. `built` and
`revision` are both kept -- they are not redundant: a timestamp can only be
compared, an object name can be looked up, and an image rebuilt from an old
commit passes the timestamp check while running month-old code.

containers.yml — #3316 added `boot_smoke` to the same two matrix rows #3315
added `stamp` to, and both needed the `labels:` line. Resolution is the union:
both rows carry both flags, `labels` keeps #3316's per-row build-row handle,
and `build-args` is added alongside it.

Local: cargo test 543 passed / 0 failed / 1 pre-existing ignored (544 total,
unchanged from main); cargo clippy --all-targets -- -D warnings clean;
cargo build --all-targets clean; also green with GIT_SHA set.
test_3315_image_revision.py 40 passed; test_3316_image_boot_smoke.py 44 passed.
zizmor --offline --min-severity medium reports the same single allowlisted
dangerous-triggers advisory as main, and no new findings.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant