Skip to content

The C-94 findability guard checks 2 of 110 artefacts — and cannot see the sidecar by construction #312

Description

@Polichinel

Reported by the views-models seat from the first live UN-FAO delivery, 2026-09-29. The run completed, logged success, sent a WandB success alert, and produced a delivery views-faoapi refuses at ingest. Bucket listing: run-0 has 110 objects, this run 109. The missing object is the §5 GAUL sidecar.

Verified here against the code before filing. The report is right about the hole and understates it.

The guard checks 2 of 110 artefacts — and cannot see the sidecar at all

unfao/managers/unfao.py:393 calls the C-94 preflight with two ids:

{"forecast": summary["manifest_file_id"], "historical": hist_file_id}

That is the commit marker and the historical leg. Not the sidecar, not the 108 shards.

But adding them to that dict would not work either, and that is the part worth reading twice. The guard resolves each entry with:

port.latest_file_id({"name": consumer_name, "category": category})

In contract/wire/sink.py:227 every wire object is uploaded with the same common dict:

common = {"name": consumer_name, "category": "forecast", "loa": "pgm"}

Shards, sidecar and manifest all carry category="forecast". So a query for the newest forecast document returns the manifest — which is uploaded last, by design (sink.py:252). The guard is structurally incapable of observing the sidecar. It is not a missing entry in a dict; the lookup key is the wrong key.

What actually failed

The sidecar's bytes were byte-identical to run-0's, the store deduplicated, update_document ran against the old file, and the uploader logged the new filename as uploaded. _ContractStorePort.upload returned a real file_id — for the wrong document. Nothing in this repository asks whether that id resolves under the name the manifest references.

views-faoapi resolves by filename. It finds nothing. It refuses. Correctly.

The consumer seat's phrasing, which belongs in the register entry:

Manifest-last protects against a run that stops early. It does not protect against an upload that returns success having written nothing.

That is C-79 one level down. C-79 is about an upload that reports failure with the file already written; this is an upload that reports success with the file not written. The guard's own docstring — "every call reported success in run-0 too, and the historical leg still stranded" — describes this exact failure and then checks the two legs that did not exhibit it.

Two things in the report that are NOT findings

Checked, because acting on them would mean changing code that is already correct:

  1. "The sidecar must land before the manifest, always." It already does. sink.py:215 — shards → sidecar → manifest LAST — and sink.py:246-252 implements it. ADR-013 already carries the rule as a dated MINOR clarification (2026-07-19, §4 audit, register C-51): "the manifest uploads only after every shard AND the sidecar." The ordering was obeyed. The manifest is in the bucket without the sidecar not because it went first, but because the sidecar's upload lied about having landed. Encoding an ordering rule again would treat a correct mechanism as the cause.

  2. The curation was right. Manifest declares expected_cell_count: 64742, land_gaul, unmapped_count: 0. Saying so because it was the pre-registered most-likely silent failure and it did not happen.

The shape of the fix — design mine, stated so it can be argued with

Check resolvability by name, for every artefact the manifest references.

  • The manifest is in hand at the call site and already carries sidecar.name + sha256 (run_manifest.py:56) and shards[*].name. Nothing new needs plumbing.
  • filename is already passed to upload_data (store_port.py:66-76), so it is a queryable field — latest_file_id passes filters straight through.
  • An id is not evidence. The id returned here was real. The check must be "does a document resolvable by this filename exist", not "did I get a non-empty id back". A guard written the second way passes this incident.
  • 110 lookups per delivery is the obvious cost objection. It is one query per object against a store the run has just written to, on a path that runs once a month. Worth measuring before optimising, not worth pre-optimising into a sampled check that would miss exactly this.

Open for review: whether the shard set is verified per-object or by count-plus-spot-check. I lean per-object — the failure mode is one object silently absent, which a count catches only if the store also fails to create a document, and here it did create one.

Relationship to the other half

views-pipeline-core is separately fixing the upload-that-reports-success. Complementary, not alternatives, and neither should wait on the other: theirs stops the lie at the source, this one stops us believing a lie from any source. A store we do not own can regress; the guard is ours.

Register: extends C-94. Related: C-79, C-99 (fail-closed on an unrecognised result), C-105 (the torn-run ledger, which also did not fire — nothing failed).

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions