Skip to content

§5.1 ruling: is float64 for GAUL codes a wire decision or a pandas artefact? (nullable int64 not considered) #278

Description

@Polichinel

FAO's contractor has now asked twice for GAUL codes as integers — in March 2026, when we said yes and shipped the change, and again on 14 August after the global re-upload reverted it. The current answer we are preparing to send is "cast to integer on read", which is defensible but is the second time we tell them no on the same point.

Before that goes out, I want §5.1's authors to rule on a third option that I cannot find considered anywhere in the repo.

The two options that have been considered

float64 (current). views_postprocessing/contract/gaul_schema.py states the ruling directly:

Wire dtypes are the §5.1 ruling: codes are ALWAYS float64 (one stable schema regardless of whether a run contains a missing code)

int64 + sentinel. Already rejected, and rightly. docs/fao_excluded_cells.md says the uncovered cells are dropped at region level "rather than shipped with a placeholder — shipping them would attribute partner rows to a non-country (-1)". These are identifiers, not quantities: a -1 GAUL code is a fake place that groups, joins and aggregates like a real one. Silent corruption. No argument from me.

The option I cannot find discussed: nullable int64

Parquet and Arrow both carry nullable integers natively — values plus a validity bitmap. That gives integers on the wire, a real null, and one stable schema. All three properties §5.1 wants.

The stated justification — that an integer column cannot hold a missing value — is a property of NumPy-backed pandas, which upcasts int64 to float64 the moment a NaN appears. It is not a property of Parquet, which is what we actually deliver.

So the question is whether §5.1's float ruling is a considered wire decision, or an artefact of the build stack that got written down as one. I genuinely do not know which, and the docstring does not say.

What makes this non-trivial, and why I am asking rather than doing

  1. Column order and dtype are normative under ADR-013 §5.1, byte-pinned by the §10 golden fixture. Changing dtype re-cuts the fixture deliberately.
  2. The consumer hard-validates: views-faoapi handlers.py, FAO_PGMDataset._METADATA_COLS. Both ends move in lockstep or the delivery fails validation.
  3. Round-trip behaviour needs checking, not assuming. If FAO read the parquet with default pandas and a nullable Int64 column degrades back to float or object on their side, we have added contract churn for nothing. That is the empirical question I would want answered before anyone writes code.

The part that makes this awkward to defend as-is

For this delivery the code column never contains a null. The 76 GAUL-uncovered cells are removed before delivery, the exclusion list is frozen in delivery/coverage.py, and tests/test_delivery_coverage.py asserts it against the producer.

So the float type is defending against a case this product structurally cannot produce. That is a coherent position — schema should not depend on data — but it is a hard one to explain to a consumer who has never once received a missing code and has asked twice for integers.

What I need

A ruling, not a patch:

  • Is the float64 choice a considered wire decision, or inherited from pandas dtype behaviour?
  • Does nullable int64 satisfy §5.1's stability requirement? If not, why not — I would like the reason written down where the next person finds it.
  • If it does: what does the change actually cost across gaul_schema.py, the §10 fixture, and the faoapi consumer, and is it worth it?
  • If it does not: a sentence in the gaul_schema.py docstring saying nullable int was considered and rejected, so this question is not reopened every time a consumer asks.

Either answer is fine. I would just rather tell FAO "fixed" than tell them no twice, and if the answer is no, I want to give them a real reason rather than one that only holds in pandas.

Related: views-postprocessing#272, views-faoapi#407.

Activity

  1. added a commit that references this issue on Aug 17, 2026
  2. Polichinel commented on Aug 21, 2026

    @Polichinel
    CollaboratorAuthor

    Shipped to main in #283 (via #279).

    The ruling: float64 stays, but not for the reason we had been giving. You were right that the option was unconsidered, and right that the stated justification was a pandas property rather than a parquet one. Arrow carries nullable integers natively, and this repo's own data/gaul_lookup.parquet stores all three code columns as int64 with zero nulls — the float is introduced by our writer.

    It survives on a measured ground instead. pyarrow 23.0.1 / pandas 3.0.5:

    written pandas default read
    int64, no nulls int64
    int64, one null float64

    So nullable int64 does not remove the float — it moves it from our writer to the consumer's reader and makes it appear only sometimes, which is exactly the data-dependent schema §5.1 exists to prevent. Independently, faoapi's wire_reader reindexes the sidecar and calls .to_numpy(); both yield float64 from a nullable integer column, so it would not have reached them as integers anyway.

    Recorded where you asked for it: ADR-013 §5.1a, with contract/gaul_schema.py's docstring stating it at the declaration and pointing there — so the next person asked finds the answer where they are standing rather than reopening it a third time. No contract_version bump: it records a rejected alternative, it does not change the rule.

    The useful half, which is what went to FAO: because the delivered region drops the GAUL-uncovered cells, no delivered code is ever missing — held in CI by tests/test_gaul_lookup_fidelity.py::test_lookup_has_no_nulls — so astype("int64") on read is lossless for this product. That is a property of the region, not the contract, and the mail says so.

    On "treat the 'was it ever actually fixed' question as live": it was fixed in March, then deliberately changed back on 2026-07-19, and nobody told them. The reply carries that, with the apology attached to the not-telling rather than to the type.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions