fix(iri): mint R + base62 IRIs, and repair the collision guard - #19
Conversation
generate_iri() minted a base64url-derived token with no R prefix. That matches 2 of the 18,325 published FOLIO concepts. 11,427 use R + base62, including every concept a new sibling is likely to sit beside, so a freshly minted IRI was visibly foreign to the ontology it was joining. Two defects came out of the same code path: - The uniqueness check tested a bare local name for membership in FOLIO.iri_to_index, whose keys are full IRIs taken from rdf:about. No candidate was ever rejected; the MAX_IRI_ATTEMPTS loop could not run. - Non-alphanumeric characters were stripped after base64 encoding, so an unlucky uuid4 lost characters. About 1 mint in 29,000 fell to <= 16 characters, which the library's own test asserts against. uuid ffffffff-ffff-4fff-bfff-ffffffffffff encodes to the 2-character "Tw". The generator moves to folio/iri.py, standard library only, so minting no longer requires constructing a FOLIO graph and downloading 18 MB of OWL to produce one 23-character string. The 127-bit width is recovered from the published data, not chosen: body lengths of 20/21/22 occur at 0.39%/25.75%/73.85%, and base62 of 127 bits predicts 0.39%/25.31%/74.29%. 128 bits predicts 0.21%/12.9%/86.9% and is excluded. It is pinned by a test. No existing IRI changes. generate_iri() only mints new ones. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_019g78dzGEqYn2iTNzR11Jfv
The new files failed `ruff format --check` and added two `ruff check` diagnostics (UP035, UP045) that main does not have. CI only runs pytest, so neither showed up as a failing check. After this the branch matches main's lint baseline exactly: 188 ruff errors (all pre-existing), `ruff format --check` clean, `ty check folio/` at 26 diagnostics. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
Confirmed — dropping the I traced the history before answering, since your question was whether there was a decision behind the current shape. There wasn't one that's recorded: The one piece of visible intent is the original comment, truncated mid-sentence: # only use alphanumeric characters from restricted b64 encdoding toSo "alphanumeric only" does look deliberate. Worth noting explicitly in your favor: base62 preserves that intent while removing the post-hoc stripping that causes the short-name bug. This PR isn't reversing the one decision that was actually made here. I reproduced all four quantitative claims independently against
The 127-bit fit depends on excluding the 4,751 length-23 names as SALI hex legacy, so I checked that too: all 4,751 are case-insensitive hex. Since 127 bits caps base62 output at 22 digits, those names provably can't share a generator with the rest — the exclusion is sound rather than convenient. Two things worth adding, because they make the case better than the majority-share argument does. First, the real justification for Second, the blast radius is zero, which is why I'm comfortable taking the behaviour change. Family counts are byte-identical between the ontology's initial commit ( One thing to flag for next time: the new files failed Shipping as 0.4.0. |
Ships #19. Dates the 0.4.0 changelog entry and bumps pyproject.toml, folio/__init__.py, and the lockfile to 0.4.0. Adds one changelog entry that #19 did not claim for itself: minted local names previously began with a digit about 16% of the time, which is not a valid XML NCName, so an RDF/XML serializer could not form a QName without splitting the namespace mid-token (protegeproject/webprotege#18). The R prefix removes this by construction. That is the strongest argument for the change and it belongs in the record. The lockfile diff is a marker normalization from a newer uv plus the version bump; no dependency versions moved and no hashes changed. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@mjbommar — flagging you directly because the first change here is a judgment call
that's really yours, not mine.
generate_iri()mints a base64url-derived token with noRprefix. I went lookingfor the convention while adding a single class to
FOLIO.owland found that shapematches 2 of the 18,325 published concepts. The other 18,323 look nothing like
it. Breakdown of
owl:Classlocal names atalea-institute/FOLIO@8ebf17b:R+ base62R+ 23 hex characters (SALI legacy)R+ base64urlgenerate_iri()The intent of this PR is to continue the scheme the vast majority of FOLIO items
already use, so a concept minted today is indistinguishable from the siblings it
sits beside. Concretely:
Investment FundsisRCzxVprwB3RwZ8EMewt18JtandHedge FundisR7e7pNl5IOMFbxKN2GV2C41; a new child of theirs minted by thecurrent code would be
07BzhNmkTx6PPCobTF1ufw.If dropping the
Rprefix was deliberate rather than incidental, say so andI'll close this. That's a real possibility — the docstring says the function
"approximates the WebProtégé IRI generation algorithm", and WebProtégé genuinely
does emit bare base64url. But it emits it with
-and_and always at 22characters, which is the 1,992-member family, not what this function produces. So
either way the current output matches nothing. What shouldn't happen is both
schemes staying in use, which is the situation today.
Deriving the width
127 bits isn't a guess. Published
R-family body lengths of 20/21/22 charactersoccur at 0.39% / 25.75% / 73.85% (n = 11,427). Base62 of a uniform 127-bit
integer predicts 0.39% / 25.31% / 74.29%. A 128-bit generator predicts
0.21% / 12.9% / 86.9% and is excluded. It's pinned by a test, because that length
profile is the only surviving evidence of how these were originally generated.
Two defects fixed along the way
Both fell out of the same function:
1. The uniqueness check never fires.
iri_to_indexis keyed by full IRI, straight fromrdf:about(graph.pyline 782). A bare local name is never a key, so no candidate is ever rejected and
the
MAX_IRI_ATTEMPTSloop can't iterate. Verified against the real ontology:Collision risk with uuid4 is negligible, so this is latent rather than harmful —
but the guard is dead code.
generate_iri()now accepts either spelling.2. Short local names. Non-alphanumeric characters are stripped after
base64 encoding, so an unlucky uuid4 loses characters:
About 1 mint in 29,000 falls to ≤16 characters — which trips the library's own
assert len(b64_token) > 16intest_iri_generation. Base62 has no strippingstep, so the failure mode is gone rather than guarded.
Structure
The generator moves to a new
folio/iri.py, standard library only. Mintingpreviously required constructing a
FOLIOgraph, i.e. downloading and parsing18 MB of OWL to produce one 23-character string.
FOLIO.generate_iri()keeps itssignature and delegates;
MAX_IRI_ATTEMPTSstays importable fromfolio.graph.(
folio/__init__.pystill eagerly imports.graph, so importingfolio.irithrough the package pulls httpx and lxml anyway. Making the package init lazy
would fix that too — happy to do it here or separately if you want it.)
Compatibility
No existing IRI changes;
generate_iri()only mints new ones. Public API isunchanged apart from the added module.
CHANGES.mdproposes 0.4.0 on thebehaviour change — version bump left to you.
Tests
tests/test_iri.py, 15 tests: base62 round-tripping and digit order, outputshape, the pinned entropy width, both collision-index spellings, retry bounds,
distinctness, and the shape of four real published IRIs.
test_iri_generationin
tests/test_folio.pygains anR-prefix assertion and a real collision check.Verification note
The full suite wasn't run in my environment —
httpx/lxmlweren't installed,so
tests/test_iri.pywas exercised againstfolio/iri.pydirectly (15 passed).folio/iri.pyis standard-library-only, so nothing in it depends on those. Thegraph-level test needs a normal install to run.
🤖 Generated with Claude Code
https://claude.ai/code/session_019g78dzGEqYn2iTNzR11Jfv