Skip to content

fix(iri): mint R + base62 IRIs, and repair the collision guard - #19

Merged
mjbommar merged 2 commits into
mainfrom
iri-minting-r-base62
Aug 18, 2026
Merged

mjbommar merged 2 commits into
mainfrom
iri-minting-r-base62

Conversation

@damienriehl

Copy link
Copy Markdown
Contributor

@mjbommar — flagging you directly because the first change here is a judgment call
that's really yours, not mine.

generate_iri() mints a base64url-derived token with no R prefix. I went looking
for the convention while adding a single class to FOLIO.owl and found that shape
matches 2 of the 18,325 published concepts. The other 18,323 look nothing like
it. Breakdown of owl:Class local names at alea-institute/FOLIO@8ebf17b:

family count share
R + base62 11,427 62.4%
R + 23 hex characters (SALI legacy) 4,751 25.9%
raw base64url of a uuid4, 22 chars 1,992 10.9%
R + base64url 151 0.8%
bare alphanumeric — current generate_iri() 2 0.01%
other 2 0.01%

The intent of this PR is to continue the scheme the vast majority of FOLIO items
already use, so a concept minted today is indistinguishable from the siblings it
sits beside. Concretely: Investment Funds is RCzxVprwB3RwZ8EMewt18Jt and
Hedge Fund is R7e7pNl5IOMFbxKN2GV2C41; a new child of theirs minted by the
current code would be 07BzhNmkTx6PPCobTF1ufw.

If dropping the R prefix was deliberate rather than incidental, say so and
I'll close this.
That's a real possibility — the docstring says the function
"approximates the WebProtégé IRI generation algorithm", and WebProtégé genuinely
does emit bare base64url. But it emits it with - and _ and always at 22
characters, which is the 1,992-member family, not what this function produces. So
either way the current output matches nothing. What shouldn't happen is both
schemes staying in use, which is the situation today.

Deriving the width

127 bits isn't a guess. Published R-family body lengths of 20/21/22 characters
occur at 0.39% / 25.75% / 73.85% (n = 11,427). Base62 of a uniform 127-bit
integer predicts 0.39% / 25.31% / 74.29%. A 128-bit generator predicts
0.21% / 12.9% / 86.9% and is excluded. It's pinned by a test, because that length
profile is the only surviving evidence of how these were originally generated.

Two defects fixed along the way

Both fell out of the same function:

1. The uniqueness check never fires.

if base64_value in self.iri_to_index:   # bare local name
    continue

iri_to_index is keyed by full IRI, straight from rdf:about (graph.py
line 782). A bare local name is never a key, so no candidate is ever rejected and
the MAX_IRI_ATTEMPTS loop can't iterate. Verified against the real ontology:

full IRI is a key:        True
bare local name is a key: False
keys that are bare local names: 0 of 18,327

Collision risk with uuid4 is negligible, so this is latent rather than harmful —
but the guard is dead code. generate_iri() now accepts either spelling.

2. Short local names. Non-alphanumeric characters are stripped after
base64 encoding, so an unlucky uuid4 loses characters:

>>> generate(uuid.UUID("ffffffff-ffff-4fff-bfff-ffffffffffff"))
'Tw'

About 1 mint in 29,000 falls to ≤16 characters — which trips the library's own
assert len(b64_token) > 16 in test_iri_generation. Base62 has no stripping
step, so the failure mode is gone rather than guarded.

Structure

The generator moves to a new folio/iri.py, standard library only. Minting
previously required constructing a FOLIO graph, i.e. downloading and parsing
18 MB of OWL to produce one 23-character string. FOLIO.generate_iri() keeps its
signature and delegates; MAX_IRI_ATTEMPTS stays importable from folio.graph.

(folio/__init__.py still eagerly imports .graph, so importing folio.iri
through the package pulls httpx and lxml anyway. Making the package init lazy
would fix that too — happy to do it here or separately if you want it.)

Compatibility

No existing IRI changes; generate_iri() only mints new ones. Public API is
unchanged apart from the added module. CHANGES.md proposes 0.4.0 on the
behaviour change — version bump left to you.

Tests

tests/test_iri.py, 15 tests: base62 round-tripping and digit order, output
shape, the pinned entropy width, both collision-index spellings, retry bounds,
distinctness, and the shape of four real published IRIs. test_iri_generation
in tests/test_folio.py gains an R-prefix assertion and a real collision check.

Verification note

The full suite wasn't run in my environment — httpx/lxml weren't installed,
so tests/test_iri.py was exercised against folio/iri.py directly (15 passed).
folio/iri.py is standard-library-only, so nothing in it depends on those. The
graph-level test needs a normal install to run.


🤖 Generated with Claude Code

https://claude.ai/code/session_019g78dzGEqYn2iTNzR11Jfv

generate_iri() minted a base64url-derived token with no R prefix. That
matches 2 of the 18,325 published FOLIO concepts. 11,427 use R + base62,
including every concept a new sibling is likely to sit beside, so a
freshly minted IRI was visibly foreign to the ontology it was joining.

Two defects came out of the same code path:

- The uniqueness check tested a bare local name for membership in
  FOLIO.iri_to_index, whose keys are full IRIs taken from rdf:about. No
  candidate was ever rejected; the MAX_IRI_ATTEMPTS loop could not run.

- Non-alphanumeric characters were stripped after base64 encoding, so an
  unlucky uuid4 lost characters. About 1 mint in 29,000 fell to <= 16
  characters, which the library's own test asserts against. uuid
  ffffffff-ffff-4fff-bfff-ffffffffffff encodes to the 2-character "Tw".

The generator moves to folio/iri.py, standard library only, so minting no
longer requires constructing a FOLIO graph and downloading 18 MB of OWL to
produce one 23-character string.

The 127-bit width is recovered from the published data, not chosen: body
lengths of 20/21/22 occur at 0.39%/25.75%/73.85%, and base62 of 127 bits
predicts 0.39%/25.31%/74.29%. 128 bits predicts 0.21%/12.9%/86.9% and is
excluded. It is pinned by a test.

No existing IRI changes. generate_iri() only mints new ones.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019g78dzGEqYn2iTNzR11Jfv
The new files failed `ruff format --check` and added two `ruff check`
diagnostics (UP035, UP045) that main does not have. CI only runs pytest,
so neither showed up as a failing check.

After this the branch matches main's lint baseline exactly: 188 ruff
errors (all pre-existing), `ruff format --check` clean, `ty check folio/`
at 26 diagnostics.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@mjbommar

Copy link
Copy Markdown
Contributor

Confirmed — dropping the R prefix was incidental, not deliberate. Merging.

I traced the history before answering, since your question was whether there was a decision behind the current shape. There wasn't one that's recorded: generate_iri() arrived fully formed in 9ee7801 (v0.1.0, 2024-09-01) with no preceding commits or design notes, and the only later change is 57a345b, which swapped the soli. domain for folio. and touched nothing else. soli-python isn't a separate history to check — that repo name redirects here.

The one piece of visible intent is the original comment, truncated mid-sentence:

# only use alphanumeric characters from restricted b64 encdoding to

So "alphanumeric only" does look deliberate. Worth noting explicitly in your favor: base62 preserves that intent while removing the post-hoc stripping that causes the short-name bug. This PR isn't reversing the one decision that was actually made here.

I reproduced all four quantitative claims independently against FOLIO.owl, and they hold:

Claim Result
Family distribution reproduced exactly (11,427 / 4,751 / 1,992 / 152 / 2)
127-bit width observed 0.40/25.75/73.85% vs predicted 0.415/25.27/74.31%; 128 bits predicts 0.21/12.6/87.1% and is excluded
Dead collision guard confirmed — iri_to_index is keyed by full IRI (graph.py:592 → :782); the guard at :2498 tests a bare local name and bypasses normalize_iri
~1 in 29,000 short names 1 in 29,684 analytically, 1 in 28,571 over 200k simulated mints

The 127-bit fit depends on excluding the 4,751 length-23 names as SALI hex legacy, so I checked that too: all 4,751 are case-insensitive hex. Since 127 bits caps base62 output at 22 digits, those names provably can't share a generator with the rest — the exclusion is sound rather than convenient.

Two things worth adding, because they make the case better than the majority-share argument does.

First, the real justification for R is XML validity, not convention. WebProtégé issue #18 documents it: a local name starting with a digit forces the RDF/XML serializer to split the namespace mid-token, producing wp:gQAvsFVAjAj7tQR77pJGi from a prefix ending in 8 — "quite yucky," in their words. The current generate_iri() emits a digit-leading local name 16.1% of the time. All 16,330 R-prefixed FOLIO names avoid this; all 311 published digit-leading names come from the unstripped bare-base64url family. That's a correctness argument, and it's the one I'd lead with.

Second, the blast radius is zero, which is why I'm comfortable taking the behaviour change. Family counts are byte-identical between the ontology's initial commit (c16f652, Aug 2024) and today — everything was inherited from SALI. The single class added since (ea7908a, Refund) got RtQgrbKIvgHvOXm_6lIpaY172, i.e. R+base64url, so it was minted by hand or by WebProtégé rather than by this library. Nothing in our other repos calls generate_iri() either. This function has never minted a published IRI, so there is no installed base to be compatible with.

One thing to flag for next time: the new files failed ruff format --check and added two ruff check diagnostics (UP035, UP045) that main doesn't have. CI only runs pytest, so nothing surfaced it. I pushed 8d2b7d7 to fix it — the branch now matches main's baseline exactly (188 pre-existing ruff errors, format clean, ty at 26). Full suite is 79 passed locally, and your folio.iri tests pass under a normal install.

Shipping as 0.4.0.

@mjbommar
mjbommar merged commit 2a113ce into main Aug 18, 2026
5 of 6 checks passed
@mjbommar
mjbommar deleted the iri-minting-r-base62 branch August 18, 2026 01:34
mjbommar added a commit that referenced this pull request Aug 18, 2026
Ships #19. Dates the 0.4.0 changelog entry and bumps pyproject.toml,
folio/__init__.py, and the lockfile to 0.4.0.

Adds one changelog entry that #19 did not claim for itself: minted local
names previously began with a digit about 16% of the time, which is not a
valid XML NCName, so an RDF/XML serializer could not form a QName without
splitting the namespace mid-token (protegeproject/webprotege#18). The R
prefix removes this by construction. That is the strongest argument for
the change and it belongs in the record.

The lockfile diff is a marker normalization from a newer uv plus the
version bump; no dependency versions moved and no hashes changed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants