Skip to content

docs: identifier durability — recorded cases and the machine-readable remedy - #1

Open
damienriehl wants to merge 1 commit into
CatholicOS:mainfrom
damienriehl:docs/identifier-durability
Open

docs: identifier durability — recorded cases and the machine-readable remedy#1
damienriehl wants to merge 1 commit into
CatholicOS:mainfrom
damienriehl:docs/identifier-durability

Conversation

@damienriehl

Copy link
Copy Markdown

Adds an Identifier durability section to docs/schema-proposal.md plus a new open question: the recorded cases (89 IDs at odds with the scheme's own strip rule across ten languages, the circ:it-opus-deicirc:int-opus-dei rename, the Xinjiang ordinals) and the machine-readable remedy with the current slugs as permanent aliases.

No existing identifier is changed, deprecated, or renamed by this PR. All argument lives in the position paper — this PR only demonstrates this repository's recorded cases and the remedied form: CatholicOS/foundation-docs#58 (discussion: CatholicOS/foundation-docs#57).

🤖 Generated with Claude Code

https://claude.ai/code/session_01MoZuH8vKiwL2rsX3a1zSqc

…readable remedy

Cites the position paper 'Identifier Durability: Machine-Readable Canonical IRIs'.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MoZuH8vKiwL2rsX3a1zSqc
@coderabbitai

coderabbitai Bot commented Jul 26, 2026

Copy link
Copy Markdown

Warning

Review limit reached

@damienriehl, you've reached your PR review limit, so we couldn't start this review.

Next review available in: 59 minutes

Enable usage-based reviews in Billing to review now. Otherwise, wait until the next included review is available.
You're only billed for reviews past your plan's rate limits ($0.25/file).

How can I continue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews.

How do review limits work?

CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability.

For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro Plus

Run ID: cb622599-37d0-43eb-a7f6-122a3f1d3aa9

📥 Commits

Reviewing files that changed from the base of the PR and between 5067fb8 and 0096bdb.

📒 Files selected for processing (1)
  • docs/schema-proposal.md
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@JohnRDOrazio

Copy link
Copy Markdown
Member

Thanks for taking the time to assemble the cases — the underlying question (opaque canonical IDs with permanent human-readable aliases) is a serious one and I want it on the record as an open question. But the three cases as recorded don't support the weight put on them, and one of them rests on a misreading of what the first segment is. Taking them in turn.

The 89 figure joins two unlike cohorts

The 58 are real, and they're now fixed. scripts/generate_seed.py stripped type words with ^(arch)?diocesi di |^(arch)?diocese of . Italian forms the archdiocese as arcidiocesi, not archdiocesi, so the alternation missed every Italian archdiocese and exactly 58 slugs kept the word. That is a one-character defect in a generator, not a property of the identifier scheme. Corrected and regenerated: circ:it-arcidiocesi-di-acerenzacirc:it-acerenza, 58 IDs changed, no collisions, 2,935 still unique.

The 31 ordinariates are not violations of rule 1. Rule 5 already states that military ordinariates are national by nature and names one in exactly this form: circ:it-ordinariato-militare. A military ordinariate has no see. Its territory is a body of armed forces, not a place, so there is no place name underneath the type word to strip down to. Deutsches Militärordinariat is not the type word Militärordinariat decorating an identity — it is the identity. Removing it leaves circ:de-.

So the strip rule is broken in one language by one regex, and observed 31 times in the one cohort where the scheme deliberately keeps the word. Sweeping the corrected registry for surviving type words now returns 8 entries, all Italian territorial prelatures and abbacies (circ:it-abbazia-territoriale-di-montecassino, circ:it-prelatura-territoriale-di-loreto). Whether those should reduce to circ:it-montecassino / circ:it-loreto is a fair question and I'd take a PR raising it — but it's a naming decision about eight entries, not evidence of identifier drift.

The first segment is a nation, not a language

"31 ordinariates carry theirs in ten languages" treats the variation across circ:de-, circ:pl-, circ:lt-, circ:ba- as if language were a free variable that could shift under a stable identifier. It isn't. A circumscription exists within a nation — that is what makes it a circumscription — so the nation segment is the most stable element in the whole identifier, and the vernacular of the see name is downstream of it rather than independent of it. circ:pl-ordynariat-polowy-wojska-polskiego is not in Polish by an editorial choice that some later revision might reverse; it is in Polish because the ordinariate is Poland's.

There is a real inconsistency in that cohort, but it is a different one: the upstream index styles some ordinariates in the vernacular and others in English (circ:au-military-ordinariate-of-australia beside circ:de-deutsches-militarordinariat). That is the name-authority question already recorded as open question 4, and it's a data normalization problem, not a durability problem.

Opus Dei

Two things about circ:it-opus-deicirc:int-opus-dei.

First, the cause. The rename was not the identifier scheme failing to anticipate a supranational structure — rule 5 anticipated it, which is why int existed to move it to. It was the upstream source, world_dioceses.json, listing the personal prelature under Italy and mislabelling the row "Diocesi di Lanusei". The repository already records this in three places: the entry's own note field, docs/schema-proposal.md § "Seed and its limits", and the upstream fix at Liturgical-Calendar/LiturgicalCalendarAPI#718. A wrong value corrected to a right one is the registry working, and no identifier scheme immunizes against bad input — under the proposed shape the same correction still has to happen, it just lands in nation while a permanently resolvable alias continues to assert that Opus Dei is Italian.

Second, "the identifier that was supposed to be stable." Stability is a promise made to downstream consumers at publication. This registry is at commit two of an unreviewed proposal — not peer-reviewed, not published, no consumers, explicitly labelled draft in the seed's own $comment ("All IDs are drafts pending committee review"). Nothing has been promised to anyone yet, and the period before first publication is precisely when identifiers are supposed to be revisable. Treating a pre-publication draft's corrections as breaches of a stability guarantee sets a standard under which no registry could ever be drafted at all.

Xinjiang

Agreed that -1/-2 are arbitrary; that's why both the data note and the proposal flag them as pending proper disambiguation rather than presenting them as settled. Worth noting the proposed remedy doesn't dissolve this one. Two sees still need to be told apart in the human-readable layer, so under opaque canonical IDs the disambiguation moves into the alias rather than disappearing — and the aliases are the part downstream consumers actually read and write.

Where this leaves the PR

I'd like to keep open question 6 and take the position paper seriously on its merits — opaque canonical identifiers with guaranteed multilingual labels is a real design tradition (DOI, ORCID, Wikidata, LEI) and deserves a committee decision rather than a default.

What I'd ask is that the "Identifier durability" section be reworked before it goes in, because as written the record it establishes is inaccurate: the 89 should separate the 58 generator defect (now fixed) from the 31 rule-5 ordinariates, the language framing of the nation segment should come out, and the Opus Dei case should reflect that its cause was an upstream data error in a pre-publication draft. An argument this good doesn't need the case count.

JohnRDOrazio added a commit that referenced this pull request Aug 7, 2026
The slugify() strip rule used `^(arch)?diocesi di |^(arch)?diocese of `.
Italian forms the archdiocese as "arcidiocesi", not "archdiocesi", so the
optional `arch` prefix never matched the Italian styled form and all 58
Italian archdioceses kept the type word in their slug, contrary to rule 1
of the schema proposal ("the type is an attribute, not part of the
identity").

Split the alternation so each language keeps its own form and regenerate
the seed. 58 IDs change, e.g. circ:it-arcidiocesi-di-acerenza ->
circ:it-acerenza. No collisions: 2,935 entries, 2,935 unique IDs.

Sweeping the regenerated seed for surviving type words leaves 8 Italian
territorial prelatures and abbacies (circ:it-abbazia-territoriale-di-
montecassino, circ:it-prelatura-territoriale-di-loreto). Whether those
reduce to the bare see name is a separate naming question for the
committee and is left unchanged here.

All IDs remain drafts pending committee review (#1).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants