Skip to content

docs(rfcs): draft polymorphic types RFC - #884

Open
ragnorc wants to merge 1 commit into
ModernRelay:mainfrom
ragnorc:rfc/polymorphic-types
Open

ragnorc wants to merge 1 commit into
ModernRelay:mainfrom
ragnorc:rfc/polymorphic-types

Conversation

@ragnorc

@ragnorc ragnorc commented Oct 7, 2026

Copy link
Copy Markdown
Contributor

What & why

Adds a draft RFC for first-class polymorphism: interfaces and inline unions become abstract node types that queries can bind, traverse to and mutate through, and edge endpoints may be interfaces or unions. Today interface/implements parse and persist but nothing after the schema compiler reads them; typed edge alternation left heterogeneous endpoint unions out of scope.

The proposal keeps one concrete type and one table per node. Abstract types are type sets resolved once per invocation and lowered to per-type pieces the engine already runs. Polymorphic edges gain __src_type/__dst_type columns holding the endpoint's StableTypeId, added by a staged Lance Operation::Merge (no data rewrite; add_columns stays forbidden).

Backing issue / RFC

Checklist

  • Change is focused (one logical change)
  • Tests added/updated for behavior changes (or N/A) — N/A, docs only; the RFC lists the test owners to extend
  • Public docs updated if user-facing surface changed (or N/A) — N/A, draft RFC
  • Reviewed against docs/dev/invariants.md — the RFC's Invariants section maps each affected invariant and the reviewed deny-list items

Local verification

  • python3 scripts/check-docs.py — Documentation OK (204 Markdown files checked)
  • bash scripts/check-agents-md.sh — passed
  • typos docs/rfcs/2026-10-07-polymorphic-types.md — clean

Notes for reviewers

  • Evidence: before drafting, 134 claims behind the design were checked against this commit and the pinned Lance 11.0.0, DataFusion 54.0.0 and arrow-json 58.3.0 sources. Probes confirmed that a detached Operation::Merge adds a nullable column without touching data files or indexes, that BTREE and Bitmap indexes serve a UInt64 tag, and that BM25 scores differ across datasets for the same document.
  • Unresolved questions for review: keyed-edge id spelling for polymorphic endpoints, bare-id endpoint resolution, wide whole-node projection, historical abstract bindings, interface-named policy rules.
  • Found during the investigation and tracked separately: a named cross-type edge with a multi-hop bound (e.g. worksAt{1,2}) is accepted and silently runs one hop.

Interfaces and inline unions become abstract node types that queries can bind, traverse to and mutate through, and edge endpoints may be interfaces or unions. Every node keeps one concrete type and table; polymorphic edges gain StableTypeId endpoint tags added by a staged Lance Operation::Merge.

@ragnorc ragnorc left a comment

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Recommendation: keep this RFC in draft and resolve the two contract conflicts below before acceptance. This review covers ee55aa75948c7d911f765685ad2182f9a4205898. The PR changes documentation only. These findings concern the proposed contract, not new runtime failures in this PR.

What this PR means

Today, each graph node belongs to one concrete type. Applications need separate edge families or query variants when the same relationship can involve several types. This RFC lets an interface or union name that group. A query could then traverse one Subject -> Subject relationship across people and organizations.

Each node still lives in its concrete table. The compiler resolves an abstract type into a set of concrete types for one accepted snapshot. Queries read those tables and combine their results. Polymorphic edges store a stable type ID beside each endpoint ID. One graph publication still makes the complete change visible.

This could simplify application schemas and queries. It also makes some formerly single-table work depend on the number of concrete types. The RFC does not implement that behavior yet.

Changes needed before acceptance

  1. Define how endpoint generalization preserves keyed-edge identity. Adding a tag changes the canonical key tuple. The proposed metadata-only migration leaves the old edge ID unchanged. The unresolved choice of tag spelling does not resolve this migration conflict.
  2. Define direction for equal endpoint sets. The rule that rejects a source which fits both ends also rejects the promised Subject -> Subject recursion. The current compiler chooses outgoing traversal for equal concrete endpoint types.

The inline comments give the examples, source references, and required checks.

Tradeoffs and long-term liability

  • Keeping concrete tables and one graph publication limits liability. It reuses the existing storage and visibility model. It avoids a second registry that maps each node to its type.
  • Type sets can remove repeated application edge definitions and query branches. The engine must then support type sets across scans, traversal, mutation, constraints, serialization, and historical reads. That is a substantial maintenance cost, even though this PR only adds prose.
  • A stored type tag makes endpoint identity explicit and avoids searching every possible node table. Nullable tags avoid rewriting existing edge data. In exchange, every consumer must interpret an absent or null tag through the same implicit-type rule. Five similar features should reuse one logical endpoint identity and one type-set resolver. Separate rules for cascade, keys, export, and traversal would multiply liability.
  • For graph workloads with many small concrete tables, opening N tables can dominate an abstract query. Larger scans also need bounded memory and I/O across those tables. The proposed budgets and one edge scan per window fit this workload. They do not establish a latency improvement. Interface-wide uniqueness adds reads across member tables and graph-level conflict validation.
  • Refusing cross-table full-text scores is a sound initial boundary. Lance computes BM25 statistics within each dataset. A later shared search projection would add another derived artifact and its maintenance cost.

The design can reduce application liability, but it increases engine liability. Acceptance should depend on closing the identity and direction rules and keeping the added mechanisms shared. The net addition of 705 documentation lines does not measure the eventual software cost.

Evidence and limits

I checked the repository invariants, testing guide, first-principles guide, and Lance alignment material. I traced the current compiler, keyed mutation and load paths, keyed write matching, test owners, and relevant storage code.

The pinned substrate is Lance 11.0.0, source commit ab6b5bbe46009ed78746b444df8db59a8bc5d842. Its AllNulls path retains existing fragments. Its Merge validation preserves existing field bindings and requires new field IDs above the prior maximum. This supports the tag-column mechanism. It does not change existing logical edge IDs.

I also checked DataFusion 54.0.0 UnionExec and OmniGraph's partition-zero execution. The RFC's schema normalization and requirement to stream all arms have a concrete basis.

Local checks at the reviewed head:

  • All 434 compiler library tests passed. These include the existing outgoing-direction test for a same-type edge.
  • Documentation validation passed for 204 Markdown files. Agent-guide links and diff whitespace checks passed.
  • An additional key-codec unit run did not reach test execution during its dependency rebuild. I stopped that review-owned build. The identity finding rests on code inspection, not a reproduced polymorphic migration. No temporary source edits remain.

GitHub's GQT run passed and records this exact head SHA. These checks validate existing behavior and documentation. They do not validate the proposed polymorphic implementation. I did not reproduce the RFC author's compression, throughput, or detached-merge measurements. No cross-type performance claim is independently established by this review.

| Rename interface (`@rename_from`) | supported | none |
| Add `implements` when the node already declares every inherited property compatibly | supported | none; satisfaction links change |
| Add `implements` with missing nullable properties | supported | today's add-property path for those properties |
| Generalize an endpoint from `C` to an interface or union containing `C` | supported | staged `Operation::Merge` adds the tag column; `implicit_type = C` |

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P2] Define keyed-edge identity before promising metadata-only generalization

This row marks generalization as supported with only a new nullable tag column. But the keyed-edge rule above also adds that tag to the canonical ID tuple. Consider an existing @key(@src, @dst) edge whose ID is ["alice","acme"]. Adding a type component produces a different ID, regardless of whether the component uses a name or a stable type ID.

The current mutation path derives this ID with canonical_key_id. Keyed writes match on that ID. A metadata-only Merge retains the old ID. A later typed upsert therefore cannot match the old row by its new canonical ID. The loader also rejects an explicit ID that differs from its canonical key, so exported old rows need a defined round-trip rule too.

Choose an identity-preserving encoding, define an explicit identity migration, or exclude keyed edges from this supported migration. The unresolved spelling question does not cover existing IDs. Require an existing keyed edge to survive generalization, typed upsert, export/load, and restart with one logical row and the intended stable identity.

Comment on lines +197 to +199
alternation. Direction is still inferred. When a source's set fits both ends of
an edge (for example `Mentions: Named -> Note` with `Note implements Named`),
the traversal is refused unless the other endpoint's declared type decides it.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P2] Exempt equal endpoint sets from the ambiguous-direction refusal

For the next paragraph's RelatedTo: Subject -> Subject, every valid source fits both ends. Any valid destination also fits both ends, so its declared type cannot decide the direction. The stated rule therefore rejects the very recursive traversal that the RFC promises. It would also change ordinary same-type traversal if applied to singleton type sets.

The current resolve_member chooses Direction::Out when the source matches the edge's source type, including equal endpoint types. The existing test_traversal_direction_out asserts that behavior and passed locally.

Define a direction rule for equal endpoint sets, such as retaining the outgoing default, while rejecting genuinely ambiguous overlaps between unequal sets. Include one-hop and recursive tests for equal sets, plus a refusal test for the unequal overlapping case.

@ragnorc
ragnorc marked this pull request as ready for review October 8, 2026 13:56
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant