Skip to content

Latest commit

 

History

7 Commits

Folders and files

Repository files navigation

relational

relational is a write-side relational compiler for PostgreSQL. It derives a relational decomposition for ordinary Rust structs and enums, generates its DDL, and lowers one value or a homogeneous collection into one parameterized statement made from data-modifying CTEs.

It is intended for systems with a central ingestion component that receives large amounts of structured, application-defined data and materializes it as ordinary normalized PostgreSQL tables. Once written, the data has no dependency on this crate: consumers can query it with handwritten SQL, SQLx, Diesel, reporting tools, or any other PostgreSQL client.

structured Rust values
        |
        v
relational ingestion
        |
        v
normalized PostgreSQL tables
        |
        +--> handwritten SQL
        +--> Rust SQL crates
        +--> reporting and analytics tools

The leaf-to-root CTE composition is extracted and generalized from the ADT-to-CTE work in lostman/fuel-indexer-ng. This project is a clean, standalone implementation: it has no Fuel, Sway, blockchain, VM, ABI, Prisma, indexer, ORM, or database-dialect dependency.

Where it fits

The crate concentrates complexity at the write boundary. The ingestion service owns normalization, identity, structural deduplication, dependency ordering, conflict behavior, and atomic graph insertion. Readers see a conventional relational schema and remain free to choose their own access patterns and tools.

This model is particularly useful for indexers, archival and provenance systems, compilers, telemetry pipelines, scientific data, normalized API snapshots, and other workloads that ingest nested values centrally but expose the resulting data to many independent consumers. Content identities make repeated immutable values share storage, while natural and generated identities cover entities and distinct occurrences.

relational is deliberately not an ORM, query DSL, database driver, or migration framework. It does not generate a read API or try to own downstream queries. The generated tables, names, identities, and relationships should therefore be treated as a public data contract, with schema and identity changes managed explicitly by the application.

The derived schema is fixed by Rust types at compile time. The crate stores user data that conforms to those application-defined types; it does not create arbitrary schemas invented by users at runtime without a corresponding code-generation and compilation step.

Each generated query is one atomic PostgreSQL statement. Applications ingesting very large graphs should use bounded batches so statement size, parameter count, and planning cost remain controlled. Driver-specific execution and retry policy belong to the surrounding ingestion service.

Quick start

use relational::{NaturalIdConflict, Relational, ToCte, WriteOptions};

#[derive(Relational)]
#[relational(table = "users")]
struct User {
    #[relational(id)]
    user_id: i32,
    name: String,
}

#[derive(Relational)]
#[relational(
    table = "points",
    identity_domain = "geometry.point",
    identity_version = 1
)]
struct Point {
    x: i32,
    y: i32,
}

#[derive(Relational)]
#[relational(table = "shapes", identity_domain = "geometry.shape")]
enum Shape {
    Point(Point),
    Circle { centre: Point, radius: f64 },
    Empty,
}

fn main() -> Result<(), Box<dyn std::error::Error>> {
    let shapes = vec![
        Shape::Point(Point { x: 10, y: 20 }),
        Shape::Circle {
            centre: Point { x: 3, y: 4 },
            radius: 5.0,
        },
        Shape::Empty,
    ];

    let query = shapes.to_cte()?;
    println!("{}", query.sql());
    println!("{:?}", query.params());
    println!("{}", Shape::schema().to_sql()?);

    let user = User { user_id: 7, name: "Ada".to_owned() };
    let query = user.to_cte_with(WriteOptions {
        natural_id_conflict: NaturalIdConflict::Upsert,
    })?;
    println!("User query:\n{}", query.sql());
    println!("User parameters: {:?}", query.params());
    Ok(())
}

Run the complete example with:

cargo run -p relational --example basic

Composition model

Traversal is deterministic and post-order. A symbolic reference always points to a final ready CTE, not merely to a row that happened to be inserted.

nested Rust value
    -> child base/relationship CTEs
    -> child ready CTEs
    -> parent base-row CTE
    -> enum payload and ordered-collection CTEs
    -> parent ready CTE
    -> ordered root result (ordinal, id, ok)

Every scalar becomes a SqlValue parameter. SQL contains PostgreSQL $1, $2, ... placeholders and never interpolates runtime scalar data. Alias, placeholder, and root-row order are stable for deterministic input and a deterministic injected UUID source.

CteQuery::expected_root_count() is the number of input occurrences. The final projection returns one row for each occurrence, including repeated identities:

column meaning
ordinal zero-based input position
id expected root identity
ok whether the root's final ready CTE produced a row

A failed verification is therefore observable as ok = false. An empty root collection produces a typed query returning zero rows.

Identity

Identity is never a hidden universal integer. It is determined before SQL is rendered and has one of three strategies.

Natural identity: entities

Exactly one #[relational(id)] field makes that field the actual PostgreSQL primary key. Version one accepts integer types, String, and uuid::Uuid. There is no surrogate BIGSERIAL column.

Within one input graph, equal natural IDs and equal complete stored content share one table insertion while every root and collection occurrence remains. The same natural ID with different content is a BuildError in every conflict mode, before SQL execution.

Content identity: values

An ID-less type uses a 32-byte BLAKE3 digest as:

"id" BYTEA PRIMARY KEY CHECK (octet_length("id") = 32)

The digest covers the complete stored relational content. Equal values in the same domain and version share a row. A field, variant, option tag, nested identity, collection length, order, or repeated element can change the digest. #[relational(skip)] fields affect neither storage nor identity.

If different canonical byte sequences produce the same digest in one build, construction returns BuildError::DigestCollision; it never silently aliases the values.

Generated identity: occurrences

#[relational(identity = "generated")] allocates a client-side UUID v7 and stores it as UUID PRIMARY KEY. Allocation happens during child-first query construction, before a content-addressed parent is hashed. Thus a parent can include the occurrence identity without waiting for PostgreSQL. Separately traversed equal generated values remain distinct.

The UUID source and content hasher are injectable through hidden testing hooks, which keeps golden and collision tests deterministic.

Canonical hash contract

The project owns a platform-independent binary format; it does not use Rust Hash, DefaultHasher, Debug, JSON, Serde, bincode, memory layout, or compiler discriminants.

Version 1 begins with:

  1. the bytes relational\0;
  2. a big-endian u16 library-format version (1);
  3. a big-endian u64 length and UTF-8 identity-domain bytes;
  4. the big-endian user identity_version (u32, default 1);
  5. the identity-strategy tag; and
  6. a struct or enum tag, with an explicit UTF-8 enum variant name.

Stored fields follow in declaration order. Each has an index and a length-prefixed UTF-8 Rust field name boundary, followed by an explicit shape and scalar type tag. Integers use fixed-width big-endian encodings. Strings are length-prefixed UTF-8; byte arrays are length-prefixed raw bytes. Options use explicit None/Some tags. Collections include presence where optional, their length, and every ordered element. A nested value contributes its domain, identity version, identity-strategy tag, and canonical identity bytes.

f32 and f64 use raw IEEE-754 bits. Consequently -0.0 and 0.0 have different identities, and distinct NaN payloads remain distinct. Golden test vectors pin this format.

The default domain is module_path::TypeName. Set identity_domain to preserve identity across Rust module/type changes, and increment identity_version for an intentional identity migration. SQL table and column names are not part of the canonical format. Adding, removing, or semantically changing stored fields changes content identity. With default domains, moving or renaming a Rust type changes identity too.

Relational mapping

Identifiers are always PostgreSQL-quoted.

Rust shape PostgreSQL representation
scalar field non-null column
Option<scalar> nullable column
nested derived field <field>_id foreign-key column
Option<derived> nullable foreign-key column
Box<T> transparent
#[relational(skip)] absent from DDL, writes, and identity
Vec<u8> BYTEA by default
Vec<T> / [T; N] field ordered join table
Option<Vec<T>> presence column plus ordered join table

Collection tables are named <owner_table>__<field>, contain owner_id, zero-based ordinal, and value or child_id, and enforce UNIQUE (owner_id, ordinal). Order, length, repeated elements, and optional element tags are retained. Use #[relational(collection)] to force a Vec<u8> field to be an ordered collection rather than BYTEA.

Enums use one identity table containing id and variant, plus one normalized payload table per non-unit variant named <enum_table>__<variant>. Payload collections get their own ordered tables. Unit, tuple, and struct variants may contain scalar, optional, nested, and collection payload fields.

Scalar mappings are:

Rust PostgreSQL
bool BOOLEAN
small signed/unsigned integers SMALLINT or INTEGER
i32 INTEGER
i64, u32, isize BIGINT
i128, u64, u128, usize exact NUMERIC via decimal text
f32, f64 REAL, DOUBLE PRECISION
String, &str TEXT
Vec<u8> BYTEA
uuid::Uuid UUID

Schema::to_sql() walks nested dependencies first, deduplicates equal table definitions, and rejects incompatible definitions that claim the same table.

Conflict policies and ready CTEs

NaturalIdConflict::Error is the default. It emits a plain natural-key insert, so an existing database row produces PostgreSQL's ordinary uniqueness error.

Verify uses insert-or-select behavior. It accepts an existing natural entity only if every non-ID base column and every owned collection row is exactly equal. Comparisons are null-safe and reject missing, differing, or extra ordinals. Floating-point comparisons use PostgreSQL's binary send representation so signed zero and NaN payload bits follow the canonical identity contract.

Upsert updates every non-ID base column and makes every owned collection relation exactly match the input, including deleting stale ordinals and clearing a relation for an empty input collection.

Content rows always use equality-aware insert-or-select: ON CONFLICT DO NOTHING RETURNING, followed by resolution of a pre-existing row only when its stored columns and dependent relations match. This avoids a no-op update and does not hide digest collisions or corrupted/incomplete relations. A node's ready CTE is empty until its base row, enum payload, and all collections have been inserted or verified.

Attributes

#[relational(table = "custom_table")]
#[relational(rename = "custom_column")]
#[relational(skip)]
#[relational(id)]
#[relational(collection)]
#[relational(identity = "generated")]
#[relational(identity_domain = "stable.domain", identity_version = 2)]

The derive reports contradictory/misplaced attributes and generated SQL-name collisions at compile time. Cargo dependency aliases are supported.

Version-one boundaries

Version one deliberately rejects composite natural keys, natural-ID enums, maps, sets and unordered containers, cyclic/recursive relational types, Rc/Arc, borrowed relational children, database-generated identities, database dialects other than PostgreSQL, and derived types with type or const generic parameters. Lifetime parameters used by supported &str fields are accepted. Scalar type aliases may require using their underlying built-in type.

Direct self-recursion is diagnosed by the derive. Indirect and type-aliased cycles are detected during schema validation and return SchemaError::RecursiveType before SQL construction.

Schema/identity evolution is an explicit migration concern. Changing content identity can leave old rows unreferenced; this crate generates definitions and writes, not a migration or deletion framework.

Testing

Run the required checks:

cargo fmt --check
cargo clippy --workspace --all-targets --all-features -- -D warnings
cargo test --workspace --all-features

The PostgreSQL integration test is feature-gated and skips when TEST_DATABASE_URL is absent. To run it explicitly in PowerShell:

$env:TEST_DATABASE_URL = "postgresql://postgres:postgres@localhost/relational_test"
cargo test -p relational --features postgres-tests --test postgres

The test creates a uniquely named schema and removes it when it finishes; use a disposable test database account with create/drop privileges.

Benchmarks

An end-to-end PostgreSQL benchmark compares relational with handwritten SQLx bulk inserts and SeaORM bulk ActiveModel inserts over the same normalized schema and nested data. It is isolated from the main workspace so the benchmark can use current SQLx and SeaORM releases without raising the library's Rust 1.85 minimum.

See the PostgreSQL comparison for the workload, fairness notes, database setup, and command to run it.

About

A write-side relational compiler for Rust that materializes structured values as normalized, content-addressed PostgreSQL data using parameterized CTEs

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages