Skip to content

fix(wasix): honor caller fsync settings across database lifecycle #249

Description

@f0rr0

Goal and agreed scope

Make caller-selected fsync behavior consistent across the WASIX database lifecycle: preparation, publication, serving, and close/reopen. Use the existing SDK PostgreSQL configuration APIs rather than an additional durability flag or phase-specific overrides.

This is the separate WASIX follow-up agreed when limiting #241 to native scope. #241 remains the native serving-default fix; merged #246 is the private-initdb startup optimization. Coordinate dependency work with #247. Preserve their scope and avoid folding this work back into the startup-only change.

Current behavior and root causes

Reviewed main: 15e4766.

  1. Configuration reaches serving after storage preparation. Rust startup_guc/startup_gucs populate PostgresConfig; TypeScript exposes startupGUCs and forwards it through Node-API. The builder calls prepare_database with a DatabasePlan that does not carry those settings. Preparation therefore cannot consistently honor the caller's selected fsync behavior.
  2. Ordinary flush and explicit synchronization share one hook. SyncHostFile::poll_flush calls host sync_all(). In the consumed WASIX 0.702.1 API, fd_sync uses AsyncWrite::flush and fd_datasync uses a FlushPoller that calls poll_flush; VirtualFile has no distinct synchronization methods. Changing poll_flush to an ordinary flush would also weaken explicit sync requests. Ordinary writes must not be described as universally calling fsync: the inspected fd_write flush branch is conditional on stdio.
  3. Publication is unconditionally durable even with serving fsync disabled. Seed/fresh-cluster publication syncs the staging tree and post-rename parent; root-descriptor publication also syncs its file and directory. Host namespace operations contain additional parent-directory barriers. This is the lifecycle inconsistency explicitly retained by perf(wasix): reduce initdb latency with private initialization #246.

The dependency upgrade alone does not remove the blocker. Inspected published virtual-fs/wasmer-wasix 0.705.0, proposed by #247, retain the same routing: VirtualFile, fd_sync, and fd_datasync.

Proposed direction

  • Resolve effective fsync through the existing validated PostgreSQL configuration path and pass that decision through lifecycle preparation/publication. Preserve PostgreSQL defaults, persisted configuration precedence, and caller overrides; avoid a second partial GUC parser or an independent public durability option.
  • Provide a supported, fallible distinction between ordinary flush and explicit file/data synchronization in the dependency contract. Forward it through relevant filesystem wrappers. Coordinate with feat(wasix): refresh Wasmer engines and browser SDK #247 and use a distributable dependency change; an unpublished local Cargo patch is not a completed SDK fix.
  • Ordinary flush drains pending writes without forcing crash-durability barriers. Explicit sync requests remain real checked operations; caller fsync=off must not silently convert an explicitly requested sync into a no-op.
  • Route initialization, seed installation, root-descriptor publication, applicable restore/publication paths, and filesystem metadata barriers consistently through the resolved lifecycle contract. Keep staging, complete PGDATA/WAL copies, locking, private permissions, rename ordering, cleanup, and commit-state reporting intact. Classify barriers by their purpose rather than mechanically deleting every sync call.
  • Private in-memory bootstrap may remain relaxed because its state is unpublished; perf(wasix): reduce initdb latency with private initialization #246 restores serving settings before publication. Temporary bootstrap settings must not escape into serving.
  • Keep full_page_writes and synchronous_commit independent. fsync=off accepts power-loss corruption risk; normal functional correctness, visible write errors, and safe failure handling still apply. Define any supported configuration-reload behavior so host policy cannot silently diverge from PostgreSQL.

Acceptance and evidence

  • Verify default/on/off, persisted settings, caller overrides, PostgreSQL boolean forms, and invalid-value handling through Rust sync/async and supported Node-API root/direct/worker paths. Scope browser/Postmaster compatibility work to actual dependency consumers; do not assume identical filesystem contracts.
  • Exercise fresh unseeded, seeded, reopened memory/directory databases, standard/ICU resources, queries/transactions, backup/restore, close/reopen, and supported runtime configuration changes.
  • Prove the syscall distinction with operation-level or syscall evidence: ordinary flush versus explicit fd_sync/fd_datasync, plus applicable directory/publication barriers. SHOW fsync and successful reopening alone do not establish that distinction or crash durability.
  • Inject file-write/sync, directory-barrier, rename, and descriptor-publication failures. Preserve retryable pre-publication cleanup and explicit unknown commit state after uncertain publication; never swallow errors or report partial success. Label process-crash tests separately from power-loss evidence.
  • Measure initialization, copy, publication, serving-ready and first-query time separately under both fsync configurations. Include raw and prepopulated storage, cold/warm cache state, and exact SDK/runtime/artifact identities. Compare PGlite using its measured settings and matching timing boundaries rather than using it to select our defaults.
  • Run affected owner checks, clean SDK/package consumers, and required platform qualification for changed host/dependency products. Preserve genuine platform limitations and exact artifact compatibility.

#246 reduced fresh directory startup from roughly 5–6 seconds to a measured 1,058ms median (21 warm standalone Node samples, 927–1,152ms range); seeded directory was 468ms and fresh memory 850ms, each only three samples. Those measurements excluded import/explicit seed reads and used serving fsync, full_page_writes, and synchronous_commit all on. They motivate measuring this follow-up but do not prove sub-second readiness or release qualification. The remaining latency target is separate from completing the fsync contract.

Related: #241, #246, #247; broader audit context #203 and #201.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't workingrustPull requests that update rust code

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions