Skip to content

DualStore: upload VSS writes in order, monitors before the manager - #11

Merged
kaloudis merged 4 commits into
zeusfrom
dual-store-ordered-writes
Oct 3, 2026
Merged

kaloudis merged 4 commits into
zeusfrom
dual-store-ordered-writes

Conversation

@kaloudis

@kaloudis kaloudis commented Oct 2, 2026

Copy link
Copy Markdown

Fixes #10. This is the mobile counterpart of ZeusLN/zeus-vls#273 (lnrod).

The base is rl-0.2.7-security-bump (#8), the next branch ZEUS mobile will ship. That branch only changes Cargo.toml, Cargo.lock and Package.swift, so this PR does not conflict with it. zeus has the same DualStore with small differences; I can open a separate PR for it.

Problem

DualStore spawned one thread per VSS write and remove. VssStore's per-key version check orders writes to the same key, but three gaps remained:

  • Same-key order can flip. The version is assigned on the spawned thread, so two writes of one key could swap.
  • The bulk sync can overwrite a newer write. It reads SQLite before VssStore assigns its version, so a live write landing in between loses to the older blob.
  • Nothing orders monitors before the manager. VSS could end up holding a manager ahead of a monitor it treats as persisted.

LDK refuses to restore that last state (DecodeError::DangerousValue). This affects every ZEUS mobile user with VSS backup.

Change (src/io/dual_store.rs)

One worker thread drains a set of dirty keys. It reads the current SQLite value when it uploads, so the newest value wins and superseded writes are never sent.

Order within each round:

  1. Read the manager before taking the batch.
  2. Upload monitor keys first: full monitors, MonitorUpdatingPersister updates and archived monitors.
  3. Remove keys that were deleted locally.
  4. Upload the manager read in step 1, only if every monitor key in the batch went through.

Removes of monitor keys wait until the round's monitor writes have succeeded. This matters because MonitorUpdatingPersister consolidation writes a full monitor and then deletes the update keys it replaces, so VSS must never lose an update key before it holds the new full monitor.

Retries: failed keys stay dirty and are retried with backoff from 1 s to 60 s.

Push safety gate: it now runs in the worker before each round. Disabled drops the batch; Undetermined keeps the keys dirty and retries.

Startup bulk sync: it marks every local key dirty on the same worker. It is still skipped in restore mode, and the Background bulk sync complete — n/n keys log line is kept.

Shutdown: dropping the store waits up to 5 s for pending uploads. This is shorter than lnrod's 10 s because the drop can happen on an app thread. Anything left over goes out with the next startup sync.

The read, list and restore paths are unchanged.

Tests

11 new unit tests drive the worker over in-memory stores, wired the same way DualStore wires it:

Test What it checks
newer_write_wins_when_an_older_upload_is_slow A slow older upload never lands after a newer write.
writes_queued_behind_an_upload_collapse_to_the_newest Writes queued behind an in-flight upload are sent once, as the newest value.
manager_waits_for_failed_monitor_upload The manager stays off VSS while a monitor upload fails.
manager_is_uploaded_after_monitor_keys_in_a_round Within a round the manager goes after monitors and update keys.
update_keys_are_removed_only_after_the_full_monitor_is_on_vss Consolidation never deletes an update key before its full monitor is on VSS.
remove_reaches_vss A local remove deletes the key from VSS.
startup_sync_uploads_existing_keys_in_order The startup sync uploads existing keys, monitors before the manager.
poisoned_local_store_uploads_nothing The push gate in Disabled still blocks every upload.
undetermined_gate_keeps_keys_until_vss_answers An unreachable VSS keeps keys dirty until it answers.
shutdown_flushes_dirty_keys A clean shutdown uploads what is still dirty.
shutdown_does_not_wait_forever_on_a_failing_vss Shutdown is bounded when VSS keeps failing.
  • All 34 lib tests pass, and the new ones passed 25 runs out of 25. Clippy reports nothing in dual_store.rs.
  • The old code has no worker to drive, so these tests can't run against it. The same scenarios fail on lnrod's old DualStore (ZeusLN/zeus-vls#273).
  • vss-integration.yml runs integration_tests_vss.rs against a real VSS server on this PR. Those tests build nodes through DualStore.

@kaloudis

kaloudis commented Oct 2, 2026

Copy link
Copy Markdown
Author

VSS integration tests, run locally because CI's vss-integration.yml can't run them on this base (cd vss-server/rust: No such file or directory, the same failure as on #8; fixed by #9):

  • vss-server 0a638f7 (lightningdevkit/vss-server HEAD), built with --cfg noop_authorizer --no-default-features, against Postgres 16, the same setup as Fix fork CI #9's workflow.
  • RUSTFLAGS="--cfg vss_test" cargo test --test integration_tests_vss in rust:1.96.1-bookworm (linux/amd64), at be4bf89:
    • channel_full_cycle_with_vss_store ok
    • vss_node_restart ok
    • vss_v0_schema_backwards_compatibility ok
    • 3 passed, 0 failed. vss-server logged 480 INSERT INTO statements during the run.

Other red checks:

  • The rustfmt and Documentation failures are the same as on Re-pin rust-lightning to v0.2.7 + Zeus patches #8. Only dual_store.rs is touched here. It now passes stable rustfmt --check, and the only rustdoc errors left in it are the two existing vss_push_verdict links, which Fix fork CI #9 removes.
  • df016e1 only changes doc comments.

@kaloudis
kaloudis force-pushed the rl-0.2.7-security-bump branch from 5f4194f to 6e0f337 Compare October 3, 2026 00:20
DualStore spawned one thread per VSS write and remove. VssStore's
per-key version check orders same-key writes, but the version is
assigned on the spawned thread, so two writes of one key could still
swap; the bulk sync read SQLite before taking its version, so it could
overwrite a newer live write; and nothing kept the manager behind its
monitors. A VSS copy holding a manager ahead of a monitor it treats as
persisted cannot be restored: LDK returns DecodeError::DangerousValue.

One worker thread now drains a set of dirty keys, reading the current
SQLite value at upload time. It reads the manager before taking each
batch, uploads monitor keys (full monitors, MonitorUpdatingPersister
updates, archived monitors) first, then removes keys deleted locally,
and uploads that manager last, only if every monitor key in the batch
went through. Monitor-key removes wait for the batch's monitor writes,
so an update key is deleted from VSS only once the full monitor that
supersedes it is there. Failed keys stay dirty and retry with backoff
(1 s to 60 s).

The push safety gate moves into the worker: Disabled drops the batch,
Undetermined keeps the keys dirty. The startup bulk sync marks every
local key dirty on the same worker (still skipped in restore mode).
Dropping the store waits up to 5 s for pending uploads.

Fixes #10.
@kaloudis
kaloudis force-pushed the dual-store-ordered-writes branch from df016e1 to 0787125 Compare October 3, 2026 00:24
Base automatically changed from rl-0.2.7-security-bump to zeus October 3, 2026 01:21
…t read

The worker read the manager, then took the dirty set and dropped the
manager's mark with it, so a manager write between the two was never
uploaded (found in review of ZeusLN/zeus-vls#273, which this ports).
The worker now clears the manager's mark before reading it and leaves
any later mark in place. If the push verdict is undetermined, the
consumed mark is put back.

Two regression tests: a manager write paused between the snapshot read
and the batch reaches VSS by shutdown (fails before this change with
v1 on VSS), and the manager mark survives an undetermined verdict.
@kaloudis
kaloudis merged commit 0426002 into zeus Oct 3, 2026
42 of 46 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant