Skip to content

sync: closest devices first, by measured link speed - #27

Open
ChakraFusion wants to merge 7 commits into
Liquid-co:mainfrom
ChakraFusion:pr/network-aware
Open

ChakraFusion wants to merge 7 commits into
Liquid-co:mainfrom
ChakraFusion:pr/network-aware

Conversation

@ChakraFusion

Copy link
Copy Markdown

Part of #21.

Based on the sync-speed PR (#26), which is itself based on the "Keep mine" PR. Only the last four commits are new here.

Why

With devices on a mix of home network, VPN (Tailscale) and the internet relay, OpenSave synced to devices in no particular order. A new save could go first to the slowest device over the relay, while a PC on the same LAN waited. Separately, block fetches on direct addresses had a fixed 30-second limit. "Direct" includes VPN addresses, so a few MB of blocks at a VPN relay's 100-odd KB/s outlasted it: the request failed, and with it the whole pull, over and over.

What

Measuring links

  • Each device records how fast its connection to every other device is:
    • from the pulls it makes, measured over each 30-second window while a long pull runs, so a pull lasting an hour doesn't leave the link "not measured";
    • from a short speed test (/api/p2p/speedtest, 2 MB) when nothing has measured it, at most every 6 hours.
  • A failed test, for example against a peer on an older build, is retried after 10 minutes instead of counting as done.
  • A test over a fast home network finishes in a fraction of a second. It is then repeated at 8 MB, and it counts however short it is.
  • Before anything is measured, the address kind stands in: home network fast, VPN slower, relay slowest.

Ordering

  • A sync goes to the fastest devices first.
  • Having asked a close device to take a newer save, a sync waits for it to finish (sync-complete) before asking a much slower one, bounded by how long the transfer should take. The save reaches nearby devices at full speed, and slower ones can then take it from whichever device is closest to them.
  • The device list shows each connection's kind and measured speed.

Block-fetch timeout

  • Block fetches on direct addresses get a limit that grows with the amount requested, as relay fetches already did.

Storage

  • A new table, peer_links, in migration 0038_peer_links.sql (the next number after 0037_game_root_mapped).

Compatibility

A peer without the speed-test route simply stays "estimated by address kind" and is retried every 10 minutes. Nothing else changes for older peers.

Tests

🤖 Generated with Claude Code

ChakraFusion and others added 7 commits October 4, 2026 01:29
…te a large part of a save

"Keep mine" (and "Keep both", which ends the same way) recorded every local
file as shared with the peer before the peer had pulled any of it. When that
pull did not finish - a large save, a request that timed out, a device that
went away - the next sync read the files the peer was missing as deletions
made over there, and deleted them on the device whose version had just been
kept. On a save of ~240k files this removed tens of thousands; the folders
then differed again, the conflict came back, and the same answer did it
again.

- markResolvedLocal records as shared only what both sides verifiably hold
  (the intersection of the two manifests). A file only this device has is
  new to the peer, not deleted by it. If the peer cannot be asked, nothing
  is recorded as shared. Same for extra save locations
  (ResolveRootConflict).
- A sync that would delete more than 100 files and more than 10% of a
  save, on either side, is held as a conflict instead of applied. A deletion
  confirmed on purpose (the emptied-folder flow) is not held.
- The conflict carries uncapped counts of what differs (diffFiles stops at
  100), and the dialog says what each answer does: "Keep mine"/"Keep both"
  delete nothing; "Keep theirs" deletes N files only this device has, shown
  in red when that is many. Older builds without the counts fall back to
  counting the capped list.

TestKeepLocal_PeerThatNeverPulled_DeletesNothingHere fails on v2.4.1
(300 files -> 100); either change alone makes it pass.
Also adds tests that an interrupted pull or a peer holding part of a save
never turns into deletions.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…its own

Performance
- Batched pulls (protocol revision 2): small files are fetched up to 256
  per request (4MB), 4 requests in flight, over a new /api/p2p/files
  route. Each file is still written through the PatchWriter (verified,
  fsynced, renamed). One request per file cost ~200ms on a real LAN; on
  loopback 5000 files went from 42s to 12s. Peers below revision 2 and
  relay connections keep the per-file path.
- Unchanged games: when this device still holds the agreed base, the
  manifest request carries ifHash and a peer that still holds it too
  answers "unchanged" instead of sending the manifest. Older peers ignore
  the parameter; their full answer is used, not fetched twice.
- Manifests are gzipped for clients that accept it (every Go client does,
  so older versions read them too): ~5.6x smaller with real hashes.
- Manifests and batches get a 10-minute client timeout instead of 30s; a
  large manifest over a VPN never arrived inside 30s, so those games
  never synced.
- An in-sync pass also mirrors the peer's latest snapshot, which only a
  pull did before; games identical everywhere from the start showed
  history on one device only.
- A batch is bounded by what the responder counts (each requested block at
  the file's full block size, at most the file's real size), so a batch of
  many small files with a few larger ones is never refused as too large; a
  batch refused anyway is fetched again as two halves.

Robustness
- Presence: paired LAN peers are probed every 10s on their own ticker.
  Probes used to ride along with the minute-long reconcile, which waits
  for a whole sync-all pass; peers reached over a VPN (no UDP broadcast)
  showed online/offline minutes late. Probe rounds are serialized.
- Games whose save folder is not on this device are skipped by sync-all
  and dropped from the retry queue instead of failing every 20s.
- "syncing X with Y" is logged only when there is something to do; most
  passes produced only "already in sync" lines, which rotated everything
  else out of the log within hours.
- The interrupted-pull safety test also runs over batched pulls.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Every peer started each run with whatever status the database held, and a
peer counts as offline only after three failed probes. For a PC switched
off overnight that meant being shown online after every start until three
probes - the first ten seconds after the background loops started - had
failed.

- The presence loop probes at once on start, then every 10s.
- The three-miss tolerance is for devices heard from during this run
  (answered a probe, or seen by discovery, a request or the relay since
  start). A peer with no sign of life since start goes offline on its
  first unanswered probe.

Adds an opt-in e2e measurement (OPENSAVE_PRESENCE_TIMING) of how long a
stopped device is shown online: 30s, by design, for one seen this run.
…tches

Network-aware order:
- Each device records how fast its connection to every other one is: from
  the pulls it makes, and from a short speed test (/api/p2p/speedtest, 2 MB,
  at most every 6 hours) when nothing has measured it. Before anything is
  measured, the address kind stands in: home network fast, VPN (Tailscale's
  range) slower, the internet relay slowest.
- A sync goes to the fastest devices first. A device that is behind takes
  the newer save from the fastest device holding it.
- Having asked a close device to take a newer save, a sync waits for it to
  finish (it reports sync-complete) before asking a much slower one, bounded
  by how long the transfer should take. The save reaches the close devices
  at full speed, and the slow ones can then take it from whichever device is
  nearest to them.
- The device list shows each connection's kind and measured speed.

Block fetches on direct addresses had a fixed 30-second limit. "Direct"
includes VPN addresses, and a few MB of blocks at a VPN relay's hundred-odd
KB a second outlasted it: the request failed, and with it the whole pull,
over and over. The limit now grows with what is asked for, as it already did
over the internet relay.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A speed test that failed - a peer not yet on a build that answers one -
counted as done for six hours, so every link stayed unmeasured and the order
fell back to address guesses all that time. A link never measured is now
tested again ten minutes after a failed test.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… end

A pull of a large save over a slow link takes an hour, and the link's speed
was recorded once, when it ended: all that time the device list said "not
measured" and the sync order fell back to address guesses. The link is now
measured over each 30-second window while the pull runs, and the rest when
it ends.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Over a home network the 2 MB speed test is over in a fraction of a second,
and the rule that ignores transfers under half a second - meant for syncs,
whose small transfers say more about round trips than about the link -
threw it away: the fastest link stayed unmeasured, and with nothing measured
it sorted behind the measured slow ones. A test that quick is now run again
at 8 MB, and a test counts however short it is.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@vercel

vercel Bot commented Oct 4, 2026

Copy link
Copy Markdown

@ChakraFusion is attempting to deploy a commit to the sivadaboi's projects Team on Vercel.

A member of the Team first needs to authorize it.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant