Repository navigation
sync: closest devices first, by measured link speed - #27
Open
ChakraFusion wants to merge 7 commits into
Open
ChakraFusion wants to merge 7 commits into
ChakraFusion wants to merge 7 commits into
Conversation
…te a large part of a save "Keep mine" (and "Keep both", which ends the same way) recorded every local file as shared with the peer before the peer had pulled any of it. When that pull did not finish - a large save, a request that timed out, a device that went away - the next sync read the files the peer was missing as deletions made over there, and deleted them on the device whose version had just been kept. On a save of ~240k files this removed tens of thousands; the folders then differed again, the conflict came back, and the same answer did it again. - markResolvedLocal records as shared only what both sides verifiably hold (the intersection of the two manifests). A file only this device has is new to the peer, not deleted by it. If the peer cannot be asked, nothing is recorded as shared. Same for extra save locations (ResolveRootConflict). - A sync that would delete more than 100 files and more than 10% of a save, on either side, is held as a conflict instead of applied. A deletion confirmed on purpose (the emptied-folder flow) is not held. - The conflict carries uncapped counts of what differs (diffFiles stops at 100), and the dialog says what each answer does: "Keep mine"/"Keep both" delete nothing; "Keep theirs" deletes N files only this device has, shown in red when that is many. Older builds without the counts fall back to counting the capped list. TestKeepLocal_PeerThatNeverPulled_DeletesNothingHere fails on v2.4.1 (300 files -> 100); either change alone makes it pass. Also adds tests that an interrupted pull or a peer holding part of a save never turns into deletions. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…its own Performance - Batched pulls (protocol revision 2): small files are fetched up to 256 per request (4MB), 4 requests in flight, over a new /api/p2p/files route. Each file is still written through the PatchWriter (verified, fsynced, renamed). One request per file cost ~200ms on a real LAN; on loopback 5000 files went from 42s to 12s. Peers below revision 2 and relay connections keep the per-file path. - Unchanged games: when this device still holds the agreed base, the manifest request carries ifHash and a peer that still holds it too answers "unchanged" instead of sending the manifest. Older peers ignore the parameter; their full answer is used, not fetched twice. - Manifests are gzipped for clients that accept it (every Go client does, so older versions read them too): ~5.6x smaller with real hashes. - Manifests and batches get a 10-minute client timeout instead of 30s; a large manifest over a VPN never arrived inside 30s, so those games never synced. - An in-sync pass also mirrors the peer's latest snapshot, which only a pull did before; games identical everywhere from the start showed history on one device only. - A batch is bounded by what the responder counts (each requested block at the file's full block size, at most the file's real size), so a batch of many small files with a few larger ones is never refused as too large; a batch refused anyway is fetched again as two halves. Robustness - Presence: paired LAN peers are probed every 10s on their own ticker. Probes used to ride along with the minute-long reconcile, which waits for a whole sync-all pass; peers reached over a VPN (no UDP broadcast) showed online/offline minutes late. Probe rounds are serialized. - Games whose save folder is not on this device are skipped by sync-all and dropped from the retry queue instead of failing every 20s. - "syncing X with Y" is logged only when there is something to do; most passes produced only "already in sync" lines, which rotated everything else out of the log within hours. - The interrupted-pull safety test also runs over batched pulls. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Every peer started each run with whatever status the database held, and a peer counts as offline only after three failed probes. For a PC switched off overnight that meant being shown online after every start until three probes - the first ten seconds after the background loops started - had failed. - The presence loop probes at once on start, then every 10s. - The three-miss tolerance is for devices heard from during this run (answered a probe, or seen by discovery, a request or the relay since start). A peer with no sign of life since start goes offline on its first unanswered probe. Adds an opt-in e2e measurement (OPENSAVE_PRESENCE_TIMING) of how long a stopped device is shown online: 30s, by design, for one seen this run.
…tches Network-aware order: - Each device records how fast its connection to every other one is: from the pulls it makes, and from a short speed test (/api/p2p/speedtest, 2 MB, at most every 6 hours) when nothing has measured it. Before anything is measured, the address kind stands in: home network fast, VPN (Tailscale's range) slower, the internet relay slowest. - A sync goes to the fastest devices first. A device that is behind takes the newer save from the fastest device holding it. - Having asked a close device to take a newer save, a sync waits for it to finish (it reports sync-complete) before asking a much slower one, bounded by how long the transfer should take. The save reaches the close devices at full speed, and the slow ones can then take it from whichever device is nearest to them. - The device list shows each connection's kind and measured speed. Block fetches on direct addresses had a fixed 30-second limit. "Direct" includes VPN addresses, and a few MB of blocks at a VPN relay's hundred-odd KB a second outlasted it: the request failed, and with it the whole pull, over and over. The limit now grows with what is asked for, as it already did over the internet relay. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A speed test that failed - a peer not yet on a build that answers one - counted as done for six hours, so every link stayed unmeasured and the order fell back to address guesses all that time. A link never measured is now tested again ten minutes after a failed test. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… end A pull of a large save over a slow link takes an hour, and the link's speed was recorded once, when it ended: all that time the device list said "not measured" and the sync order fell back to address guesses. The link is now measured over each 30-second window while the pull runs, and the rest when it ends. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Over a home network the 2 MB speed test is over in a fraction of a second, and the rule that ignores transfers under half a second - meant for syncs, whose small transfers say more about round trips than about the link - threw it away: the fastest link stayed unmeasured, and with nothing measured it sorted behind the measured slow ones. A test that quick is now run again at 8 MB, and a test counts however short it is. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
@ChakraFusion is attempting to deploy a commit to the sivadaboi's projects Team on Vercel. A member of the Team first needs to authorize it. |
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Part of #21.
Based on the sync-speed PR (#26), which is itself based on the "Keep mine" PR. Only the last four commits are new here.
Why
With devices on a mix of home network, VPN (Tailscale) and the internet relay, OpenSave synced to devices in no particular order. A new save could go first to the slowest device over the relay, while a PC on the same LAN waited. Separately, block fetches on direct addresses had a fixed 30-second limit. "Direct" includes VPN addresses, so a few MB of blocks at a VPN relay's 100-odd KB/s outlasted it: the request failed, and with it the whole pull, over and over.
What
Measuring links
/api/p2p/speedtest, 2 MB) when nothing has measured it, at most every 6 hours.Ordering
sync-complete) before asking a much slower one, bounded by how long the transfer should take. The save reaches nearby devices at full speed, and slower ones can then take it from whichever device is closest to them.Block-fetch timeout
Storage
peer_links, in migration0038_peer_links.sql(the next number after0037_game_root_mapped).Compatibility
A peer without the speed-test route simply stays "estimated by address kind" and is retried every 10 minutes. Nothing else changes for older peers.
Tests
linkspeed_test.go: address kinds, order by speed, tiny sync transfers ignored, waiting for a close device, a failed test retried soon, a fast link measured.🤖 Generated with Claude Code