Repository navigation
sync: decide by save versions, not by guessing from files - #30
Draft
ChakraFusion wants to merge 23 commits into
Draft
ChakraFusion wants to merge 23 commits into
ChakraFusion wants to merge 23 commits into
Conversation
…te a large part of a save "Keep mine" (and "Keep both", which ends the same way) recorded every local file as shared with the peer before the peer had pulled any of it. When that pull did not finish - a large save, a request that timed out, a device that went away - the next sync read the files the peer was missing as deletions made over there, and deleted them on the device whose version had just been kept. On a save of ~240k files this removed tens of thousands; the folders then differed again, the conflict came back, and the same answer did it again. - markResolvedLocal records as shared only what both sides verifiably hold (the intersection of the two manifests). A file only this device has is new to the peer, not deleted by it. If the peer cannot be asked, nothing is recorded as shared. Same for extra save locations (ResolveRootConflict). - A sync that would delete more than 100 files and more than 10% of a save, on either side, is held as a conflict instead of applied. A deletion confirmed on purpose (the emptied-folder flow) is not held. - The conflict carries uncapped counts of what differs (diffFiles stops at 100), and the dialog says what each answer does: "Keep mine"/"Keep both" delete nothing; "Keep theirs" deletes N files only this device has, shown in red when that is many. Older builds without the counts fall back to counting the capped list. TestKeepLocal_PeerThatNeverPulled_DeletesNothingHere fails on v2.4.1 (300 files -> 100); either change alone makes it pass. Also adds tests that an interrupted pull or a peer holding part of a save never turns into deletions. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…its own Performance - Batched pulls (protocol revision 2): small files are fetched up to 256 per request (4MB), 4 requests in flight, over a new /api/p2p/files route. Each file is still written through the PatchWriter (verified, fsynced, renamed). One request per file cost ~200ms on a real LAN; on loopback 5000 files went from 42s to 12s. Peers below revision 2 and relay connections keep the per-file path. - Unchanged games: when this device still holds the agreed base, the manifest request carries ifHash and a peer that still holds it too answers "unchanged" instead of sending the manifest. Older peers ignore the parameter; their full answer is used, not fetched twice. - Manifests are gzipped for clients that accept it (every Go client does, so older versions read them too): ~5.6x smaller with real hashes. - Manifests and batches get a 10-minute client timeout instead of 30s; a large manifest over a VPN never arrived inside 30s, so those games never synced. - An in-sync pass also mirrors the peer's latest snapshot, which only a pull did before; games identical everywhere from the start showed history on one device only. - A batch is bounded by what the responder counts (each requested block at the file's full block size, at most the file's real size), so a batch of many small files with a few larger ones is never refused as too large; a batch refused anyway is fetched again as two halves. Robustness - Presence: paired LAN peers are probed every 10s on their own ticker. Probes used to ride along with the minute-long reconcile, which waits for a whole sync-all pass; peers reached over a VPN (no UDP broadcast) showed online/offline minutes late. Probe rounds are serialized. - Games whose save folder is not on this device are skipped by sync-all and dropped from the retry queue instead of failing every 20s. - "syncing X with Y" is logged only when there is something to do; most passes produced only "already in sync" lines, which rotated everything else out of the log within hours. - The interrupted-pull safety test also runs over batched pulls. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Every peer started each run with whatever status the database held, and a peer counts as offline only after three failed probes. For a PC switched off overnight that meant being shown online after every start until three probes - the first ten seconds after the background loops started - had failed. - The presence loop probes at once on start, then every 10s. - The three-miss tolerance is for devices heard from during this run (answered a probe, or seen by discovery, a request or the relay since start). A peer with no sign of life since start goes offline on its first unanswered probe. Adds an opt-in e2e measurement (OPENSAVE_PRESENCE_TIMING) of how long a stopped device is shown online: 30s, by design, for one seen this run.
…tches Network-aware order: - Each device records how fast its connection to every other one is: from the pulls it makes, and from a short speed test (/api/p2p/speedtest, 2 MB, at most every 6 hours) when nothing has measured it. Before anything is measured, the address kind stands in: home network fast, VPN (Tailscale's range) slower, the internet relay slowest. - A sync goes to the fastest devices first. A device that is behind takes the newer save from the fastest device holding it. - Having asked a close device to take a newer save, a sync waits for it to finish (it reports sync-complete) before asking a much slower one, bounded by how long the transfer should take. The save reaches the close devices at full speed, and the slow ones can then take it from whichever device is nearest to them. - The device list shows each connection's kind and measured speed. Block fetches on direct addresses had a fixed 30-second limit. "Direct" includes VPN addresses, and a few MB of blocks at a VPN relay's hundred-odd KB a second outlasted it: the request failed, and with it the whole pull, over and over. The limit now grows with what is asked for, as it already did over the internet relay. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A speed test that failed - a peer not yet on a build that answers one - counted as done for six hours, so every link stayed unmeasured and the order fell back to address guesses all that time. A link never measured is now tested again ten minutes after a failed test. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… end A pull of a large save over a slow link takes an hour, and the link's speed was recorded once, when it ended: all that time the device list said "not measured" and the sync order fell back to address guesses. The link is now measured over each 30-second window while the pull runs, and the rest when it ends. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Over a home network the 2 MB speed test is over in a fraction of a second, and the rule that ignores transfers under half a second - meant for syncs, whose small transfers say more about round trips than about the link - threw it away: the fastest link stayed unmeasured, and with nothing measured it sorted behind the measured slow ones. A test that quick is now run again at 8 MB, and a test counts however short it is. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
On a save of ~240k small files (Project Zomboid map chunks) a repeat manifest build took 13-19s in 2.4.0; now ~1.3s, with an identical hash. - Walk with filepath.WalkDir instead of filepath.Walk. Walk's extra Lstat per entry opens every file on Windows and dominated a warm build. - Reuse block read buffers (sync.Pool) instead of allocating 64KB+ per file: a cold build allocated ~15GB of garbage, now ~0.6GB. - Sort with sort.Slice when evicting from the hash cache. The insertion sort was quadratic in the entry count (map order is random). - Raise the cache budget floor to 256MB and scale it at 1MB per game, ceiling 512MB. 64MB was filled by one large save on its own, so the cache evicted and re-read that save on every pass. A budget, not an allocation: small libraries never come near it. - Share the whole-file hash string with the block hash for single-block files (most save files). - FileEntryForInfo: a cache lookup with the FileInfo a walk already has.
…compressing them again Every snapshot stays a complete zip that restores on its own. What changes is how it is written: a file whose content (SHA-256, as already recorded per snapshot and held by the hash cache) matches the previous snapshot of the same branch is copied in as its already-compressed entry (zip.Writer.Copy). Only new and changed files are compressed. On a 204k-file save: a full snapshot 80.5s, a follow-up 8.9s, same size. The archive walk also moves to filepath.WalkDir for the same reason as the manifest walk.
…no warning - A missing save folder is looked for at the identical path on every other drive letter (a Steam library on D: on one PC and E: on another). Exactly one match is adopted for this device and logged; none or several are logged once. Searched at most once an hour. - A save folder on a drive this device does not have is shown as "Not on this device" (muted) instead of a warning, kept out of the warning headline and the notification bell. A folder gone from a drive that exists is still a warning. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
OpenSave runs in the background, and with Go's default (let the heap double before collecting) a large save's manifests showed up as twice their size in Task Manager. GOGC 50 and a 1 GiB soft memory limit make the runtime collect and return memory sooner, for a little CPU. GOGC and GOMEMLIMIT set in the environment still win. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
New package internal/owntouch records the paths OpenSave itself writes, removes, creates or re-dates: pulled files, a peer's deletions, the folders a pull creates, the files "keep theirs" removes and the mtimes "keep mine" touches. The watcher classifies every event by it: a burst made only of OpenSave's own changes moves the recorded hash and nothing else - no auto-snapshot, and no OnChanged sending the sync straight back out. Before, every change a sync applied looked exactly like the game saving. A peer deleting files one request at a time produced an auto-snapshot every couple of seconds, each of a save that was neither the old one nor the new one; they filled the retention budget and pushed out the snapshots that mattered. - a file counts as OpenSave's only if nothing wrote it after the mark (its mtime is not later), so a game saving straight after a sync is still the game's change. - once OpenSave is done with a file (renamed into place, re-dated to the peer's time) it records the time and size it left it with, and from then on the file is OpenSave's only while it has exactly those - a game saving within the two seconds the time comparison allows is still the game's. - a path found gone counts as OpenSave's only if OpenSave removed it (MarkRemoved): a file it pulled and the game or a person then deleted is their change. - the rescan after new folders come under watch is OpenSave's own when it created those folders (a pull of a save with subfolders), and someone else's otherwise. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The batched pull re-dates each file to the peer's time like the per-file pull does, and like it marks the file first and settles it after, so the watcher takes a batch for OpenSave's own change. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Every device now records, per game, which version of the save it holds as a version vector: one counter per device, moved only when the save really changes there (never by a sync; see internal/owntouch). The version travels in the manifest response, and two devices compare versions before anything else: - one older than the other: that device is behind. It takes the newer save whole - every file and every deletion - and only from a device holding it completely; only then does it adopt that version. Until then it is a source for nobody: nothing is taken from it, pushed by it, or deleted on its word. A pull that stops part-way is remembered across restarts, so the half-copy is never mistaken for a change made there. - a device remembers the newest version it has heard of (its target). Two devices that both know a newer one exists do nothing with each other until a device holding it is back. This is what stopped a half-finished copy being passed on, with everything it lacked read as deletions. - an empty folder holds no version, so it fills itself and never raises a conflict. - neither newer: both changed independently. Only this asks. Resolution stays mutual: "keep mine" on one device does not make the other's save simply older - the other device is asked in turn, and the first is not asked again (the answer travels with the version). If both keep their own, the later answer stands. Every sync and every manifest served first checks the save against the hash recorded with its version, so a change the watcher has not reported yet (it waits for the game to finish writing) is a version before anyone compares - otherwise a device with unrecorded work would look merely behind and take the other's save over it. Changes to excluded files are not versions, and neither is writing an exclusion rule: the recorded hash is tagged with the rules it was taken under and simply re-taken when they change. A device fetching back a save it was asked to put back after it was emptied (hold.go) follows the hold's rules. Saves from before versions get one named after their content, so devices that already agreed agree at once; ones that did not are compared as before until they agree. Games with automatic sync off are not watched, so they keep the old file comparison; so do peers that predate this. Also: - the mass-deletion guard is removed. It held syncs that would delete many files for a decision - a stopgap for deletions read from a half-finished copy, which versions now rule out at the source. - the sync engine marks what it writes and removes (owntouch), and a burst of deletions a peer asks for takes one safety snapshot, at most every ten minutes, instead of one auto-snapshot per file.
…is save everywhere"
Saves that existed before version tracking get a starting version named
after their content, and nothing recorded says which device's is newer.
Between two of those, the old file-by-file comparison used to decide - and
it either asked, or merged the two file by file into a save no game wrote.
Now, between saves from before versions:
- if the two last agreed on a state and one still holds exactly it, the
other is newer;
- otherwise if one is newer in every file that differs (newer, or missing on
the other side) and the other holds no file written after its newest,
that one is newer;
- the older device takes the newer save whole, never file by file. The side
that gives way was older in every differing file, so it loses nothing
newer than what it takes.
- anything else is asked about.
"Use this save everywhere" (game header, POST /api/games/{id}/use-everywhere)
is for states nobody can rank - half-arrived copies, mixtures. The device's
version then covers every save from before versions (entry "g:*"), so every
such device takes it whole, keeping a snapshot first. A change made on
another device since versions existed is not overruled: that is a conflict.
An open conflict no longer blocks a game for good once a version newer than
both sides of it arrives: it was settled elsewhere, and is cleared.
…amage
The transition to save versions happens one device at a time, and a device
still on an older build decides by the old file comparison: whatever its copy
lacks it reads as deleted, so a copy that arrived half-way spread its gaps as
deletions to the devices that had the whole save. That is how a restored save
lost 6,927 files again to a device not yet updated.
- Nothing is taken from a peer that does not keep save versions (status
peer_outdated_app); it can still take from this device.
- Deletions it asks for are refused, on the LAN and over the relay. Requests
from builds that keep versions say so ("versioned": true).
- The device list marks such a peer "needs update".
- The conflict dialog offers "Use this device's save everywhere" for when
neither side is the right save.
Peers signal that they keep versions with "versions": true in the manifest
answer (and LAN protocol revision 3). The relay path now also carries the
save's version: until now, relay syncs never got one and always fell back
to the old file comparison.
… not asked about Before versions, devices that never ran a game held copies relayed from the others at different times, and syncs file by file had left some holding mixtures. Two such copies, neither newer in every file, were asked about - a question about states nobody made, whose right answer is always the copy with the newer work. Now, after the two existing rules (unchanged since the last agreed state; newer in every differing file), the save holding the most recent work wins: the latest-written of the files where the two differ, then the one newer in more of them. The other takes it whole, never file by file, keeping a snapshot of its own first. Only an exact tie is asked about. Changes made since versions existed are not decided this way: they are versions of their own, and a real conflict between them is still asked.
A save kept as a single file is keyed by its name, so the same save named differently on two devices (an emulator's <profile>.sav) shares no path with the other. The rule settling saves from before versions by their most recent work then made the newer one the save, and the other device would have removed its own file for one with a different name. Saves with no file in common are not two versions of the same files: they are asked about. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
"<peer> has an older version - asking it to take this one" was logged on every pass while the peer was still taking it, every 20 seconds during a long pull. Logged once per state now. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…nished Taking a peer's version checks, once its files are written, that the save is now exactly the peer's; if not, the pull counts as stopped part-way and is resumed on the next sync. A game saving in that moment - after the pull wrote its file, before the check - made the save differ, and the resumed pull then replaced the new save with the peer's. Only the snapshot taken before replacing it still held it, and both devices called themselves in sync. (CI: TestSwitch_TheSameGameUnderDifferentIDsSyncs, about one run in fifteen locally.) When every difference is a file the pull wrote or removed and someone other than OpenSave has written it since (owntouch), the pull did finish: the device takes the peer's version, records the change as a version of its own made from it, and asks the peer to take that. Anything else is still a pull stopped part-way. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…hange here The check after a pull counted only files the pull had written as possibly changed since. A file the pull did not touch - the same on both sides when the pull was planned - deleted here the moment the pull was done made the save differ from the peer's, the pull counted as stopped part-way, and the resumed pull fetched the deleted file back (CI: TestFullSyncFlow, Windows). A file the pull did not touch that differs now was changed here meanwhile: it counts like a file the pull wrote that someone else has written since. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Taking a newer version after following a peer onto a branch read the save back through the gate the same sync still held, and waited for itself until its time ran out (30 minutes). It reads it directly. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The API lists a branch's snapshots oldest first, so the newest - what a branch switch puts back - is the last, not the first its comment meant. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
@ChakraFusion is attempting to deploy a commit to the sivadaboi's projects Team on Vercel. A member of the Team first needs to authorize it. |
This was referenced Oct 4, 2026
Open
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Part of #21.
Design and motivation: see issue #29 (Proposal: save versions).
Stacked on: the "Keep mine", sync-speed, network-aware, large-save-speed, save-on-other-drive, memory-footprint and own-writes PRs. The new commits start at "sync: decide by save versions". I'll rebase as the earlier PRs land.
Commits
POST /api/games/{id}/use-everywhere).TestSwitch_TheSameGameUnderDifferentIDsSyncs, about one run in fifteen;TestFullSyncFlow).Note on the "Keep mine" PR
This PR removes the mass-deletion guard that the "Keep mine" PR adds. The guard was a stopgap against deletions read from half-finished copies, and versions rule those out at the source.
Storage
internal/store/migrations/0039_game_versions.sql: one table,game_versions.Compatibility
"versions": true, LAN protocol revision 3).Tests
version_test.go, 17 scenario tests:changedsince_test.go: the pull races (a save written over a pulled file, a file the pull did not touch deleted right after it, a pull that really stopped part-way).deletion_race_test.go: a deletion requested by an OpenSave without versions is refused (HTTP 409).🤖 Generated with Claude Code