Skip to content

sync: decide by save versions, not by guessing from files - #30

Draft
ChakraFusion wants to merge 23 commits into
Liquid-co:mainfrom
ChakraFusion:pr/save-versions
Draft

ChakraFusion wants to merge 23 commits into
Liquid-co:mainfrom
ChakraFusion:pr/save-versions

Conversation

@ChakraFusion

Copy link
Copy Markdown

Part of #21.

Design and motivation: see issue #29 (Proposal: save versions).
Stacked on: the "Keep mine", sync-speed, network-aware, large-save-speed, save-on-other-drive, memory-footprint and own-writes PRs. The new commits start at "sync: decide by save versions". I'll rebase as the earlier PRs land.

Commits

  1. decide by save versions: version vectors, targets, the whole-save pull, mutual conflict resolution, and unreported changes checked before every comparison. Bursts of peer-requested deletions take one safety snapshot (at most every 10 minutes) instead of one per file.
  2. saves from before versions: settled by the last agreed state or by files that are newer in every respect. Adds "Use this save everywhere" (POST /api/games/{id}/use-everywhere).
  3. devices without versions can't do damage: nothing is taken from them, and their deletion requests are refused. They're marked "needs update", with protocol revision 3 and the version carried over the relay.
  4. most recent work wins between old saves; only an exact tie is asked about.
  5. no file in common → ask, instead of replacing an emulator save that has a different name.
  6. "older version" logged once per state, not every 20 s.
  7. a save made right after a pull is a new version. A save written, or a file deleted, the moment a pull finished was taken for an unfinished pull and undone on the next sync. CI found both races (TestSwitch_TheSameGameUnderDifferentIDsSyncs, about one run in fifteen; TestFullSyncFlow).

Note on the "Keep mine" PR

This PR removes the mass-deletion guard that the "Keep mine" PR adds. The guard was a stopgap against deletions read from half-finished copies, and versions rule those out at the source.

Storage

internal/store/migrations/0039_game_versions.sql: one table, game_versions.

Compatibility

  • Peers announce versions ("versions": true, LAN protocol revision 3).
  • With peers that predate this, nothing is taken from them and their deletions are refused. They can still take from updated devices.
  • Games with automatic sync off keep the old file comparison.

Tests

  • version_test.go, 17 scenario tests:
    • ordering, and "Use this save everywhere" covering old saves without overruling later changes;
    • saves from before versions: clearly newer, most recent work, no common file, exact tie;
    • incomplete devices waiting for the newest one; an empty device taking the save without asking;
    • deletions travelling with the version; independent changes asking each side once, with the later answer winning;
    • an unnoticed change found when devices meet; nothing taken from an old build; pulled files not being a new version.
  • changedsince_test.go: the pull races (a save written over a pulled file, a file the pull did not touch deleted right after it, a pull that really stopped part-way).
  • deletion_race_test.go: a deletion requested by an OpenSave without versions is refused (HTTP 409).
  • CI on the fork (all jobs green): https://github.com/ChakraFusion/OpenSave/actions/runs/37204416636

🤖 Generated with Claude Code

ChakraFusion and others added 23 commits October 4, 2026 01:29
…te a large part of a save

"Keep mine" (and "Keep both", which ends the same way) recorded every local
file as shared with the peer before the peer had pulled any of it. When that
pull did not finish - a large save, a request that timed out, a device that
went away - the next sync read the files the peer was missing as deletions
made over there, and deleted them on the device whose version had just been
kept. On a save of ~240k files this removed tens of thousands; the folders
then differed again, the conflict came back, and the same answer did it
again.

- markResolvedLocal records as shared only what both sides verifiably hold
  (the intersection of the two manifests). A file only this device has is
  new to the peer, not deleted by it. If the peer cannot be asked, nothing
  is recorded as shared. Same for extra save locations
  (ResolveRootConflict).
- A sync that would delete more than 100 files and more than 10% of a
  save, on either side, is held as a conflict instead of applied. A deletion
  confirmed on purpose (the emptied-folder flow) is not held.
- The conflict carries uncapped counts of what differs (diffFiles stops at
  100), and the dialog says what each answer does: "Keep mine"/"Keep both"
  delete nothing; "Keep theirs" deletes N files only this device has, shown
  in red when that is many. Older builds without the counts fall back to
  counting the capped list.

TestKeepLocal_PeerThatNeverPulled_DeletesNothingHere fails on v2.4.1
(300 files -> 100); either change alone makes it pass.
Also adds tests that an interrupted pull or a peer holding part of a save
never turns into deletions.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…its own

Performance
- Batched pulls (protocol revision 2): small files are fetched up to 256
  per request (4MB), 4 requests in flight, over a new /api/p2p/files
  route. Each file is still written through the PatchWriter (verified,
  fsynced, renamed). One request per file cost ~200ms on a real LAN; on
  loopback 5000 files went from 42s to 12s. Peers below revision 2 and
  relay connections keep the per-file path.
- Unchanged games: when this device still holds the agreed base, the
  manifest request carries ifHash and a peer that still holds it too
  answers "unchanged" instead of sending the manifest. Older peers ignore
  the parameter; their full answer is used, not fetched twice.
- Manifests are gzipped for clients that accept it (every Go client does,
  so older versions read them too): ~5.6x smaller with real hashes.
- Manifests and batches get a 10-minute client timeout instead of 30s; a
  large manifest over a VPN never arrived inside 30s, so those games
  never synced.
- An in-sync pass also mirrors the peer's latest snapshot, which only a
  pull did before; games identical everywhere from the start showed
  history on one device only.
- A batch is bounded by what the responder counts (each requested block at
  the file's full block size, at most the file's real size), so a batch of
  many small files with a few larger ones is never refused as too large; a
  batch refused anyway is fetched again as two halves.

Robustness
- Presence: paired LAN peers are probed every 10s on their own ticker.
  Probes used to ride along with the minute-long reconcile, which waits
  for a whole sync-all pass; peers reached over a VPN (no UDP broadcast)
  showed online/offline minutes late. Probe rounds are serialized.
- Games whose save folder is not on this device are skipped by sync-all
  and dropped from the retry queue instead of failing every 20s.
- "syncing X with Y" is logged only when there is something to do; most
  passes produced only "already in sync" lines, which rotated everything
  else out of the log within hours.
- The interrupted-pull safety test also runs over batched pulls.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Every peer started each run with whatever status the database held, and a
peer counts as offline only after three failed probes. For a PC switched
off overnight that meant being shown online after every start until three
probes - the first ten seconds after the background loops started - had
failed.

- The presence loop probes at once on start, then every 10s.
- The three-miss tolerance is for devices heard from during this run
  (answered a probe, or seen by discovery, a request or the relay since
  start). A peer with no sign of life since start goes offline on its
  first unanswered probe.

Adds an opt-in e2e measurement (OPENSAVE_PRESENCE_TIMING) of how long a
stopped device is shown online: 30s, by design, for one seen this run.
…tches

Network-aware order:
- Each device records how fast its connection to every other one is: from
  the pulls it makes, and from a short speed test (/api/p2p/speedtest, 2 MB,
  at most every 6 hours) when nothing has measured it. Before anything is
  measured, the address kind stands in: home network fast, VPN (Tailscale's
  range) slower, the internet relay slowest.
- A sync goes to the fastest devices first. A device that is behind takes
  the newer save from the fastest device holding it.
- Having asked a close device to take a newer save, a sync waits for it to
  finish (it reports sync-complete) before asking a much slower one, bounded
  by how long the transfer should take. The save reaches the close devices
  at full speed, and the slow ones can then take it from whichever device is
  nearest to them.
- The device list shows each connection's kind and measured speed.

Block fetches on direct addresses had a fixed 30-second limit. "Direct"
includes VPN addresses, and a few MB of blocks at a VPN relay's hundred-odd
KB a second outlasted it: the request failed, and with it the whole pull,
over and over. The limit now grows with what is asked for, as it already did
over the internet relay.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A speed test that failed - a peer not yet on a build that answers one -
counted as done for six hours, so every link stayed unmeasured and the order
fell back to address guesses all that time. A link never measured is now
tested again ten minutes after a failed test.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… end

A pull of a large save over a slow link takes an hour, and the link's speed
was recorded once, when it ended: all that time the device list said "not
measured" and the sync order fell back to address guesses. The link is now
measured over each 30-second window while the pull runs, and the rest when
it ends.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Over a home network the 2 MB speed test is over in a fraction of a second,
and the rule that ignores transfers under half a second - meant for syncs,
whose small transfers say more about round trips than about the link -
threw it away: the fastest link stayed unmeasured, and with nothing measured
it sorted behind the measured slow ones. A test that quick is now run again
at 8 MB, and a test counts however short it is.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
On a save of ~240k small files (Project Zomboid map chunks) a repeat
manifest build took 13-19s in 2.4.0; now ~1.3s, with an identical hash.

- Walk with filepath.WalkDir instead of filepath.Walk. Walk's extra
  Lstat per entry opens every file on Windows and dominated a warm build.
- Reuse block read buffers (sync.Pool) instead of allocating 64KB+ per
  file: a cold build allocated ~15GB of garbage, now ~0.6GB.
- Sort with sort.Slice when evicting from the hash cache. The insertion
  sort was quadratic in the entry count (map order is random).
- Raise the cache budget floor to 256MB and scale it at 1MB per game,
  ceiling 512MB. 64MB was filled by one large save on its own, so the
  cache evicted and re-read that save on every pass. A budget, not an
  allocation: small libraries never come near it.
- Share the whole-file hash string with the block hash for single-block
  files (most save files).
- FileEntryForInfo: a cache lookup with the FileInfo a walk already has.
…compressing them again

Every snapshot stays a complete zip that restores on its own. What changes
is how it is written: a file whose content (SHA-256, as already recorded
per snapshot and held by the hash cache) matches the previous snapshot of
the same branch is copied in as its already-compressed entry
(zip.Writer.Copy). Only new and changed files are compressed.

On a 204k-file save: a full snapshot 80.5s, a follow-up 8.9s, same size.

The archive walk also moves to filepath.WalkDir for the same reason as
the manifest walk.
…no warning

- A missing save folder is looked for at the identical path on every
  other drive letter (a Steam library on D: on one PC and E: on another).
  Exactly one match is adopted for this device and logged; none or
  several are logged once. Searched at most once an hour.
- A save folder on a drive this device does not have is shown as "Not
  on this device" (muted) instead of a warning, kept out of the warning
  headline and the notification bell. A folder gone from a drive that
  exists is still a warning.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
OpenSave runs in the background, and with Go's default (let the heap
double before collecting) a large save's manifests showed up as twice
their size in Task Manager. GOGC 50 and a 1 GiB soft memory limit make the
runtime collect and return memory sooner, for a little CPU. GOGC and
GOMEMLIMIT set in the environment still win.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
New package internal/owntouch records the paths OpenSave itself writes,
removes, creates or re-dates: pulled files, a peer's deletions, the folders a
pull creates, the files "keep theirs" removes and the mtimes "keep mine"
touches. The watcher classifies every event by it: a burst made only of
OpenSave's own changes moves the recorded hash and nothing else - no
auto-snapshot, and no OnChanged sending the sync straight back out.

Before, every change a sync applied looked exactly like the game saving. A peer
deleting files one request at a time produced an auto-snapshot every couple of
seconds, each of a save that was neither the old one nor the new one; they
filled the retention budget and pushed out the snapshots that mattered.

- a file counts as OpenSave's only if nothing wrote it after the mark (its
  mtime is not later), so a game saving straight after a sync is still the
  game's change.
- once OpenSave is done with a file (renamed into place, re-dated to the
  peer's time) it records the time and size it left it with, and from then
  on the file is OpenSave's only while it has exactly those - a game saving
  within the two seconds the time comparison allows is still the game's.
- a path found gone counts as OpenSave's only if OpenSave removed it
  (MarkRemoved): a file it pulled and the game or a person then deleted is
  their change.
- the rescan after new folders come under watch is OpenSave's own when it
  created those folders (a pull of a save with subfolders), and someone
  else's otherwise.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The batched pull re-dates each file to the peer's time like the per-file
pull does, and like it marks the file first and settles it after, so the
watcher takes a batch for OpenSave's own change.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Every device now records, per game, which version of the save it holds as a
version vector: one counter per device, moved only when the save really
changes there (never by a sync; see internal/owntouch). The version travels in
the manifest response, and two devices compare versions before anything else:

- one older than the other: that device is behind. It takes the newer save
  whole - every file and every deletion - and only from a device holding it
  completely; only then does it adopt that version. Until then it is a source
  for nobody: nothing is taken from it, pushed by it, or deleted on its word.
  A pull that stops part-way is remembered across restarts, so the half-copy
  is never mistaken for a change made there.
- a device remembers the newest version it has heard of (its target). Two
  devices that both know a newer one exists do nothing with each other until
  a device holding it is back. This is what stopped a half-finished copy being
  passed on, with everything it lacked read as deletions.
- an empty folder holds no version, so it fills itself and never raises a
  conflict.
- neither newer: both changed independently. Only this asks. Resolution stays
  mutual: "keep mine" on one device does not make the other's save simply
  older - the other device is asked in turn, and the first is not asked again
  (the answer travels with the version). If both keep their own, the later
  answer stands.

Every sync and every manifest served first checks the save against the hash
recorded with its version, so a change the watcher has not reported yet (it
waits for the game to finish writing) is a version before anyone compares -
otherwise a device with unrecorded work would look merely behind and take the
other's save over it.

Changes to excluded files are not versions, and neither is writing an
exclusion rule: the recorded hash is tagged with the rules it was taken under
and simply re-taken when they change. A device fetching back a save it was
asked to put back after it was emptied (hold.go) follows the hold's rules.

Saves from before versions get one named after their content, so devices that
already agreed agree at once; ones that did not are compared as before until
they agree. Games with automatic sync off are not watched, so they keep the
old file comparison; so do peers that predate this.

Also:
- the mass-deletion guard is removed. It held syncs that would delete many
  files for a decision - a stopgap for deletions read from a half-finished
  copy, which versions now rule out at the source.
- the sync engine marks what it writes and removes (owntouch), and a burst
  of deletions a peer asks for takes one safety snapshot, at most every ten
  minutes, instead of one auto-snapshot per file.
…is save everywhere"

Saves that existed before version tracking get a starting version named
after their content, and nothing recorded says which device's is newer.
Between two of those, the old file-by-file comparison used to decide - and
it either asked, or merged the two file by file into a save no game wrote.

Now, between saves from before versions:
- if the two last agreed on a state and one still holds exactly it, the
  other is newer;
- otherwise if one is newer in every file that differs (newer, or missing on
  the other side) and the other holds no file written after its newest,
  that one is newer;
- the older device takes the newer save whole, never file by file. The side
  that gives way was older in every differing file, so it loses nothing
  newer than what it takes.
- anything else is asked about.

"Use this save everywhere" (game header, POST /api/games/{id}/use-everywhere)
is for states nobody can rank - half-arrived copies, mixtures. The device's
version then covers every save from before versions (entry "g:*"), so every
such device takes it whole, keeping a snapshot first. A change made on
another device since versions existed is not overruled: that is a conflict.

An open conflict no longer blocks a game for good once a version newer than
both sides of it arrives: it was settled elsewhere, and is cleared.
…amage

The transition to save versions happens one device at a time, and a device
still on an older build decides by the old file comparison: whatever its copy
lacks it reads as deleted, so a copy that arrived half-way spread its gaps as
deletions to the devices that had the whole save. That is how a restored save
lost 6,927 files again to a device not yet updated.

- Nothing is taken from a peer that does not keep save versions (status
  peer_outdated_app); it can still take from this device.
- Deletions it asks for are refused, on the LAN and over the relay. Requests
  from builds that keep versions say so ("versioned": true).
- The device list marks such a peer "needs update".
- The conflict dialog offers "Use this device's save everywhere" for when
  neither side is the right save.

Peers signal that they keep versions with "versions": true in the manifest
answer (and LAN protocol revision 3). The relay path now also carries the
save's version: until now, relay syncs never got one and always fell back
to the old file comparison.
… not asked about

Before versions, devices that never ran a game held copies relayed from the
others at different times, and syncs file by file had left some holding
mixtures. Two such copies, neither newer in every file, were asked about -
a question about states nobody made, whose right answer is always the copy
with the newer work.

Now, after the two existing rules (unchanged since the last agreed state;
newer in every differing file), the save holding the most recent work wins:
the latest-written of the files where the two differ, then the one newer in
more of them. The other takes it whole, never file by file, keeping a
snapshot of its own first. Only an exact tie is asked about.

Changes made since versions existed are not decided this way: they are
versions of their own, and a real conflict between them is still asked.
A save kept as a single file is keyed by its name, so the same save named
differently on two devices (an emulator's <profile>.sav) shares no path with
the other. The rule settling saves from before versions by their most recent
work then made the newer one the save, and the other device would have
removed its own file for one with a different name. Saves with no file in
common are not two versions of the same files: they are asked about.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
"<peer> has an older version - asking it to take this one" was logged on
every pass while the peer was still taking it, every 20 seconds during a
long pull. Logged once per state now.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…nished

Taking a peer's version checks, once its files are written, that the save
is now exactly the peer's; if not, the pull counts as stopped part-way and
is resumed on the next sync. A game saving in that moment - after the pull
wrote its file, before the check - made the save differ, and the resumed
pull then replaced the new save with the peer's. Only the snapshot taken
before replacing it still held it, and both devices called themselves in
sync. (CI: TestSwitch_TheSameGameUnderDifferentIDsSyncs, about one run in
fifteen locally.)

When every difference is a file the pull wrote or removed and someone other
than OpenSave has written it since (owntouch), the pull did finish: the
device takes the peer's version, records the change as a version of its
own made from it, and asks the peer to take that. Anything else is still a
pull stopped part-way.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…hange here

The check after a pull counted only files the pull had written as possibly
changed since. A file the pull did not touch - the same on both sides when
the pull was planned - deleted here the moment the pull was done made the
save differ from the peer's, the pull counted as stopped part-way, and the
resumed pull fetched the deleted file back (CI: TestFullSyncFlow, Windows).

A file the pull did not touch that differs now was changed here meanwhile:
it counts like a file the pull wrote that someone else has written since.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Taking a newer version after following a peer onto a branch read the save
back through the gate the same sync still held, and waited for itself until
its time ran out (30 minutes). It reads it directly.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The API lists a branch's snapshots oldest first, so the newest - what a
branch switch puts back - is the last, not the first its comment meant.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@vercel

vercel Bot commented Oct 4, 2026

Copy link
Copy Markdown

@ChakraFusion is attempting to deploy a commit to the sivadaboi's projects Team on Vercel.

A member of the Team first needs to authorize it.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant