Skip to content

Faster manifests and snapshots for saves of many files - #25

Open
ChakraFusion wants to merge 2 commits into
Liquid-co:mainfrom
ChakraFusion:pr/large-save-speed
Open

ChakraFusion wants to merge 2 commits into
Liquid-co:mainfrom
ChakraFusion:pr/large-save-speed

Conversation

@ChakraFusion

Copy link
Copy Markdown

Part of #21.

Measured on Project Zomboid (~240k small map-chunk files).

1. delta: hash cache and manifest walk that scale

A repeat manifest build went from 13–19 s to ~1.3 s, with an identical hash.

  • filepath.WalkDir instead of filepath.Walk. Walk's extra Lstat per entry opens every file on Windows, and that dominated a warm build.
  • Block read buffers are reused (sync.Pool). A cold build allocated ~15 GB of garbage; it now allocates ~0.6 GB.
  • Eviction from the hash cache sorts with sort.Slice. The old insertion sort was quadratic in the entry count.
  • The cache budget floor rises to 256 MB, scaling at 1 MB per game up to a ceiling of 512 MB. With 64 MB, one large save filled the cache on its own, so that save was evicted and re-read on every pass. It's a budget, not an allocation: small libraries never come near it.
  • Single-block files (most save files) share the whole-file hash string with the block hash.

2. snapshot: copy unchanged files from the previous snapshot instead of compressing them again

Every snapshot is still a complete zip that restores on its own. Only the way it is written changes. A file whose SHA-256 matches its entry in the branch's previous snapshot is copied in as its already-compressed entry (zip.Writer.Copy), so only new and changed files are compressed.

On a 204k-file save, a full snapshot takes 80.5 s and a follow-up 8.9 s, with the same size.

Tests

  • reuse_test.go: a follow-up snapshot reuses unchanged entries and restores byte-identical.
  • reuse_bench_test.go: an opt-in benchmark.
  • go test ./... passes. CI on the fork: https://github.com/ChakraFusion/OpenSave/actions/runs/37162381473 (all green; one e2e shard needed a re-run for TestJourney_DeviceSpecificConfigNeverTravels, which also fails intermittently on plain v2.4.1)

🤖 Generated with Claude Code

On a save of ~240k small files (Project Zomboid map chunks) a repeat
manifest build took 13-19s in 2.4.0; now ~1.3s, with an identical hash.

- Walk with filepath.WalkDir instead of filepath.Walk. Walk's extra
  Lstat per entry opens every file on Windows and dominated a warm build.
- Reuse block read buffers (sync.Pool) instead of allocating 64KB+ per
  file: a cold build allocated ~15GB of garbage, now ~0.6GB.
- Sort with sort.Slice when evicting from the hash cache. The insertion
  sort was quadratic in the entry count (map order is random).
- Raise the cache budget floor to 256MB and scale it at 1MB per game,
  ceiling 512MB. 64MB was filled by one large save on its own, so the
  cache evicted and re-read that save on every pass. A budget, not an
  allocation: small libraries never come near it.
- Share the whole-file hash string with the block hash for single-block
  files (most save files).
- FileEntryForInfo: a cache lookup with the FileInfo a walk already has.
…compressing them again

Every snapshot stays a complete zip that restores on its own. What changes
is how it is written: a file whose content (SHA-256, as already recorded
per snapshot and held by the hash cache) matches the previous snapshot of
the same branch is copied in as its already-compressed entry
(zip.Writer.Copy). Only new and changed files are compressed.

On a 204k-file save: a full snapshot 80.5s, a follow-up 8.9s, same size.

The archive walk also moves to filepath.WalkDir for the same reason as
the manifest walk.
@vercel

vercel Bot commented Oct 4, 2026

Copy link
Copy Markdown

@ChakraFusion is attempting to deploy a commit to the sivadaboi's projects Team on Vercel.

A member of the Team first needs to authorize it.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant