Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
43 commits
Select commit Hold shift + click to select a range
c06ae14
traces: fetch MSR Cambridge from the cacheMon mirror of SNIA's archives
sshaplygin Sep 13, 2026
91ce8a0
bench: gate the trace loaders and LRU against libCacheSim
sshaplygin Sep 13, 2026
0df6046
bench: replay the trace table on request-counted epochs, repeated
sshaplygin Sep 14, 2026
7b090d8
bench: retain the calibrated twelve-trace baseline and run provenance
sshaplygin Sep 27, 2026
0755f47
docs: align evidence and project claims with the repeated trace baseline
sshaplygin Sep 27, 2026
197f803
docs: correct sampling limits and preserve evidence links
sshaplygin Sep 27, 2026
8a5589a
test: expose release checks that accept uninstallable modules
sshaplygin Sep 27, 2026
7a24536
release: prepare isolated consumer checks and the v0.4 module graph
sshaplygin Sep 27, 2026
ef10f66
release: require candidate licenses to be tracked
sshaplygin Sep 27, 2026
44a3cc0
test(F4,F15,F17): pin release identity and inventory failures
sshaplygin Sep 27, 2026
e7f78db
fix(F4,F15-F18,F25,F27-F28): validate committed release candidates
sshaplygin Sep 27, 2026
06e64c0
test(F3,F26): reject corrupt cached and partial trace downloads
sshaplygin Sep 27, 2026
e569fa1
test(F8,F24): cover reference omissions and empty evidence summaries
sshaplygin Sep 28, 2026
3331d5f
fix(F1,F3,F8-F14,F19,F23-F24,F26,F29): collect unbiased reproducible …
sshaplygin Sep 28, 2026
c38a756
test(F19): reject empty evidence manifests
sshaplygin Sep 28, 2026
29ea6d9
test(F19): reject incompatible batches and untracked measurement sources
sshaplygin Sep 28, 2026
fbded7b
fix(F19): verify complete compatible measurement batches
sshaplygin Sep 28, 2026
37bd850
test(F19): reject ignored Go measurement inputs
sshaplygin Sep 28, 2026
886844d
fix(F19): reject ignored package sources during measurement
sshaplygin Sep 28, 2026
29a831f
docs(F2,F5-F7,F9-F14,F20-F22,F29): publish complete current evidence
sshaplygin Sep 28, 2026
e174263
fix(F4,R-G1,R-G2): validate release candidates from one committed sna…
sshaplygin Sep 28, 2026
03cc448
test(R-G2): verify HEAD release version during working-tree checks
sshaplygin Sep 28, 2026
4e53b67
fix(R-G5,R-G8,R-G12,R-G13): record evidence in isolated committed che…
sshaplygin Sep 28, 2026
844f9bc
fix(R-G8,R-G12): isolate tool environment and canonicalize measuremen…
sshaplygin Sep 28, 2026
541047d
test(R-G8): complete workspace isolation fixture environment
sshaplygin Sep 28, 2026
5447271
fix(R-G12): prevent ambient Make configuration from injecting measure…
sshaplygin Sep 28, 2026
15a6562
test(R-G12): make external Makefile injection regression discriminate
sshaplygin Sep 28, 2026
46c2f5f
fix(R-G3,R-G7,Q-G4): align trace contract and clean failed downloads
sshaplygin Sep 28, 2026
162db3d
test(Q-G4): reject fabricated download caller locations
sshaplygin Sep 28, 2026
00bdcb1
fix(R-G6,R-G9,R-G10,R-G11,Q-G2): retain measured configuration and pr…
sshaplygin Sep 28, 2026
0c95edd
fix(R-G4,Q-G1,C-G1,C-G3): verify generated evidence and scope public …
sshaplygin Sep 28, 2026
a10ebc8
docs(C-G2): scope shadow memory and advisory rates to their measurements
sshaplygin Sep 28, 2026
0fa5595
test(R-G4): require byte-exact generated report verification
sshaplygin Sep 28, 2026
7f2590c
fix(R-G4): compare generated report bytes without newline normalization
sshaplygin Sep 28, 2026
4a66d34
test(R-G4): pin committed report generator under hidden working changes
sshaplygin Sep 28, 2026
1f99046
fix(R-G4): refresh reports with an isolated committed generator
sshaplygin Sep 28, 2026
3ffda19
bench: replace current evidence with three verified committed-snapsho…
sshaplygin Sep 29, 2026
8494d33
test(F16): cover routine tidy with ignored clones and unpublished sib…
sshaplygin Sep 29, 2026
2a84c13
fix(F16): keep routine tidy inventory scoped to tracked modules
sshaplygin Sep 29, 2026
7846cb0
fix(V1): validate Ristretto values without requiring dropped writes
sshaplygin Sep 29, 2026
27cdd18
docs(V1): distinguish queued Ristretto writes from later admission
sshaplygin Sep 29, 2026
1cb65fe
fix(V2): retain synthetic comparisons without an adaptive performance…
sshaplygin Sep 29, 2026
c4765ca
bench: replace current evidence with runs measured at 1cb65fe
sshaplygin Oct 7, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
11 changes: 3 additions & 8 deletions .github/workflows/ci.yml
Original file line number Diff line number Diff line change
Expand Up @@ -67,12 +67,8 @@ jobs:
go run . -strategy=$strategy -epoch=1 -smoke
done

# Two things block a release and show up in no other check: a module missing
# its licence, and a sibling required at a placeholder version. The second is
# invisible locally, because a replace directive resolves it - and replace
# directives are ignored by anyone consuming the module, so the first person
# to find out is a stranger whose `go get` fails. Neither build, test nor
# lint can see either one.
# Build tracked candidate sources through a temporary module proxy with
# GOWORK=off and fresh consumer caches. External dependencies need network.
release-check:
name: release-check
runs-on: ubuntu-latest
Expand All @@ -82,5 +78,4 @@ jobs:
with:
go-version: "1.25"
cache-dependency-path: go.sum
# Reads each go.mod through `go mod edit -json`, so no network is needed.
- run: ./scripts/release-check.sh
- run: make python-check script-test evidence-check release-check
7 changes: 7 additions & 0 deletions .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -12,6 +12,9 @@ traces/
cluster0*
.DS_Store

# libCacheSim, built by scripts/verify-ref.sh at a pinned commit
.tools/

# Working material that is not part of the library. These files stay on
# disk -- they are the working notes and the article's exhibits -- but the
# repository does not ship them. Documentation a reader or an agent should
Expand All @@ -22,3 +25,7 @@ codemaps/
.reports/
.vscode/
metrics/zzz_refute_probe_test.go

# Python release-check test bytecode and local Go workspace checksums
__pycache__/
go.work.sum
182 changes: 68 additions & 114 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -4,123 +4,77 @@ All notable changes to this project are documented here. The format follows
[Keep a Changelog](https://keepachangelog.com/en/1.1.0/), and from v0.1.0 the
project follows [semantic versioning](https://semver.org/spec/v2.0.0.html).

## [Unreleased]
## [Unreleased] — v0.4.0 candidate

### Removed

- The rendered social-preview image (`site/og.png`) and image-card metadata.
The site uses a summary card while evidence materials show the current version.

- **Breaking API change:** removed the distributed bandit and the
`bandit/redis` module. This includes `Distributed`/`NewDistributed`, `Config`,
`Mode`, `EvidenceMode`, `MemStore`/`NewMemStore`, `Store`, `Bucket`, `Role`,
`ArmKey`, `ArmCounts`, `SyncRequest`, `SyncResult`, `WindowCounts`,
`ArmEvidence`, `Snapshot`, and their coordination constants and errors:
`ModeLeader`, `ModeSharedPosterior`, `EvidenceAll`, `EvidenceShadowOnly`,
`RoleShadow`, `RoleActive`, `DefaultWindow`, `DefaultDecay`,
`DefaultLocalDiscount`, `DefaultJitter`, `DefaultMaxEvidence`, `ErrNilStore`,
`ErrInvalidCoordinationEpoch`, `ErrInvalidWindow`, `ErrInvalidDecay`,
`ErrInvalidJitter`, `ErrEmptyNamespace`, `ErrShadowOnlyUnderLeader`.
Redis `Options`, `Store` and `New` are removed with that module. Local
`Thompson` and `Greedy` remain. No replacement fleet package is published by
this release; users of the old API must retain the v0.3.1 module family.
- Fleet benchmarks, Redis examples and the unused stitchfix bandit dependency.
Examples now use the repository's native bandits.

### Changed

- Bandit callbacks run outside the cache mutex, remain serialized separately,
and stale epoch results are discarded. A slow callback no longer holds the
cache lock; concurrent traffic can advance while selection is in progress.
- Eight published modules use consistent v0.4.0 sibling requirements without
local `replace` directives. A development `go.work` joins the repository's
modules. Release checks now build each module through an isolated external
consumer; a separate check verifies real versions after publication.
- `Keys()` ordering is explicitly policy-specific. Migration documentation
now describes the limits of preserving order across policies.
- Documentation and the site now describe adaptive selection as experimental,
with measured limitations and workload-dependent results. The historical
explorer dataset is removed; final materials show only current measurements.

### Fixed

- Shadow reads now insert missed keys even when the active policy hits. Add
fan-out avoids counting that fill twice, so shadow statistics represent the
policy's own read-through behavior.
- Gradual migration excludes its source from shadow filling, preventing zero
placeholder values from being promoted and returned as real cached data.
- Zero-traffic stability gates, request-counted epoch progress, request
attribution across epochs now have corrected behavior and regression tests.
- Evidence now measures misses immediately after a switch; TTL documentation
makes clear that expiry still uses wall time in request-counted replays.

### Added

- **Bug fix: a gradual migration could serve a shadow's zero value as real
data.** While a `MigrationGradual` window is open the source policy is
deliberately not demoted - it holds the only copy of every value not yet
promoted - but it is also not the active policy, so the shadow read fan-out
filled it with zero values on a miss. `promoteLocked` reads those values back
with `Peek`, which cannot tell a zero somebody wrote from a real value still
pending, so the zero was promoted into the active policy and returned to the
caller as a hit. The fan-out now skips the migration source until its window
closes. Reachable whenever the working set exceeds the capacity during a
window; a test whose working set exactly fits never evicts and so could not
catch it, which is why the existing gradual-migration test did not.

- **Bug fix: shadow policies could only acquire keys the active policy had
missed**, which made every shadow measurement unreliable whenever the active
policy was performing well, and inverted it outright behind a strong one.
`AdaptiveCache.get` fanned out only `Get` to the shadows; the sole insert
path was the caller's `Add`, which a read-through caller makes only on an
active-policy miss. Measured on a cyclic workload with a 94%-hit incumbent,
arms that truly serve 0.00% were reported above 90%, and `Advice()`
recommended switching from the best arm to the worst. A shadow that misses
now fills itself, and the `Add` fan-out skips keys a shadow already holds so
the fill is not double-counted as an access.
- Every measured number in `docs/evidence.md` was re-run. Adaptive selection
now beats the best fixed policy on **two of the six real traces** rather
than one: LIRS `loop` moved from -7.45 points to +0.12, while P3's margin
shrank from +1.13 to +0.05. The synthetic conclusion is unchanged - on
those five workloads it still never beats the best fixed policy.
- `TestShadowsMeasureWhatThePolicyWouldActuallyServe` pins the property that
was missing: a shadow's measured hit rate must equal the same policy
replayed standalone. Every deterministic arm now matches to within 0.005.

- **S3-FIFO and SIEVE arms** (`policies/fifo`, `ascache.S3FIFO` and
`ascache.SIEVE`) - a new module adapting
[scalalang2/golang-fifo](https://github.com/scalalang2/golang-fifo) v1.2.0
(MIT), kept separate so that dependency stays out of builds that do not use
these arms.
- **S3-FIFO** uses three static FIFO queues: a small queue holding a tenth of
the cache filters keys requested only once, a ghost queue remembers what it
evicted so a returning key is admitted on its second request rather than
its third, and the main queue evicts by FIFO-reinsertion over a counter
capped at three. Cite: Yang, Zhang, Qiu, Yue & Rashmi, *FIFO Queues are All
You Need for Cache Eviction*, SOSP '23.
- **SIEVE** uses one FIFO queue and a hand that sweeps it, evicting the first
entry it reaches that has not been visited since the hand last passed and
clearing the visited bit of every entry it steps over. No ghost queue, no
counters, no second queue. Cite: Zhang, Yang, Yue, Vigfusson & Rashmi,
*SIEVE is Simpler than LRU*, NSDI '24.

Both are deterministic and neither reorders anything on a hit. They share one
module and one adapter because they come from one dependency and need the
same four missing methods supplied; giving each its own module would have
duplicated ~300 lines of deadlock-sensitive glue.
- **`benchclient.DefaultArms` now includes S3-FIFO.** The set exists to be
deterministic and unencumbered, and S3-FIFO is both. This changes the arms a
replay through `benchclient` runs, and therefore its numbers. SIEVE is
deliberately not in it: arms are not free, each one thins the evidence every
other arm gets per epoch, and a default set is the wrong place to add a
second policy from the same family.
- **MSR Cambridge trace loader** (`bench.LoadMSRTrace`). Reads the SNIA IOTTA
block I/O layout `Timestamp,Hostname,DiskNumber,Type,Offset,Size,ResponseTime`,
expanding each record's byte length into the block accesses it covers and
namespacing keys by host and disk. Reads only by default; `IncludeWrites`
models a write-back cache instead. `BlockSize` defaults to 512, matching the
Caffeine simulator's reader so numbers are comparable with what is published
from it.
- **Meta kvcache trace loader** (`bench.LoadMetaKVTrace`). Reads the CacheBench
workloads, locating columns by name so both the 2022 layout
(`key,op,size,op_count,key_size`) and the 2024 one (which reordered them and
added five) are read correctly. Expands `op_count`, which is a repeat count
and not a sequence number.
- **`scripts/fetch-traces.sh` fetches a slice of the Meta trace** over a plain
HTTPS byte-range request - the published files are 5 to 10 GB and need no AWS
credentials to read partially. `AS_CACHE_META_BYTES` sets the size.
- **Trace-loader tests run in `make test`**, not only under `make evidence`.
Pinned against fixtures copied from the real files: a format misread is a
correctness bug that produces a plausible-looking workload, and every number
taken from it is wrong.

### Notes

- **MSR Cambridge cannot be fetched by script.** SNIA serves the files behind a
click-through licence and a cookie check, so `fetch-traces.sh` prints how to
get them by hand rather than pretending to download them. Any file named
`msr_<volume>.csv[.gz]` in the trace directory is picked up automatically.
- **S3-FIFO's ghost queue and the adapter's index both cost memory.** The ghost
queue remembers roughly as many keys as the cache holds, and the adapter
keeps its own copy of the key set on top of that. Values are never
duplicated. The measured total is in
[evidence](docs/evidence.md#memory-and-per-operation-cost).
- **The S3-FIFO adapter pays for four methods `golang-fifo` does not have**:
`Keys`, `Values`, `Resize`, `Cap`, and `Add`'s evicted flag. The costs are
documented on the package and worth reading before quoting this arm's
numbers. The largest is `Resize`, which rebuilds the cache and therefore
discards the ghost queue and every frequency counter - and `AdaptiveCache`
resizes a policy on every promotion and demotion. The adapter also keeps a
second copy of the key set, because the library cannot enumerate its own
contents.
- **Three upstream behaviours the adapter works around.** The eviction callback
runs under the library's mutex, so the adapter's callback must not take its
own lock (it would deadlock; this is why the adapter always builds with a TTL
of zero and therefore no expiry goroutine). S3-FIFO's `Len()` takes no lock,
so the adapter answers `Len` from its own index for both algorithms rather
than depending on which is wrapped. And neither can be built at size zero -
SIEVE panics, S3-FIFO loops waiting to evict from an empty cache and never
returns - so both constructors reject a non-positive size and return an
error, matching `NewLRU`, `NewLFU` and `NewTwoQueue`. An arm built at zero
would accept nothing and report no hits for its whole life, which is a silent
no-op rather than a policy. Resizing an existing cache to zero stays legal,
since `AdaptiveCache.Resize` passes its own capacity through to every arm.
- **Upstream counts a write as an access** for both algorithms - `Set` over a
live key raises S3-FIFO's frequency counter and sets SIEVE's visited bit -
which neither paper does. Shadow policies are driven with `Add`, so these
arms look more used on shadow duty than they should.
- `Settings.MigrationMaxRequests` bounds a gradual migration window by request
count. The default zero retains unlimited request count; see configuration
for its interaction with the existing migration limits.
- MSR Cambridge block-I/O and Meta kvcache trace loaders, fetch support, format
fixtures and independent LRU calibration against pinned libCacheSim.
- Generated [current results](bench/results/current/README.md) with input/output
hashes, sixty LRU calibration points and three consecutive evidence runs.
The tables retain all outcomes, ties, workload context and effective sampling;
they do not assert that adaptive selection always beats the worst fixed policy.
- An offline ObserveOnly sweep on all twelve traces, an object/byte-capacity
comparison for Meta, independent SIEVE diagnostics, and repeated P3 tuning
with request-counted epochs. Raw wall-clock timings are not product claims.
- Experimental S3-FIFO and SIEVE adapters in repository source and the research
suite, planned for v0.5. **The FIFO module is excluded from v0.4.0 publication.**
`benchclient.DefaultArms` retains the four released arms from v0.3.1:
LRU, LFU, 2Q and Random. Random is nondeterministic.

See [releasing and upgrading](docs/releasing.md) for the module set, checks and
publication procedure.

## [0.3.1]

Expand Down
42 changes: 34 additions & 8 deletions Makefile
Original file line number Diff line number Diff line change
Expand Up @@ -5,7 +5,7 @@ MODULES := . lfu policies policies/arc policies/fifo policies/tinylfu metrics ba
GOLANGCI_LINT_VERSION := v2.8.0

.PHONY: all
all: fmt vet lint test release-check ## Format, vet, lint, test and check releasability
all: fmt vet lint test python-check script-test evidence-check release-check ## Format, vet, lint, test and check releasability

.PHONY: lint
lint: ## Run golangci-lint across all modules
Expand Down Expand Up @@ -43,19 +43,29 @@ test: ## Run tests with the race detector across all modules
done

.PHONY: release-check
release-check: ## Check the repository could actually be released today
release-check: release-check-test ## Build eight candidate modules as external consumers
@./scripts/release-check.sh

.PHONY: release-check-test
release-check-test: ## Test the release checker against broken module fixtures
@python3 -m unittest discover -s scripts -p 'release_check_test.py'

.PHONY: release-check-published
release-check-published: ## Verify actual published tags (only after publication)
@./scripts/release-check.sh --published

.PHONY: evidence
evidence: ## Replay the workload suite and print the policy comparison tables
( cd bench && go test -count=1 -timeout 20m -v ./... )
evidence: ## Verify all 13 pinned trace files and replay the complete workload suite
@python3 scripts/trace_inputs.py "$${AS_CACHE_TRACES:?set AS_CACHE_TRACES}"
( cd bench && go test -count=1 -timeout 45m -v ./... )

.PHONY: verify-ref
verify-ref: ## Calibrate the trace loaders and LRU against libCacheSim (needs AS_CACHE_TRACES)
@./scripts/verify-ref.sh

.PHONY: tidy
tidy: ## Run go mod tidy across all modules
@set -e; for m in $(MODULES); do \
echo "==> tidy $$m"; \
( cd $$m && go mod tidy ); \
done
@python3 scripts/tidy.py

.PHONY: install-tools
install-tools: ## Install golangci-lint at the pinned version
Expand All @@ -65,3 +75,19 @@ install-tools: ## Install golangci-lint at the pinned version
help: ## Show this help
@grep -hE '^[a-zA-Z_-]+:.*?## .*$$' $(MAKEFILE_LIST) | \
awk 'BEGIN {FS = ":.*?## "}; {printf " \033[36m%-14s\033[0m %s\n", $$1, $$2}'

.PHONY: python-check
python-check: ## Lint and check formatting with pinned Ruff
@./scripts/python-check.sh

.PHONY: script-test
script-test: ## Check trace integrity and simulator export regressions
@python3 -m unittest discover -s scripts -p '*_test.py'

.PHONY: measure-bytes
measure-bytes: ## Compare Meta LRU with object and byte capacities
@python3 scripts/measure_bytes.py "$${AS_CACHE_TRACES:?set AS_CACHE_TRACES}" "$${AS_CACHE_BYTES_OUT:?set AS_CACHE_BYTES_OUT}"

.PHONY: evidence-check
evidence-check: ## Verify retained evidence and generated tables without rerunning measurements
@python3 -B scripts/record_evidence.py --verify bench/results/current
Loading
Loading