Skip to content

Prepare v0.4 with verified current evidence and installable modules - #17

Open
sshaplygin wants to merge 20 commits into
mainfrom
v04/release
Open

sshaplygin wants to merge 20 commits into
mainfrom
v04/release

Conversation

@sshaplygin

@sshaplygin sshaplygin commented Sep 27, 2026 •

Copy link
Copy Markdown
Owner

The previous evidence gate rejected unfavorable random outcomes, and the release check could package uncommitted code. This prepares the v0.4.0 candidate with a complete current dataset, explicit measurement limits, and eight modules checked as external dependencies.

  • Remove the adaptive-versus-worst assertion while retaining every comparison. Record three consecutive full runs on one clean source commit, with 15 observations per nondeterministic subject, all nine policies, twelve traces and all 10/20/50 request-epoch settings.
  • Verify all downloaded and cached traces by size and SHA256, reject missing inputs and reference coverage gaps, and calibrate 60 LRU points against pinned libCacheSim within 0.0051 percentage points.
  • Add the Meta object/byte experiment, offline ObserveOnly sweep, independent SIEVE model and FIFO control, and repeated request-counted P3 tuning. Explain effective sampling, small margins, ties, size-accounting semantics and the limits of each experiment.
  • Generate and validate manifests and tables at bench/results/current/. Replace old results and remove the historical site explorer and timeline. Remove unsupported maximum-deficit and single-run timing claims from the documentation and site.
  • Require a clean committed release candidate, discover publishable modules with explicit exclusions, test failure reasons and nested-module packaging, and build eight isolated consumers. Keep FIFO outside v0.4.0; restore benchclient.DefaultArms to the four arms published in v0.3.1. Add pinned Python linting and formatting and make prepublication tidy skips explicit.

Validation: three consecutive full make evidence runs passed on clean source 3331d5f; all 60 reference points passed. The strict manifest verifier checks all 16 artifacts and recomputes pooled results from the raw batches. Final make all passed on 29a831f, including vet/lint/race-short checks across twelve modules, 19 script tests, 10 release-check fixtures and eight external consumer builds. Repeated -race -short -count=3 tests passed in the root, bench and benchclient modules, together with basic and cold/warm/gradual migration smoke runs. GitHub CI passed all 25 checks on 29a831f. Independent QA, critic and reviewer all approved the final candidate.

This PR prepares a candidate. No v0.4.0 tags or GitHub release have been published. Human review and merge, tag publication and make release-check-published remain required. The offline ObserveOnly sweep is complete; a real-service trial remains pending a service and environment.

fetch-traces.sh could not download MSR Cambridge and printed instructions for
fetching it by hand from SNIA IOTTA trace 388. That source hands files out
only through a browser form (cookies, name, affiliation, email), and in
September 2026 iotta.snia.org did not respond at all, from two separate
networks. MSR volumes were therefore never part of a scripted evidence run.

The script now takes six volumes (hm_0, prn_0, proj_0, src1_2, usr_0, web_0,
210 MB) from the mirror the cacheMon project keeps at
cache-datasets.s3.amazonaws.com/cache_dataset_txt/2008_msr, which holds
SNIA's original msr-cambridge1.tar and msr-cambridge2.tar. The SNIA Trace
Data Files Download License v2.0 permits use and redistribution without
restriction, so this is a lawful copy. Only the listed volumes are fetched,
by byte range out of the 5.3 GB of uncompressed tar.

Each entry pins the volume's tar header offset, its size, and the MD5 given
for it in the archive's MD5.txt. Before downloading, the script reads the
member name from the tar header at that offset; after downloading, it checks
the MD5. Either mismatch fails the script and leaves no file behind. Those
checksums come from inside the mirrored archive, so they detect a corrupted
or repacked download, not deliberate tampering; the mirror has not been
compared with a copy from SNIA, because SNIA could not be reached.
docs/benchmarking.md says so.

Verified:
- the script fetches all six volumes and every MD5 matches (19.7 s);
- a wrong MD5 and a wrong offset each fail with exit 1, naming the problem
  (for the offset, the member actually found there), and leave no file;
- bash -n passes; shellcheck is not installed here and was not run;
- golangci-lint and go vet on bench: clean (the Go change is a comment).
Every number the evidence suite reports from a real trace depends on the
loader turning the file into the right request sequence and on the replay
counting hits the way other simulators do. The format fixtures prove only
that a loader reads the rows it was tested on. Nothing compared the whole
pipeline with an independent implementation, and the v0.4.0 plan requires
that comparison before the trace tables are regenerated.

make verify-ref (scripts/verify-ref.sh) does it for all twelve traces the
suite reads: Twitter cluster052, LIRS loop and 2_pools, ARC P3 and OLTP,
Meta kvcache 202206 and the six MSR volumes.

- awk in the script expands each raw file into one key per request,
  following the loader's documented rules: ARC block runs, Meta op_count
  repeats over GET rows, MSR 512-byte block ranges over reads. It does not
  call the Go loaders, so a loader bug cannot cancel itself out.
- libCacheSim's cachesim, built into .tools/ at the pinned commit
  1d7415569978330ea95c9cff06a260630406f7e3, replays that sequence through
  LRU with object sizes ignored, at 0.25x to 4x the suite's capacity.
- TestLRUMatchesReference loads the same files through the Go loaders,
  replays this repository's LRU at the same capacities, and requires the
  same request count and a miss ratio within 0.5 points. It also requires
  the reference to include the capacity the suite actually uses.

The gate fails rather than passing over nothing: libCacheSim that will not
build, a missing trace, an unset AS_CACHE_TRACES, or a skipped Go test each
exit 1 with a message. The test skips when no reference is supplied, so
make test and make evidence are unaffected. Nothing runs in CI:
libCacheSim needs a C toolchain with glib and argp, and the traces are not
committed.

Verified:
- make verify-ref: 60 points, largest miss-ratio difference 0.005 points
  (the rounding of cachesim's four-decimal output), request counts equal
  at every point;
- the test fails with a miss ratio off by 1.4 points, with a request count
  off by one, and with a reference that omits the suite's capacity, and
  passes on the correct row; it skips when the reference is unset;
- the script exits 1 with a trace missing and with AS_CACHE_TRACES unset;
- golangci-lint on bench: 0 issues. shellcheck is not installed here.

The first run of the script stopped after the third trace with status 141
and no message: once awk reached the request limit, gzip took SIGPIPE and
pipefail ended the script. Decompression now tolerates exactly that status,
and an ERR trap reports the line on which any other failure stopped.
TestTraceEvidence replayed the adaptive cache on a 2ms wall-clock epoch, once
per trace. How many epochs a replay saw depended on how fast the machine ran
it, so the result moved between runs, and the table in docs/evidence.md
(measured at 50ms, a setting that was never in the test) could not be
reproduced by make evidence. On cd8502f and on 1e2e599, before the B1 and
B2 changes, the 2ms test gave ARC P3 3.3-3.7% against the documented 11.4%.
At 50ms, five runs across both commits gave 6.2-9.6%.

The test now:
- ends epochs on request counts (EpochRequests), set as a number of epochs
  over the whole trace (10, 20 and 50), reporting all three rather than the
  best of them;
- keeps the configuration production would use: all nine arms and a shadow
  sample rate of 0.05;
- replays every subject that is not reproducible five times and reports the
  median with its range. That covers the adaptive cache, whose sampler hash
  is seeded per cache, and the Random and W-TinyLFU arms. The deterministic
  arms are replayed once;
- prints a per-trace table and a summary row per trace, and, when
  AS_CACHE_EVIDENCE_OUT names a file, writes every run as JSON with the
  commit, whether the tree was modified, the Go version, the platform and
  the settings.

The assertion is unchanged in substance: the adaptive median at each epoch
length must beat the median of the worst fixed policy. It moved to a new
file, bench/trace_evidence_test.go, because trace_test.go keeps the trace
list and the loader checks. make evidence's timeout rises from 20m to 45m
for the extra replays.

Verified: TestTraceEvidence passes on all twelve traces (Twitter, LIRS loop
and 2_pools, ARC P3 and OLTP, Meta kvcache, six MSR volumes) in 330 s, and
writes the JSON. go vet and golangci-lint on bench: clean. The full make
evidence run and the docs rewrite follow separately.
@sshaplygin
sshaplygin marked this pull request as ready for review September 27, 2026 22:05
@sshaplygin
sshaplygin marked this pull request as draft September 28, 2026 20:41
@sshaplygin sshaplygin changed the title Prepare v0.4 with calibrated trace evidence and installable modules Prepare v0.4 with verified current evidence and installable modules Sep 28, 2026
@sshaplygin
sshaplygin marked this pull request as ready for review September 28, 2026 21:20

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant