Prepare v0.4 with verified current evidence and installable modules - #17
Open
sshaplygin wants to merge 20 commits into
Open
sshaplygin wants to merge 20 commits into
sshaplygin wants to merge 20 commits into
Conversation
fetch-traces.sh could not download MSR Cambridge and printed instructions for fetching it by hand from SNIA IOTTA trace 388. That source hands files out only through a browser form (cookies, name, affiliation, email), and in September 2026 iotta.snia.org did not respond at all, from two separate networks. MSR volumes were therefore never part of a scripted evidence run. The script now takes six volumes (hm_0, prn_0, proj_0, src1_2, usr_0, web_0, 210 MB) from the mirror the cacheMon project keeps at cache-datasets.s3.amazonaws.com/cache_dataset_txt/2008_msr, which holds SNIA's original msr-cambridge1.tar and msr-cambridge2.tar. The SNIA Trace Data Files Download License v2.0 permits use and redistribution without restriction, so this is a lawful copy. Only the listed volumes are fetched, by byte range out of the 5.3 GB of uncompressed tar. Each entry pins the volume's tar header offset, its size, and the MD5 given for it in the archive's MD5.txt. Before downloading, the script reads the member name from the tar header at that offset; after downloading, it checks the MD5. Either mismatch fails the script and leaves no file behind. Those checksums come from inside the mirrored archive, so they detect a corrupted or repacked download, not deliberate tampering; the mirror has not been compared with a copy from SNIA, because SNIA could not be reached. docs/benchmarking.md says so. Verified: - the script fetches all six volumes and every MD5 matches (19.7 s); - a wrong MD5 and a wrong offset each fail with exit 1, naming the problem (for the offset, the member actually found there), and leave no file; - bash -n passes; shellcheck is not installed here and was not run; - golangci-lint and go vet on bench: clean (the Go change is a comment).
Every number the evidence suite reports from a real trace depends on the loader turning the file into the right request sequence and on the replay counting hits the way other simulators do. The format fixtures prove only that a loader reads the rows it was tested on. Nothing compared the whole pipeline with an independent implementation, and the v0.4.0 plan requires that comparison before the trace tables are regenerated. make verify-ref (scripts/verify-ref.sh) does it for all twelve traces the suite reads: Twitter cluster052, LIRS loop and 2_pools, ARC P3 and OLTP, Meta kvcache 202206 and the six MSR volumes. - awk in the script expands each raw file into one key per request, following the loader's documented rules: ARC block runs, Meta op_count repeats over GET rows, MSR 512-byte block ranges over reads. It does not call the Go loaders, so a loader bug cannot cancel itself out. - libCacheSim's cachesim, built into .tools/ at the pinned commit 1d7415569978330ea95c9cff06a260630406f7e3, replays that sequence through LRU with object sizes ignored, at 0.25x to 4x the suite's capacity. - TestLRUMatchesReference loads the same files through the Go loaders, replays this repository's LRU at the same capacities, and requires the same request count and a miss ratio within 0.5 points. It also requires the reference to include the capacity the suite actually uses. The gate fails rather than passing over nothing: libCacheSim that will not build, a missing trace, an unset AS_CACHE_TRACES, or a skipped Go test each exit 1 with a message. The test skips when no reference is supplied, so make test and make evidence are unaffected. Nothing runs in CI: libCacheSim needs a C toolchain with glib and argp, and the traces are not committed. Verified: - make verify-ref: 60 points, largest miss-ratio difference 0.005 points (the rounding of cachesim's four-decimal output), request counts equal at every point; - the test fails with a miss ratio off by 1.4 points, with a request count off by one, and with a reference that omits the suite's capacity, and passes on the correct row; it skips when the reference is unset; - the script exits 1 with a trace missing and with AS_CACHE_TRACES unset; - golangci-lint on bench: 0 issues. shellcheck is not installed here. The first run of the script stopped after the third trace with status 141 and no message: once awk reached the request limit, gzip took SIGPIPE and pipefail ended the script. Decompression now tolerates exactly that status, and an ERR trap reports the line on which any other failure stopped.
TestTraceEvidence replayed the adaptive cache on a 2ms wall-clock epoch, once per trace. How many epochs a replay saw depended on how fast the machine ran it, so the result moved between runs, and the table in docs/evidence.md (measured at 50ms, a setting that was never in the test) could not be reproduced by make evidence. On cd8502f and on 1e2e599, before the B1 and B2 changes, the 2ms test gave ARC P3 3.3-3.7% against the documented 11.4%. At 50ms, five runs across both commits gave 6.2-9.6%. The test now: - ends epochs on request counts (EpochRequests), set as a number of epochs over the whole trace (10, 20 and 50), reporting all three rather than the best of them; - keeps the configuration production would use: all nine arms and a shadow sample rate of 0.05; - replays every subject that is not reproducible five times and reports the median with its range. That covers the adaptive cache, whose sampler hash is seeded per cache, and the Random and W-TinyLFU arms. The deterministic arms are replayed once; - prints a per-trace table and a summary row per trace, and, when AS_CACHE_EVIDENCE_OUT names a file, writes every run as JSON with the commit, whether the tree was modified, the Go version, the platform and the settings. The assertion is unchanged in substance: the adaptive median at each epoch length must beat the median of the worst fixed policy. It moved to a new file, bench/trace_evidence_test.go, because trace_test.go keeps the trace list and the loader checks. make evidence's timeout rises from 20m to 45m for the extra replays. Verified: TestTraceEvidence passes on all twelve traces (Twitter, LIRS loop and 2_pools, ARC P3 and OLTP, Meta kvcache, six MSR volumes) in 330 s, and writes the JSON. go vet and golangci-lint on bench: clean. The full make evidence run and the docs rewrite follow separately.
sshaplygin
marked this pull request as ready for review
September 27, 2026 22:05
sshaplygin
marked this pull request as draft
September 28, 2026 20:41
sshaplygin
marked this pull request as ready for review
September 28, 2026 21:20
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The previous evidence gate rejected unfavorable random outcomes, and the release check could package uncommitted code. This prepares the v0.4.0 candidate with a complete current dataset, explicit measurement limits, and eight modules checked as external dependencies.
bench/results/current/. Replace old results and remove the historical site explorer and timeline. Remove unsupported maximum-deficit and single-run timing claims from the documentation and site.benchclient.DefaultArmsto the four arms published in v0.3.1. Add pinned Python linting and formatting and make prepublication tidy skips explicit.Validation: three consecutive full
make evidenceruns passed on clean source3331d5f; all 60 reference points passed. The strict manifest verifier checks all 16 artifacts and recomputes pooled results from the raw batches. Finalmake allpassed on29a831f, including vet/lint/race-short checks across twelve modules, 19 script tests, 10 release-check fixtures and eight external consumer builds. Repeated-race -short -count=3tests passed in the root, bench and benchclient modules, together with basic and cold/warm/gradual migration smoke runs. GitHub CI passed all 25 checks on29a831f. Independent QA, critic and reviewer all approved the final candidate.This PR prepares a candidate. No v0.4.0 tags or GitHub release have been published. Human review and merge, tag publication and
make release-check-publishedremain required. The offline ObserveOnly sweep is complete; a real-service trial remains pending a service and environment.