From b8e31b92f97da2a335dcc280a5c45b99323fc1ec Mon Sep 17 00:00:00 2001 From: Sam Shaplygin Date: Mon, 24 Aug 2026 23:58:04 +0200 Subject: [PATCH] doc: correct fourteen claims an audit found wrong A per-file audit against the code turned up claims that were wrong, stale, or not checkable against anything. The substantive ones: The phase-shift timeline said the bandit "identifies W-TinyLFU and holds it" for 82% of the run, and showed a plot to match. Re-running it, no arm holds more than 27% and eight of the nine take a turn. The old figure was measured while the shadow mechanism was reporting rivals at rates they could not achieve; it does not reproduce, and the conclusion drawn from it was the opposite of what the cache now does. The plot is replaced by the run's actual output and the prose says what it shows: on a workload whose arms sit within a couple of points of each other, Thompson sampling explores rather than settling, and that exploration is what the 63.77% costs. The from-scratch S3-FIFO comparison is gone. The implementation it measured was deleted, nothing in `make evidence` reproduces those columns, and no test guards them. What it taught -- that a differential run found a ghost-queue bug every property test missed -- is kept as prose. Cold migration was said to cost 30.7 points on OLTP, comparing a 2ms cold run against a 50ms warm one. At the same epoch it is 25.7. The shadow overhead table predated the fan-out skip added with the shadow fix, which took `Add` with three shadows from 193 to 142 ns/op. Smaller corrections: the capacity gate reads the active policy alone and suspends every arm's measurement rather than one arm's; LFU is no longer the best policy on synthetic zipf, having been overtaken by both FIFO arms; W-TinyLFU is not the only non-reproducible arm, Random is the other; `size/10` is not a cap on S3-FIFO's small queue but the threshold choosing which queue an eviction comes from; the fleet section's mixed-fleet row is six replicas, not eight; two of the six traces are research traces rather than production ones; and the cache applies every bandit selection unconditionally unless stability gates are opted into, rather than weighing a win against migration cost. --- README.md | 17 +++--- docs/configuration.md | 43 ++++++++------- docs/evidence.md | 125 +++++++++++++++++------------------------- 3 files changed, 83 insertions(+), 102 deletions(-) diff --git a/README.md b/README.md index fffa84a..7bcbbbd 100644 --- a/README.md +++ b/README.md @@ -8,7 +8,7 @@ Choosing a cache eviction policy is a decision most projects make once, from intuition, and never revisit. The trouble is that the right answer depends on traffic you have not seen yet, and it is not stable: replayed against six -published production traces, **four different policies win**, and the strongest +published traces, **four different policies win**, and the strongest general-purpose baseline of them all comes near the bottom on one. Guessing wrong is not a rounding error either — on one of those traces seven of the nine policies here serve **0.0%** while one serves 45%. @@ -17,16 +17,19 @@ as-cache makes the choice at runtime instead. One policy is **active** and serves every request. The others run as **shadows**: they see each key but never its value, and answer "would I have had this?" Once per epoch every arm reports its hit rate, a multi-armed bandit names the winner, and the cache -switches if the win is worth the migration. There is also an observe-only mode +switches to it — unconditionally by default, or subject to [stability +gates](docs/configuration.md#keeping-switches-stable) you opt into. There is also an observe-only mode where nothing ever switches and the library simply tells you which policy your traffic wants — often the more useful half of it. It is pre-1.0, the API may still change, and nothing here has run in production -that I know of. What it does have is measurement: every claim in these -documents comes from a reproducible run against published traces, and the -concurrency has been exercised under the race detector and adversarially -reviewed. The numbers are all in [the evidence](docs/evidence.md), so you do -not have to take "experimental" or "production-ready" on trust. +that I know of. What it does have is measurement: every number in these +documents comes from a run you can repeat with `make evidence`, over published +traces and generated workloads both, and the two arms whose results do not +repeat exactly are named wherever their numbers appear. The concurrency has +been exercised under the race detector and adversarially reviewed. It is all in +[the evidence](docs/evidence.md), so you do not have to take "experimental" or +"production-ready" on trust. ## Documentation diff --git a/docs/configuration.md b/docs/configuration.md index de12ace..c5fb35b 100644 --- a/docs/configuration.md +++ b/docs/configuration.md @@ -75,19 +75,18 @@ than what any particular policy costs: | Benchmark | shadows | sampling off | rate 0.05 | | --- | --- | --- | --- | -| `Get` | 1 | 97 ns/op | 35 ns/op | -| `Get` | 3 | 148 ns/op | 38 ns/op | -| `GetParallel` | 1 | 271 ns/op | 181 ns/op | -| `GetParallel` | 3 | 405 ns/op | 185 ns/op | -| `Add` | 1 | 110 ns/op | 52 ns/op | -| `Add` | 3 | 193 ns/op | 59 ns/op | -| `MixedParallel` | 1 | 183 ns/op | 87 ns/op | - -Read the `Get` rows down the shadow count. Unsampled, a third shadow costs -another 50ns, because every operation visits every policy. Sampled, going from -one shadow to three costs 3ns -- the fan-out happens on 5% of operations, so -adding a policy is close to free. That is what makes carrying nine arms -practical. +| `Get` | 1 | 102 ns/op | 40 ns/op | +| `Get` | 3 | 153 ns/op | 44 ns/op | +| `GetParallel` | 1 | 284 ns/op | 183 ns/op | +| `GetParallel` | 3 | 361 ns/op | 187 ns/op | +| `Add` | 1 | 98 ns/op | 53 ns/op | +| `Add` | 3 | 142 ns/op | 55 ns/op | +| `MixedParallel` | 1 | 187 ns/op | 96 ns/op | + +Read the `Get` rows down the shadow count. Unsampled, two further shadows cost +another 51ns, because every operation visits every policy. Sampled, the same +step costs 4ns -- the fan-out happens on 5% of operations, so adding a policy +is close to free. That is what makes carrying nine arms practical. Reproduce with `go test -run '^$' -bench . -benchtime=300ms .` @@ -126,8 +125,10 @@ not subtle. Measured on the ARC P3 trace with a 20k-entry cache: An epoch short enough to trigger frequent switches makes the cache copy its entire contents on every switch, so it spends its time migrating rather than serving. Cold migration is worse: it discards the cache at each switch, which -on the OLTP trace costs 30.7 points against warm migration at the same epoch -(37.2% against 67.8%). +on the OLTP trace costs 25.7 points against warm migration at the same epoch +(37.2% against 62.9%, both at 2ms). There is no 50ms cold run to compare +against; the sweep covers 2ms warm, 2ms cold, 50ms warm and 50ms warm with the +stability gates. Rules of thumb: @@ -140,8 +141,10 @@ Rules of thumb: which needs to re-adapt constantly, and 0.8 on OLTP, which does not. - `ShadowSampleRate: 0.05` is a reasonable default. Higher rates cost more and buy no better ranking. -- Set `EvictPartialCapacityFilling: true` when W-TinyLFU or S3-FIFO is one of - the arms. The capacity gate compares `Len()` against `Cap()` for exact - equality, and neither arm holds that reliably: otter reports an approximate - size, and S3-FIFO can drop about a tenth of the cache in a single write. An - arm sitting below its capacity has its epochs skipped entirely. +- Set `EvictPartialCapacityFilling: true` when W-TinyLFU or S3-FIFO could be + the **active** arm. The gate reads that one policy and compares its `Len()` + against `Cap()` for exact equality, and neither of those two holds that + reliably: otter reports an approximate size, and S3-FIFO can drop about a + tenth of the cache in a single write. When it fires, the whole epoch is + skipped — no arm is reported and none is reset — so one under-filled active + policy suspends every measurement, not just its own. diff --git a/docs/evidence.md b/docs/evidence.md index 9c1e4d2..7b21ffe 100644 --- a/docs/evidence.md +++ b/docs/evidence.md @@ -43,9 +43,10 @@ TTL shares a column with LRU because these workloads carry no notion of staleness and its TTL is longer than any run, so it measures its LRU behaviour exactly -- identically, to the hundredth of a point, on all five. -Every arm here reproduces to the hundredth of a point between runs except -W-TinyLFU, which does not: on `loop` it has measured 88.7% and 94.5% within a -single process. Read its column, and any delta computed against it, with that +Every arm here reproduces to the hundredth of a point between runs except two. +W-TinyLFU does not: on `loop` it has measured 88.7% and 94.5% within a single +process. Neither does `Random`, which seeds itself from the global source — +that is what a control arm is for, but it means its column moves too. Read its column, and any delta computed against it, with that in mind — the cause is in [reproducible replays](benchmarking.md). Two things stand out. LRU and LFU both score **exactly zero** on `loop`, where a @@ -185,32 +186,31 @@ against the hit rate. The real-trace figures below are far tighter than this table, because those replays use a tuned 50ms epoch rather than the 2ms one held fixed across every workload here. -The timeline explains why. Replaying `phase-shift` and sampling `ActivePolicy()` -throughout: +The timeline says why, and it is not the answer this section used to give. +Replaying `phase-shift` and sampling `ActivePolicy()` throughout: ```text -phase Z------L------Z------L------Z------L------Z------L------ (Z = zipf, L = loop) -LRU ### -TwoQueue ##### -SIEVE ### -LFU ## ### ##### -S3FIFO ## ## -TinyLFU ############################################# - -share of time active: LRU 2%, LFU 6%, TwoQueue 3%, TinyLFU 82%, S3FIFO 2%, SIEVE 3% -hit rate 77.05% +share of time active: LRU 6%, LFU 4%, TwoQueue 15%, ARC 14%, TTL 4%, + TinyLFU 27%, S3FIFO 15%, SIEVE 16% +hit rate 63.77% ``` -The bandit works exactly as designed: it explores, identifies W-TinyLFU, and -holds it for 82% of the run. It does not oscillate at phase boundaries, because -there is no crossover to exploit -- W-TinyLFU is the best arm in *both* regimes. -The remaining gap is the price of exploring and of migrating between arms. - -That timeline is one draw, not a fixed result: this test uses a wall-clock -epoch, so the number of epochs in a run varies with machine load and which -also-ran arms collect a sliver of exploration varies with it. Across six -consecutive unmodified runs S3-FIFO appeared in four and SIEVE in five. Read -the 82% as the finding and the slivers as noise. +**The bandit does not settle.** Eight of the nine arms take a turn, the best of +them holds only 27% of the run, and the cache spends the rest of it changing +its mind. That is not a defect in the bandit; it is what honest evidence looks +like on this workload. `phase-shift` alternates between two regimes every +20,000 requests, several arms sit within a couple of points of each other in +both, and Thompson sampling explores exactly as it should when the posteriors +overlap. The cost of that exploration is the gap between 63.77% here and a +fixed W-TinyLFU's 82.6%. + +Two things are worth saying plainly about this block. Earlier versions of this +document showed W-TinyLFU holding 82-90% of the same run and concluded that the +bandit "identifies W-TinyLFU and holds it"; that was measured while the shadow +mechanism was reporting rivals at rates they could not achieve, and it does not +reproduce. And the run still uses a wall-clock epoch, so the number of epochs +varies with machine load — read the shape (no arm dominates) as the finding and +the individual percentages as one draw. So the case for this library is not "it beats the best policy." It is: @@ -317,11 +317,13 @@ Then there is `loop`, where it serves **0.00%** -- tied with LRU, LFU, 2Q, TTL, ARC and SIEVE, and beaten by random eviction. That is the algorithm behaving exactly as designed. `loop` cycles through 1011 keys with a 500-entry cache, so every reuse distance is 1011 requests. A key has to be requested again while it -is still in the small queue (50 entries) or still in the ghost queue to be -promoted. The library sizes that ghost queue at the *whole* cache capacity -- -500 here, not the main queue's 450 -- so the horizon is roughly 1000 requests -wide, and a reuse distance of 1011 falls just outside it. Every key goes round -the small queue forever and nothing is ever promoted. Only two arms survive the +is still resident or still in the ghost queue to be promoted. The library sizes +that ghost queue at the *whole* cache capacity -- 500 here -- so the horizon is +roughly 1000 requests wide, and a reuse distance of 1011 falls just outside it. +Nothing is ever promoted to the main queue, so the small queue holds the whole +cache and every key cycles through it forever. (`size/10` is not a cap on that +queue: upstream uses it only to decide which of the two queues an eviction +comes from.) Only two arms survive the trace at all: W-TinyLFU's sketch, which ages rather than expiring, and random eviction, which has no order to defeat. @@ -333,51 +335,23 @@ distance. ### What the library adapter costs -This arm wraps [scalalang2/golang-fifo](https://github.com/scalalang2/golang-fifo) -rather than implementing the algorithm here, and the wrapper is not free. The +These arms wrap [scalalang2/golang-fifo](https://github.com/scalalang2/golang-fifo) +rather than implementing the algorithms here, and the wrapper is not free. The adapter has to supply `Keys`, `Values`, `Resize` and `Cap`, none of which exist upstream, which means a second copy of the key set maintained through the -library's eviction callback and a full rebuild on every resize. Measured -against a from-scratch implementation of the same algorithm, replaying the same -traces at the same capacities. - -**Read the right-hand columns as history, not as a measurement you can repeat.** -The from-scratch implementation was written first, measured against the library -here, and then deleted when the library replaced it — so nothing in `make -evidence` reproduces that side of this table, and no test guards it. It is kept -because it is the evidence behind choosing the dependency. The `via the -library` columns are reproducible and are re-measured with everything else. - -| Trace | via the library | ns/op | from scratch | ns/op | -| --- | --- | --- | --- | --- | -| Twitter | 59.73% | 450 | 61.47% | 162 | -| Meta kvcache | 69.05% | 351 | 69.75% | 137 | -| ARC OLTP | 67.79% | 385 | 68.32% | 157 | -| ARC P3 | 10.75% | 709 | 9.98% | 291 | - -**Two to three times the per-operation cost**, and between half a point worse -and three quarters of a point better on hit rate. The time goes on the index -the adapter maintains alongside the library and on the library's linked-list -allocations; the hit-rate differences come from the two implementations sizing -the ghost queue differently and from upstream counting a write as an access. - -The trade bought is not owning an eviction algorithm. That is worth something: -the from-scratch version shipped with a real bug in its ghost queue, found only -by differential testing against a second implementation, and every property -test passed while it was there. - -LFU is the sharpest illustration of why synthetic workloads mislead. It is the -**best** policy on synthetic `zipf` (73.5%) and the **worst** on both large real -traces (41.4% on Twitter, 45.4% on OLTP). Synthetic Zipf holds popularity -*stationary*, which is exactly the assumption classic LFU makes; real traffic -shifts, and an entry that was hot once keeps a frequency count that holds it -resident long after it stops being useful. That is the failure W-TinyLFU's aged -frequency sketch exists to avoid, and it is invisible until you replay real -traffic. - -How much configuration moves these numbers is in -[configuration](configuration.md#tuning-measured); the short version is that a -too-short epoch costs up to 7 points and 30x the per-operation time. +library's eviction callback and a full rebuild on every resize. The per-operation +columns in the table above are what that costs: 371 to 770 ns/op for S3-FIFO +against 80 to 130 for a bare LRU on the same traces. + +The trade bought is not owning an eviction algorithm, and that is worth +something. An earlier from-scratch S3-FIFO in this repository shipped with a +real bug in its ghost queue -- the paper's virtual-timestamp approximation +leaves a dead slot behind whenever an entry is removed early, so the queue +steadily held fewer keys than its capacity claimed and dropped them just before +they came back. Every property test passed; only a differential run against a +second implementation found it. That implementation is gone, so its numbers are +not quoted here: nothing in `make evidence` reproduces them and no test guards +them. ## Does sampling distort the comparison? @@ -437,8 +411,9 @@ effective rate rather than letting a miniature shrink into noise. ## Does pooling across a fleet help? **Only in the regime it was built for, and it is worth checking you are in that -regime before turning it on.** All figures are 8 replicas, cache capacity 300 -to 500, `make evidence`. The mechanism is described in [running a +regime before turning it on.** Figures are 8 replicas at cache capacity 300 to +500 unless the row says otherwise -- the mixed-fleet row below is 6 replicas at +400 -- and all come from `make evidence`. The mechanism is described in [running a fleet](fleet.md). The case it exists for is a replica that sees too little traffic per epoch to