diff --git a/README.md b/README.md index fffa84a..7bcbbbd 100644 --- a/README.md +++ b/README.md @@ -8,7 +8,7 @@ Choosing a cache eviction policy is a decision most projects make once, from intuition, and never revisit. The trouble is that the right answer depends on traffic you have not seen yet, and it is not stable: replayed against six -published production traces, **four different policies win**, and the strongest +published traces, **four different policies win**, and the strongest general-purpose baseline of them all comes near the bottom on one. Guessing wrong is not a rounding error either — on one of those traces seven of the nine policies here serve **0.0%** while one serves 45%. @@ -17,16 +17,19 @@ as-cache makes the choice at runtime instead. One policy is **active** and serves every request. The others run as **shadows**: they see each key but never its value, and answer "would I have had this?" Once per epoch every arm reports its hit rate, a multi-armed bandit names the winner, and the cache -switches if the win is worth the migration. There is also an observe-only mode +switches to it — unconditionally by default, or subject to [stability +gates](docs/configuration.md#keeping-switches-stable) you opt into. There is also an observe-only mode where nothing ever switches and the library simply tells you which policy your traffic wants — often the more useful half of it. It is pre-1.0, the API may still change, and nothing here has run in production -that I know of. What it does have is measurement: every claim in these -documents comes from a reproducible run against published traces, and the -concurrency has been exercised under the race detector and adversarially -reviewed. The numbers are all in [the evidence](docs/evidence.md), so you do -not have to take "experimental" or "production-ready" on trust. +that I know of. What it does have is measurement: every number in these +documents comes from a run you can repeat with `make evidence`, over published +traces and generated workloads both, and the two arms whose results do not +repeat exactly are named wherever their numbers appear. The concurrency has +been exercised under the race detector and adversarially reviewed. It is all in +[the evidence](docs/evidence.md), so you do not have to take "experimental" or +"production-ready" on trust. ## Documentation diff --git a/docs/configuration.md b/docs/configuration.md index de12ace..c5fb35b 100644 --- a/docs/configuration.md +++ b/docs/configuration.md @@ -75,19 +75,18 @@ than what any particular policy costs: | Benchmark | shadows | sampling off | rate 0.05 | | --- | --- | --- | --- | -| `Get` | 1 | 97 ns/op | 35 ns/op | -| `Get` | 3 | 148 ns/op | 38 ns/op | -| `GetParallel` | 1 | 271 ns/op | 181 ns/op | -| `GetParallel` | 3 | 405 ns/op | 185 ns/op | -| `Add` | 1 | 110 ns/op | 52 ns/op | -| `Add` | 3 | 193 ns/op | 59 ns/op | -| `MixedParallel` | 1 | 183 ns/op | 87 ns/op | - -Read the `Get` rows down the shadow count. Unsampled, a third shadow costs -another 50ns, because every operation visits every policy. Sampled, going from -one shadow to three costs 3ns -- the fan-out happens on 5% of operations, so -adding a policy is close to free. That is what makes carrying nine arms -practical. +| `Get` | 1 | 102 ns/op | 40 ns/op | +| `Get` | 3 | 153 ns/op | 44 ns/op | +| `GetParallel` | 1 | 284 ns/op | 183 ns/op | +| `GetParallel` | 3 | 361 ns/op | 187 ns/op | +| `Add` | 1 | 98 ns/op | 53 ns/op | +| `Add` | 3 | 142 ns/op | 55 ns/op | +| `MixedParallel` | 1 | 187 ns/op | 96 ns/op | + +Read the `Get` rows down the shadow count. Unsampled, two further shadows cost +another 51ns, because every operation visits every policy. Sampled, the same +step costs 4ns -- the fan-out happens on 5% of operations, so adding a policy +is close to free. That is what makes carrying nine arms practical. Reproduce with `go test -run '^$' -bench . -benchtime=300ms .` @@ -126,8 +125,10 @@ not subtle. Measured on the ARC P3 trace with a 20k-entry cache: An epoch short enough to trigger frequent switches makes the cache copy its entire contents on every switch, so it spends its time migrating rather than serving. Cold migration is worse: it discards the cache at each switch, which -on the OLTP trace costs 30.7 points against warm migration at the same epoch -(37.2% against 67.8%). +on the OLTP trace costs 25.7 points against warm migration at the same epoch +(37.2% against 62.9%, both at 2ms). There is no 50ms cold run to compare +against; the sweep covers 2ms warm, 2ms cold, 50ms warm and 50ms warm with the +stability gates. Rules of thumb: @@ -140,8 +141,10 @@ Rules of thumb: which needs to re-adapt constantly, and 0.8 on OLTP, which does not. - `ShadowSampleRate: 0.05` is a reasonable default. Higher rates cost more and buy no better ranking. -- Set `EvictPartialCapacityFilling: true` when W-TinyLFU or S3-FIFO is one of - the arms. The capacity gate compares `Len()` against `Cap()` for exact - equality, and neither arm holds that reliably: otter reports an approximate - size, and S3-FIFO can drop about a tenth of the cache in a single write. An - arm sitting below its capacity has its epochs skipped entirely. +- Set `EvictPartialCapacityFilling: true` when W-TinyLFU or S3-FIFO could be + the **active** arm. The gate reads that one policy and compares its `Len()` + against `Cap()` for exact equality, and neither of those two holds that + reliably: otter reports an approximate size, and S3-FIFO can drop about a + tenth of the cache in a single write. When it fires, the whole epoch is + skipped — no arm is reported and none is reset — so one under-filled active + policy suspends every measurement, not just its own. diff --git a/docs/evidence.md b/docs/evidence.md index 9c1e4d2..7b21ffe 100644 --- a/docs/evidence.md +++ b/docs/evidence.md @@ -43,9 +43,10 @@ TTL shares a column with LRU because these workloads carry no notion of staleness and its TTL is longer than any run, so it measures its LRU behaviour exactly -- identically, to the hundredth of a point, on all five. -Every arm here reproduces to the hundredth of a point between runs except -W-TinyLFU, which does not: on `loop` it has measured 88.7% and 94.5% within a -single process. Read its column, and any delta computed against it, with that +Every arm here reproduces to the hundredth of a point between runs except two. +W-TinyLFU does not: on `loop` it has measured 88.7% and 94.5% within a single +process. Neither does `Random`, which seeds itself from the global source — +that is what a control arm is for, but it means its column moves too. Read its column, and any delta computed against it, with that in mind — the cause is in [reproducible replays](benchmarking.md). Two things stand out. LRU and LFU both score **exactly zero** on `loop`, where a @@ -185,32 +186,31 @@ against the hit rate. The real-trace figures below are far tighter than this table, because those replays use a tuned 50ms epoch rather than the 2ms one held fixed across every workload here. -The timeline explains why. Replaying `phase-shift` and sampling `ActivePolicy()` -throughout: +The timeline says why, and it is not the answer this section used to give. +Replaying `phase-shift` and sampling `ActivePolicy()` throughout: ```text -phase Z------L------Z------L------Z------L------Z------L------ (Z = zipf, L = loop) -LRU ### -TwoQueue ##### -SIEVE ### -LFU ## ### ##### -S3FIFO ## ## -TinyLFU ############################################# - -share of time active: LRU 2%, LFU 6%, TwoQueue 3%, TinyLFU 82%, S3FIFO 2%, SIEVE 3% -hit rate 77.05% +share of time active: LRU 6%, LFU 4%, TwoQueue 15%, ARC 14%, TTL 4%, + TinyLFU 27%, S3FIFO 15%, SIEVE 16% +hit rate 63.77% ``` -The bandit works exactly as designed: it explores, identifies W-TinyLFU, and -holds it for 82% of the run. It does not oscillate at phase boundaries, because -there is no crossover to exploit -- W-TinyLFU is the best arm in *both* regimes. -The remaining gap is the price of exploring and of migrating between arms. - -That timeline is one draw, not a fixed result: this test uses a wall-clock -epoch, so the number of epochs in a run varies with machine load and which -also-ran arms collect a sliver of exploration varies with it. Across six -consecutive unmodified runs S3-FIFO appeared in four and SIEVE in five. Read -the 82% as the finding and the slivers as noise. +**The bandit does not settle.** Eight of the nine arms take a turn, the best of +them holds only 27% of the run, and the cache spends the rest of it changing +its mind. That is not a defect in the bandit; it is what honest evidence looks +like on this workload. `phase-shift` alternates between two regimes every +20,000 requests, several arms sit within a couple of points of each other in +both, and Thompson sampling explores exactly as it should when the posteriors +overlap. The cost of that exploration is the gap between 63.77% here and a +fixed W-TinyLFU's 82.6%. + +Two things are worth saying plainly about this block. Earlier versions of this +document showed W-TinyLFU holding 82-90% of the same run and concluded that the +bandit "identifies W-TinyLFU and holds it"; that was measured while the shadow +mechanism was reporting rivals at rates they could not achieve, and it does not +reproduce. And the run still uses a wall-clock epoch, so the number of epochs +varies with machine load — read the shape (no arm dominates) as the finding and +the individual percentages as one draw. So the case for this library is not "it beats the best policy." It is: @@ -317,11 +317,13 @@ Then there is `loop`, where it serves **0.00%** -- tied with LRU, LFU, 2Q, TTL, ARC and SIEVE, and beaten by random eviction. That is the algorithm behaving exactly as designed. `loop` cycles through 1011 keys with a 500-entry cache, so every reuse distance is 1011 requests. A key has to be requested again while it -is still in the small queue (50 entries) or still in the ghost queue to be -promoted. The library sizes that ghost queue at the *whole* cache capacity -- -500 here, not the main queue's 450 -- so the horizon is roughly 1000 requests -wide, and a reuse distance of 1011 falls just outside it. Every key goes round -the small queue forever and nothing is ever promoted. Only two arms survive the +is still resident or still in the ghost queue to be promoted. The library sizes +that ghost queue at the *whole* cache capacity -- 500 here -- so the horizon is +roughly 1000 requests wide, and a reuse distance of 1011 falls just outside it. +Nothing is ever promoted to the main queue, so the small queue holds the whole +cache and every key cycles through it forever. (`size/10` is not a cap on that +queue: upstream uses it only to decide which of the two queues an eviction +comes from.) Only two arms survive the trace at all: W-TinyLFU's sketch, which ages rather than expiring, and random eviction, which has no order to defeat. @@ -333,51 +335,23 @@ distance. ### What the library adapter costs -This arm wraps [scalalang2/golang-fifo](https://github.com/scalalang2/golang-fifo) -rather than implementing the algorithm here, and the wrapper is not free. The +These arms wrap [scalalang2/golang-fifo](https://github.com/scalalang2/golang-fifo) +rather than implementing the algorithms here, and the wrapper is not free. The adapter has to supply `Keys`, `Values`, `Resize` and `Cap`, none of which exist upstream, which means a second copy of the key set maintained through the -library's eviction callback and a full rebuild on every resize. Measured -against a from-scratch implementation of the same algorithm, replaying the same -traces at the same capacities. - -**Read the right-hand columns as history, not as a measurement you can repeat.** -The from-scratch implementation was written first, measured against the library -here, and then deleted when the library replaced it — so nothing in `make -evidence` reproduces that side of this table, and no test guards it. It is kept -because it is the evidence behind choosing the dependency. The `via the -library` columns are reproducible and are re-measured with everything else. - -| Trace | via the library | ns/op | from scratch | ns/op | -| --- | --- | --- | --- | --- | -| Twitter | 59.73% | 450 | 61.47% | 162 | -| Meta kvcache | 69.05% | 351 | 69.75% | 137 | -| ARC OLTP | 67.79% | 385 | 68.32% | 157 | -| ARC P3 | 10.75% | 709 | 9.98% | 291 | - -**Two to three times the per-operation cost**, and between half a point worse -and three quarters of a point better on hit rate. The time goes on the index -the adapter maintains alongside the library and on the library's linked-list -allocations; the hit-rate differences come from the two implementations sizing -the ghost queue differently and from upstream counting a write as an access. - -The trade bought is not owning an eviction algorithm. That is worth something: -the from-scratch version shipped with a real bug in its ghost queue, found only -by differential testing against a second implementation, and every property -test passed while it was there. - -LFU is the sharpest illustration of why synthetic workloads mislead. It is the -**best** policy on synthetic `zipf` (73.5%) and the **worst** on both large real -traces (41.4% on Twitter, 45.4% on OLTP). Synthetic Zipf holds popularity -*stationary*, which is exactly the assumption classic LFU makes; real traffic -shifts, and an entry that was hot once keeps a frequency count that holds it -resident long after it stops being useful. That is the failure W-TinyLFU's aged -frequency sketch exists to avoid, and it is invisible until you replay real -traffic. - -How much configuration moves these numbers is in -[configuration](configuration.md#tuning-measured); the short version is that a -too-short epoch costs up to 7 points and 30x the per-operation time. +library's eviction callback and a full rebuild on every resize. The per-operation +columns in the table above are what that costs: 371 to 770 ns/op for S3-FIFO +against 80 to 130 for a bare LRU on the same traces. + +The trade bought is not owning an eviction algorithm, and that is worth +something. An earlier from-scratch S3-FIFO in this repository shipped with a +real bug in its ghost queue -- the paper's virtual-timestamp approximation +leaves a dead slot behind whenever an entry is removed early, so the queue +steadily held fewer keys than its capacity claimed and dropped them just before +they came back. Every property test passed; only a differential run against a +second implementation found it. That implementation is gone, so its numbers are +not quoted here: nothing in `make evidence` reproduces them and no test guards +them. ## Does sampling distort the comparison? @@ -437,8 +411,9 @@ effective rate rather than letting a miniature shrink into noise. ## Does pooling across a fleet help? **Only in the regime it was built for, and it is worth checking you are in that -regime before turning it on.** All figures are 8 replicas, cache capacity 300 -to 500, `make evidence`. The mechanism is described in [running a +regime before turning it on.** Figures are 8 replicas at cache capacity 300 to +500 unless the row says otherwise -- the mixed-fleet row below is 6 replicas at +400 -- and all come from `make evidence`. The mechanism is described in [running a fleet](fleet.md). The case it exists for is a replica that sees too little traffic per epoch to