Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
51 changes: 21 additions & 30 deletions cache.go
Original file line number Diff line number Diff line change
Expand Up @@ -73,12 +73,11 @@ type AdaptiveCache[K comparable, V any] struct {
// after each report, so an answer about the traffic has to be accumulated
// somewhere.
//
// It is cleared for both policies involved in a switch. Pooling a policy's
// active tenure with its shadow tenure would mix two different measurement
// regimes - full capacity over all traffic against miniature capacity over
// a sample - and, worse, would leave the just-demoted policy's long good
// history outweighing the promoted one's short history, so Advice would
// recommend reverting a switch the cache had just made correctly.
// It is cleared for both policies in a switch. Pooling a policy's active
// tenure with its shadow tenure mixes full capacity over all traffic with
// a miniature over a sample, and leaves the demoted policy's long history
// outweighing the promoted one's short one -- so Advice would recommend
// reverting a switch the cache had just made correctly.
tenureStats map[PolicyType]PolicyStats

// reportingEpochs counts only the epochs that actually measured something.
Expand Down Expand Up @@ -125,10 +124,9 @@ func (c *AdaptiveCache[K, V]) recordActiveSample(sampled, hit bool) {
// Get returns the value stored for key by the active policy, feeding the same
// lookup to every shadow policy that samples the key.
//
// When Settings.EpochRequests is set, the call that completes an epoch runs it
// here, after every lock this method took has been released - runEpoch needs
// the write lock, and a Get still holding the read lock would deadlock against
// it.
// With Settings.EpochRequests set, the call completing an epoch runs it here,
// after every lock this method took is released: runEpoch needs the write lock
// and would deadlock against a Get still holding the read lock.
func (c *AdaptiveCache[K, V]) Get(key K) (V, bool) {
value, found := c.get(key)
c.countRequest()
Expand All @@ -153,12 +151,10 @@ func (c *AdaptiveCache[K, V]) get(key K) (V, bool) {
}
c.mu.RUnlock()

// Gradual migration window: resolve the whole lookup under the write lock,
// promoting an eligible key into the active policy BEFORE its Get is
// counted. The active policy then records a hit for a request the cache
// serves; promoting after the Get would leave a spurious miss in the
// active arm's stats for a served request, skewing both Stats() and the
// bandit's posterior toward the demoted policy.
// Gradual window: resolve the lookup under the write lock, promoting an
// eligible key BEFORE its Get is counted. Promoting after would record a
// miss for a request the cache served, skewing Stats() and the bandit's
// posterior toward the demoted policy.
c.mu.Lock()
defer c.mu.Unlock()

Expand Down Expand Up @@ -187,16 +183,11 @@ func (c *AdaptiveCache[K, V]) Add(key K, value V) bool {
continue
}

// Only a key the shadow does not already hold. A shadow's value is
// always the zero value, so re-adding a key it has carries no
// information - but it is not free: for a policy whose eviction
// state is a counter or a single bit, a write counts as an access.
// SIEVE would mark every freshly filled key as visited, defeating
// exactly the one-hit-wonder filtering it is carried for, and
// S3-FIFO's counter would run ahead of the algorithm.
//
// Peek rather than Contains or Get, because it must not disturb
// that state either.
// Only a key the shadow does not hold: its value is always zero, so
// re-adding carries no information but does count as an access for
// a policy whose eviction state is a counter or a bit. SIEVE would
// mark every filled key visited, defeating the one-hit-wonder
// filtering it is carried for. Peek, so the check disturbs nothing.
if _, held := policy.Peek(key); held {
continue
}
Expand Down Expand Up @@ -275,10 +266,10 @@ func (c *AdaptiveCache[K, V]) Purge() {
// miniature capacity that corresponds to size rather than to size itself, so
// they stay faithful simulations of a cache of the requested capacity.
//
// The sample rate itself is fixed for the life of the cache: changing it would
// change which keys are sampled, invalidating every shadow's accumulated state.
// The miniature capacity therefore follows the rate directly here, without the
// MinShadowCapacity floor that construction applies - see scaledCapacity.
// The sample rate is fixed for the life of the cache -- changing it would
// change which keys are sampled and invalidate every shadow's state -- so the
// miniature capacity follows the rate directly, without the MinShadowCapacity
// floor construction applies. See scaledCapacity.
func (c *AdaptiveCache[K, V]) Resize(size int) int {
c.mu.Lock()
defer c.mu.Unlock()
Expand Down
10 changes: 9 additions & 1 deletion docs/configuration.md
Original file line number Diff line number Diff line change
Expand Up @@ -90,7 +90,10 @@ is close to free. That is what makes carrying nine arms practical.

Reproduce with `go test -run '^$' -bench . -benchtime=300ms .`

Sampling is off by default. Very small caches disable it automatically, since a
Sampling is off by default. `MinShadowCapacity` (256 unless you set it) is the
floor on a miniature: when the rate would shrink a shadow below it, the
*effective rate* is raised rather than the capacity alone, and on a cache small
enough that the floor exceeds its nominal size, sampling disables itself. A
miniature of a handful of entries measures noise rather than a policy.

Sampling does not distort which policy wins — that was measured directly, see
Expand All @@ -111,6 +114,11 @@ migration. Three settings damp that, all inactive at their zero value:
}
```

`MinEpochRequests` counts the requests **the bandit sees**, which under
`ShadowSampleRate` are sampled requests: at a rate of 0.05 a threshold of 100
is reached after roughly 2000 real ones. Set it against the sampled stream, not
against your traffic.

## Tuning, measured

The epoch duration is the setting that matters most, and the failure mode is
Expand Down
11 changes: 8 additions & 3 deletions docs/design.md
Original file line number Diff line number Diff line change
Expand Up @@ -59,9 +59,14 @@ On each request:
anything: a read-through caller only calls `Add` when the *active* policy
missed, so without it a shadow could never acquire a key the incumbent was
already serving, and the better the incumbent performed the less its rivals
were allowed to learn. Shadows hold keys and eviction bookkeeping, never
data, which is why N policies do not cost N times the memory — and why no
caller can ever be handed a shadow's zero.
were allowed to learn. The drift that causes is not a small bias: measured
on a cyclic workload behind a 94%-hit incumbent, arms that truly serve 0.00%
reported over 90%, because a starved shadow's contents go static and a
static cache covering most of a small keyspace looks excellent. Its sign
depends on which arm is incumbent, so it does not cancel — `Advice()`
recommended switching from the best arm to the worst. Shadows hold keys and
eviction bookkeeping, never data, which is why N policies do not cost N
times the memory — and why no caller can ever be handed a shadow's zero.

Then once per epoch:

Expand Down
33 changes: 33 additions & 0 deletions docs/policies.md
Original file line number Diff line number Diff line change
Expand Up @@ -208,6 +208,21 @@ fan-out skips a key the shadow already holds, so the counters no longer run
ahead on shadow duty. It still applies to a caller whose own traffic rewrites
live keys.

**Demotion disturbs their eviction state more than it disturbs the others'.**
When a policy stops being active it is rewritten to zero values in `Keys()`
order, which for a recency policy re-establishes the same order and for a
frequency policy adds one access to every surviving key, leaving the relative
order alone. Neither holds here. SIEVE treats a write as setting the visited
bit, and that bit is its whole eviction criterion, so rewriting every key sets
it on every key and erases the ordering rather than preserving it. S3-FIFO's
counter saturates at three, so a key already at the cap gains nothing while a
key at zero gains one, compressing the ordering instead of shifting it
uniformly. The effect is a bias in the demoted policy's first shadow epochs
rather than a standing loss — the queues are untouched and ordinary traffic
rewrites the bits soon after — but a policy whose eviction state is a single
saturating bit per entry should not be demoted this way without measuring what
it costs.

Two smaller notes. The adapter always builds with a TTL of zero, which is
load-bearing: a non-zero TTL starts a background goroutine that would invoke
the eviction callback from a goroutine the adapter never entered, and the
Expand Down Expand Up @@ -243,3 +258,21 @@ Note that `Resize` on an adapted cache rebuilds it, discarding whatever
adaptation the algorithm had learned. `AdaptiveCache` resizes shadow policies
when its own capacity changes, so adapted policies are heavier arms to carry
than natively resizable ones.

Three rules an arm has to honour, each of which a real implementation has
broken here:

- **Never return a zero value with `true`.** A `Get` or `Peek` that reports a
hit for an entry it no longer holds hands the caller a value nobody stored.
This library's central invariant is that a shadow's zero is never observable,
and one arm returning `(zeroValue, true)` for an expired-but-unreaped entry
defeats it from below. `hashicorp/golang-lru/v2/expirable` does exactly that,
which is why `policies.NewTTL` is written over a plain LRU with lazy expiry
rather than wrapping it.
- **`Keys()` and `Values()` must line up.** A `Values()` padded to full length
with trailing zeros does not correspond to `Keys()`, and warm migration
copies through both.
- **Size 0 means empty, not unlimited.** Shadows are resized automatically, so
an arm reading 0 as "no limit" turns a bounded miniature into an unbounded
cache. It must also not start a goroutine it gives you no way to stop: a
reaper per cache with no `Close` leaks the goroutine and the cache with it.
35 changes: 17 additions & 18 deletions epoch.go
Original file line number Diff line number Diff line change
Expand Up @@ -25,15 +25,13 @@ func (c *AdaptiveCache[K, V]) runAdaptiveSelect() {
}

// countRequest advances the request-driven epoch clock and runs the epoch on
// the call that completes it.
// the call that completes it. Caller must hold no lock: runEpoch takes the
// write lock.
//
// It must be called with no lock held: runEpoch takes the write lock.
//
// Exactly one caller per epoch observes the count equal to the limit, so
// exactly one epoch runs however many goroutines are in Get at once. The limit
// is then subtracted rather than the counter reset, so requests that arrived
// during the crossing are still counted towards the next epoch instead of
// being dropped.
// Exactly one caller per epoch sees the count equal the limit, so exactly one
// epoch runs however many goroutines are in Get. The limit is subtracted
// rather than the counter reset, so requests arriving mid-crossing still
// count towards the next epoch.
func (c *AdaptiveCache[K, V]) countRequest() {
limit := c.settings.EpochRequests
if limit <= 0 {
Expand Down Expand Up @@ -108,16 +106,17 @@ func (c *AdaptiveCache[K, V]) tryChangePolicy() PolicyType {
return c.selectPolicyLocked()
}

// selectPolicyLocked reports every policy's stats to the bandit — the active
// policy included, so its posterior does not go stale — and returns the
// bandit's chosen policy for the next epoch. When
// EvictPartialCapacityFilling is false and the active policy is not yet full,
// it returns early without reporting or resetting anything; counters then
// accumulate until the next reporting epoch. On a reporting epoch counters
// are reset after delivery; the active policy's counts are folded into
// globalStats first so Stats() stays cumulative and no active-tenure counts
// leak into a policy's first shadow epoch after demotion. It must be called
// while the write lock is held.
// selectPolicyLocked reports every policy's stats to the bandit -- the active
// policy included, so its posterior does not go stale -- and returns the arm
// chosen for the next epoch.
//
// With EvictPartialCapacityFilling false and the active policy not yet full it
// returns early, reporting and resetting nothing; counters accumulate until
// the next reporting epoch. Otherwise counters reset after delivery, the
// active policy's folded into globalStats first so Stats() stays cumulative
// and no active-tenure count leaks into a first shadow epoch after demotion.
//
// Caller must hold the write lock.
func (c *AdaptiveCache[K, V]) selectPolicyLocked() PolicyType {
currentPolicy := c.activePolicy

Expand Down
14 changes: 5 additions & 9 deletions migration.go
Original file line number Diff line number Diff line change
Expand Up @@ -113,15 +113,11 @@ func (c *AdaptiveCache[K, V]) drainOneKey() {
// (promotion here or an earlier Remove). It must be called while the write
// lock is held during a gradual migration window.
//
// A note on why the source is trustworthy here. It is not the active policy,
// so anything walking c.policies and skipping only activePolicy would treat it
// as a shadow and fill it with zero values - and the Peek below cannot tell
// such a zero from a real value still pending, so it would promote the zero
// and serve it to a caller as a hit. fanOutReadLocked therefore skips the
// source while a window is open. "Not active" is not the same as "is a
// shadow": for the duration of a gradual window there are three roles, not
// two, and anything iterating the policies has to say what it means to do to
// this one.
// The Peek below cannot tell a zero written by a shadow fill from a real value
// still pending, so fanOutReadLocked skips the source while a window is open.
// "Not active" is not "is a shadow": during a gradual window there are three
// roles, and anything iterating c.policies has to say what it does to this
// one.
func (c *AdaptiveCache[K, V]) promoteLocked(key K) {
// Skip keys the caller has since written directly, or already promoted.
if _, ok := c.migrationRealKeys[key]; ok {
Expand Down
60 changes: 24 additions & 36 deletions sampling.go
Original file line number Diff line number Diff line change
Expand Up @@ -11,25 +11,19 @@ import (
const maxUint64AsFloat = float64(1 << 64)

// keySampler decides whether a key belongs to the deterministic subset of the
// keyspace that shadow policies track. Sampling lets a shadow estimate its hit
// rate from a small fraction of traffic instead of mirroring every operation.
// keyspace that shadow policies track.
//
// The decision is a pure function of the key and the seed, so a given key is
// either always sampled or never sampled for the lifetime of the sampler. That
// matters twice over: a sampled shadow sees a coherent access pattern for the
// keys it does track (rather than a random scatter that would destroy any
// notion of reuse), and every shadow sharing one sampler measures the same
// sub-workload, which is what makes their hit rates comparable to each other.
// The decision is a pure function of key and seed, so a key is either always
// sampled or never sampled. That matters twice: a shadow sees a coherent
// access pattern for the keys it tracks rather than a random scatter with no
// reuse, and every shadow sharing one sampler measures the same sub-workload,
// which is what makes their hit rates comparable.
//
// The seed is drawn per cache rather than fixed, so the sampled subset differs
// between processes and cannot be predicted or targeted by a caller.
// The seed is per cache, so the subset cannot be predicted or targeted.
//
// Sampled counts are never scaled back up to full-traffic magnitude before
// reaching the bandit. Scaling would restore magnitude while inventing
// confidence, handing a Beta posterior twenty times the evidence that was
// actually collected. Instead every arm, the active policy included, is
// measured over this same sampled substream, so the arms carry equal and
// honest evidence and remain directly comparable.
// Sampled counts are never scaled back up before reaching the bandit: that
// would restore magnitude while inventing confidence. Every arm, the active
// policy included, is measured over this same substream instead.
type keySampler[K comparable] struct {
seed maphash.Seed
// threshold is the exclusive upper bound on a key's hash for it to be in
Expand Down Expand Up @@ -70,18 +64,14 @@ func (s *keySampler[K]) sampled(key K) bool {
}

// scaledCapacity returns the miniature capacity corresponding to sampling rate
// of a cache of the given size, holding the identity capacity/size == rate.
// of a cache of the given size, holding capacity/size == rate.
//
// Unlike shadowCapacity it applies no floor. The floor exists to stop a cache
// from being built with a miniature too small to measure, and it works by
// raising the sample rate to match. After construction the rate can no longer
// move, so applying the floor alone would leave shadows running at a capacity
// larger than their share of the traffic - and a shadow of capacity C fed an
// r-sampled stream simulates a cache of C/r. Every shadow would then simulate a
// larger cache than the active policy actually is and report a better hit rate
// for that reason alone, which is a systematic bias against whichever policy is
// active. A miniature that is merely small is noisy; one that is inconsistent
// with its rate is wrong, so the identity wins.
// Unlike shadowCapacity it applies no floor. A shadow of capacity C fed an
// r-sampled stream simulates a cache of C/r, so raising the capacity without
// raising the rate would have every shadow simulate a larger cache than the
// active policy is and report a better hit rate for that reason alone. A
// miniature that is merely small is noisy; one inconsistent with its rate is
// wrong.
func scaledCapacity(size int, rate float64) int {
if size <= 0 || rate >= 1 {
return size
Expand All @@ -98,16 +88,14 @@ func scaledCapacity(size int, rate float64) int {
return capacity
}

// shadowCapacity returns the capacity a shadow policy should run at to
// simulate a full-size cache of nominalCap over the sampled substream, and the
// effective rate that capacity corresponds to.
// shadowCapacity returns the capacity a shadow should run at to simulate a
// full-size cache of nominalCap over the sampled substream, and the effective
// rate that capacity corresponds to.
//
// A cache of capacity rate*N fed an rate-sampled stream approximates a cache
// of capacity N fed the full stream, so the capacity has to shrink with the
// rate for the estimate to mean anything. A floor guards the degenerate end:
// a five-entry miniature measures noise, so when rate*nominalCap falls below
// minCapacity the rate itself is raised (not just the capacity) to keep the
// simulation identity intact, up to the point where sampling disables itself.
// The capacity shrinks with the rate for the estimate to mean anything. When
// rate*nominalCap falls below minCapacity the rate itself is raised, not just
// the capacity, keeping the identity intact -- up to the point where sampling
// disables itself.
func shadowCapacity(nominalCap int, rate float64, minCapacity int) (capacity int, effectiveRate float64) {
if nominalCap <= 0 || rate >= 1 {
return nominalCap, 1
Expand Down
Loading
Loading