Skip to content

refactor(llm)!: remove stale duplicate of the harness response cache - #65

Merged
senamakel merged 4 commits into
mainfrom
tinyinference-drop-stale-cache
Oct 7, 2026
Merged

senamakel merged 4 commits into
mainfrom
tinyinference-drop-stale-cache

Conversation

@senamakel

@senamakel senamakel commented Oct 7, 2026 •

Copy link
Copy Markdown
Member

Summary

tinyinference_llm::cache carried an old copy of the tinyagents harness response cache. Its module doc still opened with "Harness cache module". The harness copy (tinyagents-harness/src/cache/) has since moved on to scoped keys, a SQLite store and singleflight, and it is the one hosts actually use. This PR deletes the stale copy and keeps CachePolicy, which ModelRequest::cache_policy and the Anthropic provider depend on.

Removed (about 430 lines of source and 249 lines of tests):

  • cache_key(&ModelRequest) -> String
  • the ResponseCache trait, InMemoryResponseCache (with DEFAULT_CAPACITY, new, with_capacity) and the crate-private LruResponseMap
  • PromptCacheLayout (from_request, prefix_ids, fingerprint, is_prefix_stable_against) and CacheLayoutEvent
  • the private helpers hex_digest, fold_canonical and fnv1a_hex, plus cache/types.rs
  • the sha2 dependency of tinyinference-llm, which nothing else in the crate used (tinyinference-providers still uses it)

Kept, at the same paths:

Related issue

None.

API or behavior changes

This is a breaking change: the public items listed above are removed. No runtime behavior changes.

Before deleting anything, I checked usage with grep across this repo, libraries/tinyagents (main, its worktrees and its vendored copies), the openhuman superproject (crates/ and every vendor/ submodule, recursively), and every other libraries/* repo:

  • The only items imported through tinyinference_llm::cache:: are CachePolicy (plus intra-doc links to CachePolicy::protect_prompt_prefix) and canonical_value (the pending Export replace_text_blocks, canonical_value and the JSON schema validator #64 / tinyagents dedupe-clones pair).
  • No file imports cache_key, ResponseCache, InMemoryResponseCache, PromptCacheLayout or CacheLayoutEvent from tinyinference. Every hit for those names resolves to tinyagents_harness::cache. For example, openhuman-core imports tinyagents_harness::cache::{InMemoryResponseCache, CacheLayoutEvent}.
  • tinyagents-harness re-exports the whole tinyinference_llm crate, but nothing reaches these items through that re-export.

Validation

All commands were run in this branch at stable 1.99.0:

  • cargo fmt --all -- --check: clean
  • cargo +stable clippy --all-targets --all-features -- -D warnings: clean
  • cargo build --all-targets --all-features: covered by the clippy and test builds
  • cargo test --workspace --all-features: 1019 passed, 0 failed
  • RUSTDOCFLAGS="-D warnings" cargo doc -p tinyinference-llm --no-deps: clean, with no broken intra-doc links

As a downstream check, I made a throwaway checkout of tinyagents main (2cc39826) with vendor/tinyinference pointed at this branch, and it was deleted afterwards:

  • cargo check -p tinyagents-harness -p tinyagents-integration-tests --all-targets --all-features: clean
  • the harness cache lib tests passed (57), as did wave2_cache_store, wave2_cache_layout and wave2_cache_key_scope (7, 11 and 14). wave2_cache_store exercises the CachePolicy builders.

Tests

The old cache/mod_tests.rs only tested the removed items, so I replaced it with four tests:

  • CachePolicy defaults
  • the builder methods
  • a serde round-trip that leaves out unset options
  • canonical_value sorting keys at every depth

Documentation

The module doc now says what the module is for: carrying the per-request policy, while the host harness owns the response cache. I also added a CHANGELOG entry under Unreleased, in a new "Breaking changes" section.

Checklist

  • The change is focused on one logical change
  • No new #[allow(...)], #[ignore], or relaxed lints
  • No secrets, tokens, or .env contents in the diff or the description

Summary by CodeRabbit

  • Breaking Changes
    • Removed outdated response-cache interfaces and implementations from the LLM crate. The maintained cache implementations remain available through the harness crate.
  • Improvements
    • Cache policy options remain available, including response caching, prompt-prefix protection, optional expiration, and namespaces.
    • JSON value canonicalization is now publicly accessible.
  • Dependencies
    • The LLM crate no longer depends on SHA-256 hashing.

senamakel and others added 3 commits October 7, 2026 07:58
The cache module now only carries the per-request CachePolicy and the
canonical_value JSON helper, since the response cache, its keys and the
prompt-layout guard live in the host harness. canonical_value is exported so
host-derived cache keys canonicalize identically, and the sha2 dependency is
dropped along with the removed hashing code.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Document the breaking removal of the duplicated harness response cache from
tinyinference_llm::cache, since the maintained versions live in
tinyagents-harness and nothing used the stale copies. Also note that
canonical_value is now public and the crate no longer depends on sha2.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Remove the sha2 entry from the dependency list in Cargo.lock, reflecting that the crate is no longer a direct dependency.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
@coderabbitai

coderabbitai Bot commented Oct 7, 2026 •

Copy link
Copy Markdown

Review in Change Stack →

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration
  • Configuration used: Organization UI
  • Review profile: CHILL
  • Plan: Advanced
  • Run ID: a3daf5cd-b3c7-4cfe-9937-0b1f7146240a
📥 Commits

Reviewing files that changed from the base of the PR and between a32eec8 and 715080a.

⛔ Files ignored due to path filters (1)
  • Cargo.lock is excluded by !**/*.lock
📒 Files selected for processing (5)
  • CHANGELOG.md
  • crates/tinyinference-llm/Cargo.toml
  • crates/tinyinference-llm/src/cache/mod.rs
  • crates/tinyinference-llm/src/cache/mod_tests.rs
  • crates/tinyinference-llm/src/cache/types.rs
💤 Files with no reviewable changes (2)
  • crates/tinyinference-llm/Cargo.toml
  • crates/tinyinference-llm/src/cache/types.rs

Included review availability: This review used your included allowance. Your plan provides up to 1 included review per hour; 0 remain after this review.


📝 Walkthrough

Walkthrough

The cache module retains CachePolicy and canonical_value. It removes response-cache and prompt-layout types and implementations. The crate also removes its sha2 dependency and updates cache tests and the changelog.

Changes

Cache API surface

Layer / File(s) Summary
Retain policy and remove cache implementations
crates/tinyinference-llm/src/cache/*, crates/tinyinference-llm/Cargo.toml, CHANGELOG.md
CachePolicy remains in the cache module, and canonical_value remains available. The module removes the response-cache and prompt-layout types and implementations. Tests cover policy defaults, builder options, serde behavior, and canonicalization. The manifest removes sha2, and the changelog documents the API changes.

Priority: ➖ Normal

Estimated code review effort: 2 (Simple) | ~12 minutes

Change: Refactor

Merge Risk: ⚪ Minimal · up to 71508

The cache API removal has no established merge-blocking issue; it is ready for normal merge checks.

Security Architecture Review

Security architecture risk: 🔵 Low · up to 71508

The retained request-policy behavior is unchanged, and no new security bypass was established. Residual risk concerns downstream compatibility and cache enforcement, which could not be independently verified.

Retained concerns
No architecture-level concerns identified.

Security review details

Security Blast Radius

  • inferred — The established change surface is a library API contraction affecting dependent hosts. canonical_value neither performs I/O nor accesses cache state, so its unchanged implementation does not establish a new path to privileged execution or cross-tenant data access.

Trust Boundaries and Controls

  • observed — The local prompt-cache policy gate is unchanged: requests require a cacheable segment, and an explicit protect_prompt_prefix=false vetoes breakpoints. The Anthropic test asserts that this veto suppresses the cache-control marker; this does not verify downstream response-cache isolation or expiry.
🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title clearly identifies the breaking removal of the stale response-cache API duplicate, which is the primary change.
Docstring Coverage ✅ Passed Docstring coverage is 100.00% which is sufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 9 functions across 2 files. (1 skipped: 1 …
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
✨ Finishing Touches
📝 Generate docstrings
  • Commit to this branch
  • Create a new PR
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

I’m a rabbit with a policy to keep,
While old cache types hop off to sleep.
Nested keys line up in a row,
No sha2 seeds left here to sow.
I nibble the changelog, then bound away!

Comment @coderabbitai help to get the list of available commands.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 52278b0e37

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment on lines +8 to +9
//! cache markers; the response cache itself, its keys and the prompt-layout
//! guard live in the host harness (`tinyagents-harness`'s `cache` module).

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Keep provider-neutral response-cache APIs in TinyInference

When a consumer other than TinyAgents needs response caching, moving ResponseCache, cache-key construction, and the default in-memory implementation exclusively into tinyagents-harness forces that consumer either to depend on an agent runtime or to duplicate the contract. The repository boundary explicitly assigns provider-neutral cache APIs to TinyInference while reserving agent/runtime concerns for consumers, so retain the neutral trait and key API here while leaving harness-specific SQLite and singleflight implementations downstream.

AGENTS.md reference: AGENTS.md:L5-L8

Useful? React with 👍 / 👎.

@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Oct 7, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-10-07T06:12:41.626737Z 715080a New commits
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@tinysweeper

tinysweeper Bot commented Oct 7, 2026 •

Copy link
Copy Markdown

Tiny Sweeper review

Tiny Sweeper reviewed this change across 6 lane(s) and found 2 active actionable finding(s). Detailed lane evidence and any incomplete work are listed below.

State: Ready for maintainer review
Priority: medium
Reviewed head: 715080a72224
Updated: 1791354241 (Unix time)

Review snapshot

Change surface Files Review signal Count
Production 2 Active findings 8
Tests 1 Noted findings 0
Documentation 1 Resolved findings 0
Configuration 1 Pending checks/questions 0

Completeness: Complete
Test assessment: No supported feature-to-test mapping was available; this does not mean tests are absent or passed.

What changed

The review could not produce a supported behavioral summary; inspect the cited changed surface and lane details below.

Features

None identified with supported citations.

Tests

No supported feature-to-test mapping was produced. Test execution is not inferred.

Findings

  • medium · critique · Apply defaults when deserializing incomplete policies — The boolean fields do not have `#[serde(default)]`, so deserializing a policy that omits either flag fails instead of using the documented safe defaults. For example, `serde_json:: (crates/tinyinference\-llm/src/cache/mod\.rs:26)
  • medium · critique · Handle durations larger than the TTL field can represent — `Duration::as_millis()` returns `u128`, but this cast stores only `u64`. A duration larger than `u64::MAX` milliseconds is silently saturated/reduced to the maximum representable T (crates/tinyinference\-llm/src/cache/mod\.rs:55)
  • medium · security · Apply defaults when deserializing incomplete policies — The policy documents both flags as defaulting to `false`, but neither field has `#[serde(default)]`. Deserializing a persisted or externally supplied policy that omits either flag (crates/tinyinference\-llm/src/cache/mod\.rs:26)
  • medium · security · Handle durations larger than the TTL field can represent — A `Duration` can contain more milliseconds than `u64` can represent, while `ttl_ms` cannot. The cast silently truncates the `u128` millisecond count, so a very large requested TTL (crates/tinyinference\-llm/src/cache/mod\.rs:55)
  • medium · tests · Apply defaults when deserializing incomplete policies — Deserializing a partial `CachePolicy` — e.g. `{"response_cache_enabled": true}` from a config file missing the newer keys — fails outright, because the `bool` fields have no `#[ser (crates/tinyinference\-llm/src/cache/mod\.rs:22)
  • medium · tests · Handle durations larger than the TTL field can represent — `with_ttl` stores `ttl.as_millis() as u64`; a `Duration` whose millisecond count exceeds `u64::MAX` silently truncates (in practice `as_millis` itself saturates above ~584 million (crates/tinyinference\-llm/src/cache/mod\.rs:55)
  • medium · description · Apply defaults when deserializing incomplete policies — `CachePolicy` derives `Deserialize`, but only `ttl_ms` and `namespace` carry `#[serde(default)]`. A JSON object that omits `response_cache_enabled` or `protect_prompt_prefix` — pla (\(pull request description\))
  • medium · description · Handle durations larger than the TTL field can represent — `ttl.as_millis()` returns a `u128`; the `as u64` cast silently truncates durations above roughly 49 days, and in a release build a caller passing e.g. `Duration::from_secs(u64::MAX (\(pull request description\))

Before merge

None.

How this fits together

flowchart LR
  n0["with_capacity"]:::impacted
  n1["new"]:::impacted
  n2["put"]:::impacted
  n3["cache_key"]:::impacted
  n4["response_cache_capacity_zero_retains_last"]:::impacted
  n0 -->|calls| n1
  n1 -->|calls| n0
  n3 -->|calls| n1
  n4 -->|calls| n0
  n4 -->|tests| n0
  n4 -->|calls| n2
  n4 -->|tests| n2
  classDef changed fill:#0d4429,stroke:#238636,color:#e6edf3
  classDef impacted fill:#161b22,stroke:#6e7681,color:#c9d1d9
  classDef flagged fill:#5a1e02,stroke:#d93f0b,color:#ffffff
  classDef blocking fill:#67060c,stroke:#f85149,color:#ffffff
Loading
Agent review details

critique

  • Conclusion: Success
  • Scope reviewed: all assigned evidence
  • Lane summary: The cache policy refactor is not safe to merge because incomplete serialized policies still fail to deserialize and oversized TTL durations are silently reduced to a representable maximum. (2 observation(s) grouped into shared inline comments) _Code retrieval was unavailable (model: ladder embeddings returned 400 Bad Request: {"error":{"message":"unknown ladder vectors; known ladders are flash (also chat-v1, flash-v1), instant (also no-think, instant-v1), reasoning (also deepseek), max-reasoning (also max-reasoning-v1), deepseek-flash (also reasoning-v1, agentic-v1), deep (also luna), scribe, uncensored, vectors-oai3 (also embeddings-oai3-v1), vision (also vision-v1, multimodal-v1), image (also images-v1, image-v1), vi), so this review saw the diff alone._ _Memory was unavailable (model: cortex: v1/recall: error sending request for url (http://cortexdb:3141/v1/recall\)\), so this review ran without it._
  • Evidence: crates/tinyinference\-llm/src/cache/mod\.rs — Apply defaults when deserializing incomplete policies
  • Evidence: crates/tinyinference\-llm/src/cache/mod\.rs — Handle durations larger than the TTL field can represent

security

  • Conclusion: Success
  • Scope reviewed: all assigned evidence
  • Lane summary: The cache-policy refactor is narrowly scoped but still needs fixes before it can be merged. _Code retrieval was unavailable (model: ladder embeddings returned 400 Bad Request: {"error":{"message":"unknown ladder vectors; known ladders are flash (also chat-v1, flash-v1), instant (also no-think, instant-v1), reasoning (also deepseek), max-reasoning (also max-reasoning-v1), deepseek-flash (also reasoning-v1, agentic-v1), deep (also luna), scribe, uncensored, vectors-oai3 (also embeddings-oai3-v1), vision (also vision-v1, multimodal-v1), image (also images-v1, image-v1), vi), so this review saw the diff alone._ _Memory was unavailable (model: cortex: v1/recall: error sending request for url (http://cortexdb:3141/v1/recall\)\), so this review ran without it._
  • Evidence: crates/tinyinference\-llm/src/cache/mod\.rs — Apply defaults when deserializing incomplete policies
  • Evidence: crates/tinyinference\-llm/src/cache/mod\.rs — Handle durations larger than the TTL field can represent

tests

  • Conclusion: Success
  • Scope reviewed: all assigned evidence
  • Lane summary: This pass only deletes the stale response-cache/layout types moved to the harness, plus their tests; the retained `CachePolicy` and `canonical_value` code is already covered by `mod_tests.rs`. The two findings from the earlier review both still stand unchanged in the kept code. (1 observation(s) grouped into shared inline comments) _Code retrieval was unavailable (model: ladder embeddings returned 400 Bad Request: {"error":{"message":"unknown ladder vectors; known ladders are flash (also chat-v1, flash-v1), instant (also no-think, instant-v1), reasoning (also deepseek), max-reasoning (also max-reasoning-v1), deepseek-flash (also reasoning-v1, agentic-v1), deep (also luna), scribe, uncensored, vectors-oai3 (also embeddings-oai3-v1), vision (also vision-v1, multimodal-v1), image (also images-v1, image-v1), vi), so this review saw the diff alone._ _Memory was unavailable (model: cortex: v1/recall: error sending request for url (http://cortexdb:3141/v1/recall\)\), so this review ran without it._
  • Evidence: crates/tinyinference\-llm/src/cache/mod\.rs — Apply defaults when deserializing incomplete policies
  • Evidence: crates/tinyinference\-llm/src/cache/mod\.rs — Handle durations larger than the TTL field can represent

commits

  • Conclusion: Neutral
  • Scope reviewed: all assigned evidence
  • Lane summary: Nothing sensitive found in what this pull request commits.

description

  • Conclusion: Success
  • Scope reviewed: all assigned evidence
  • Lane summary: The deletion is accurately described: the removed items, the kept `CachePolicy`, the newly public `canonical_value`, and the `sha2` dependency removal all match the diff, and the replacement tests cover what remains. Both previously raised findings still stand because the `CachePolicy` code was kept unchanged: deserialization still fails when the bool fields are absent, and `with_ttl` still truncates large durations via `as u64`. Nothing new was introduced by the final revision. _Code retrieval was unavailable (model: ladder embeddings returned 400 Bad Request: {"error":{"message":"unknown ladder vectors; known ladders are flash (also chat-v1, flash-v1), instant (also no-think, instant-v1), reasoning (also deepseek), max-reasoning (also max-reasoning-v1), deepseek-flash (also reasoning-v1, agentic-v1), deep (also luna), scribe, uncensored, vectors-oai3 (also embeddings-oai3-v1), vision (also vision-v1, multimodal-v1), image (also images-v1, image-v1), vi), so this review saw the diff alone._ _Memory was unavailable (model: cortex: v1/recall: error sending request for url (http://cortexdb:3141/v1/recall\)\), so this review ran without it._
  • Evidence: \(pull request description\) — Apply defaults when deserializing incomplete policies
  • Evidence: \(pull request description\) — Handle durations larger than the TTL field can represent

e2e

  • Conclusion: Neutral
  • Scope reviewed: all assigned evidence
  • Lane summary: No end-to-end harness in this repository: no e2e test files and no e2e workflow.
Evidence and run details
  • Models: gpt-5.6-luna, glm-5.3-flash
  • Spend: $0.005510
  • Tokens: 129118 input · 9466 output · 12158 cached · 0 embedding
Head State Pass summary
52278b0e37c8 ready for maintainer review 2 active finding(s), 0 resolved finding(s) (at 1791350119)
715080a72224 ready for maintainer review 8 active finding(s), 0 resolved finding(s) (at 1791354241)

tinysweeper 0.1.0

@tinysweeper tinysweeper Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

tinysweeper found nothing blocking. Approving.

             $0.0028 · 152,670 in / 7,891 out · 13,443 cached (9%) · gpt-5.6-luna, glm-5.3-flash
critique:    $0.0014 · 68,391 in  / 4,683 out · 6,349 cached (9%)  · gpt-5.6-luna, glm-5.3-flash
security:    $0.0012 · 45,424 in  / 1,606 out · 3,574 cached (8%)  · gpt-5.6-luna
tests:       $0.0001 · 12,959 in  / 156 out   · 1,856 cached (14%) · glm-5.3-flash
description: $0.0001 · 13,433 in  / 196 out   · 1,536 cached (11%) · glm-5.3-flash

}
/// Both flags default to `false` (no caching / no protection) so a host is
/// safe-by-default and opts must be explicit.
#[derive(Clone, Debug, Default, PartialEq, Eq, Serialize, Deserialize)]

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

priority medium critique confident

Apply defaults when deserializing incomplete policies

Although CachePolicy documents both flags as defaulting to false and derives Default, serde does not use that implementation for missing fields. Deserializing a valid partial policy such as {} or {"ttl_ms": 1000} therefore fails with a missing-field error. Add a serde default for the struct so omitted fields receive the documented defaults.

[RULE] serde-defaults ·

Keep the trimmed cache module: upstream's side of the conflict only
touched code this branch removes, and canonical_value is already pub.

Co-authored-by: Medulla <medulla@tinyhumans.ai>
@senamakel
senamakel merged commit fdffc2a into main Oct 7, 2026
7 checks passed

@tinysweeper tinysweeper Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

tinysweeper found nothing blocking. Approving.

             $0.0055 · 129,118 in / 9,466 out · 12,158 cached (9%) · gpt-5.6-luna, glm-5.3-flash
critique:    $0.0026 · 51,248 in  / 3,291 out · 6,605 cached (13%) · gpt-5.6-luna, glm-5.3-flash
security:    $0.0025 · 37,752 in  / 3,957 out · 5,361 cached (14%) · gpt-5.6-luna
tests:       $0.0001 · 13,313 in  / 795 out   · 64 cached (0%)     · glm-5.3-flash
description: $0.0001 · 13,911 in  / 668 out   · 64 cached (0%)     · glm-5.3-flash

pub struct CachePolicy {
/// When `true`, the host may look up (and write) local response cache
/// entries before calling the provider.
pub response_cache_enabled: bool,

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

priority medium security confident

Apply defaults when deserializing incomplete policies

The policy documents both flags as defaulting to false, but neither field has #[serde(default)]. Deserializing a persisted or externally supplied policy that omits either flag therefore fails instead of applying the documented safe default. Add #[serde(default)] to both boolean fields.


Additional critique observation

priority medium confident

Apply defaults when deserializing incomplete policies

[RULE] serde-defaults

The boolean fields do not have #[serde(default)], so deserializing a policy that omits either flag fails instead of using the documented safe defaults. For example, serde_json::from_value::<CachePolicy>(serde_json::json!({"ttl_ms": 1000})) returns a missing-field error for response_cache_enabled. Add serde defaults to both boolean fields so persisted or externally supplied policies containing only optional settings remain compatible.

Suggested change for this observation (reference only)

#[serde(default)]
    pub response_cache_enabled: bool,
    /// When `true`, middleware must preserve the order and content of cacheable
    /// prefix segments, and providers that support it mark the stable prefix
    /// for their own prompt cache.
    #[serde(default)]
    pub protect_prompt_prefix: bool,

[RULE] serde-default ·

}
/// Sets the entry TTL.
pub fn with_ttl(mut self, ttl: std::time::Duration) -> Self {
self.ttl_ms = Some(ttl.as_millis() as u64);

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

priority medium security confident

Handle durations larger than the TTL field can represent

A Duration can contain more milliseconds than u64 can represent, while ttl_ms cannot. The cast silently truncates the u128 millisecond count, so a very large requested TTL is stored as a different, much smaller TTL. Reject unrepresentable durations or otherwise define and test an explicit saturation policy instead of silently changing the value.


Additional critique observation

priority medium confident

Handle durations larger than the TTL field can represent

[RULE] lossy-integer-conversion

Duration::as_millis() returns u128, but this cast stores only u64. A duration larger than u64::MAX milliseconds is silently saturated/reduced to the maximum representable TTL, so with_ttl no longer preserves the requested expiry and entries can expire far earlier than requested. Either widen ttl_ms to a type that can represent all accepted Duration values or reject/handle unrepresentable durations explicitly.


Additional tests observation

priority medium uncertain

Handle durations larger than the TTL field can represent

[RULE] lossy-duration-cast

with_ttl stores ttl.as_millis() as u64; a Duration whose millisecond count exceeds u64::MAX silently truncates (in practice as_millis itself saturates above ~584 million years, but the as u64 cast is lossy for the full u128 range), producing a wrong TTL rather than an error or saturation. Unchanged since first raised.

Suggested change for this observation (reference only)

self.ttl_ms = Some(u64::try_from(ttl.as_millis()).unwrap_or(u64::MAX));

[RULE] integer-overflow-conversion ·

///
/// Both flags default to `false` (no caching / no protection) so a host is
/// safe-by-default and opts must be explicit.
#[derive(Clone, Debug, Default, PartialEq, Eq, Serialize, Deserialize)]

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

priority medium tests likely

Apply defaults when deserializing incomplete policies

Deserializing a partial CachePolicy — e.g. {"response_cache_enabled": true} from a config file missing the newer keys — fails outright, because the bool fields have no #[serde(default)]. Only ttl_ms and namespace get defaults. Add field-level defaults so an older or partial policy still parses as safe-by-default. Unchanged since first raised; the removal of surrounding types did not touch this.

[RULE] incomplete-deserialize ·

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant