Skip to content

Add a CortexDB flag sweep to the memory eval - #197

Open
senamakel wants to merge 24 commits into
tinyhumansai:mainfrom
senamakel:cortex-flag-eval
Open

senamakel wants to merge 24 commits into
tinyhumansai:mainfrom
senamakel:cortex-flag-eval

Conversation

@senamakel

@senamakel senamakel commented Oct 4, 2026 •

Copy link
Copy Markdown
Member

Summary

This adds a harness that measures whether CortexDB's server flags change what memory is worth to an agent. It tracks accuracy, learning, surprise, conflicts, cost and latency.

  • Switchable server config. The compose file now loads flags from integration/cortexdb/flags/baseline.env, with a profile from CORTEX_FLAGS_FILE on top. There are 17 profiles. Most change one flag: graph, auto-route, HyDE/multihop, salience, surprise gate, bitemporal mode, polarity recheck, incremental layers, verifier, enrichment batching, the background learning loop, and Cohere rerank. Two are composites: max-recall and cost-optimized. Each profile's header names its target KPI and links its docs page. Every flag was checked against the v0.10.4 binary.
  • KPIs in memory_eval. A new kpi.rs turns a run into about 35 named KPIs. Cost is broken down by CortexDB role, and the --llm answerer's tokens and cost now count too. Per-scenario model usage is reported, and the JSON records the profile, its flags and the server version.
  • Comparing runs. memory_eval compare (new compare.rs) groups runs by profile, averages repeats, and marks each KPI's delta from baseline as ▲, ▼ or ~ against a noise band. That band is the repeats' spread, and never less than one probe's worth.
  • Sweep script. scripts/memory-flag-sweep.sh runs the profiles in parallel on free ports, with repeats, and writes summary.md. For MODELS=openrouter it refuses to start when the key's remaining limit can't cover the planned runs.

Findings so far (details in docs/evals/cortex-flags.md):

  • The real-model baseline costs $0.43 of CortexDB models per run, and 67% of that is enrichment.
  • Default shadow bitemporal mode raises no conflict for a planted disagreement.
  • Every lesson is captured as a fact or belief, but only 71% are answered from the pack.
  • Single runs swing by 2–4 probes even on mock models, so repeats are required.
  • Docs discrepancy: an empty CORTEX_VERIFIER_URL does not turn the verifier off in v0.10.4. no-verifier clears the keys instead; a boot-log diff confirms the verifier is the only lane that changes.

Not done: the full real-model sweep. The OpenRouter key hit its $80/day limit mid-sweep, and every later run failed on 403s. The rerun command and its ~$16 cost are in the doc.

Related issue

None.

API or behavior changes

None to any crate's public API. All changes are to the eval example, scripts and the docker harness.

The harness behaves the same by default: the baseline profile reproduces the previous compose env, and cortexdb-live.sh passes. The memory_eval example now sets test = true, so its unit tests run under cargo test.

Validation

  • cargo fmt --all -- --check: clean
  • cargo clippy --all-targets --all-features -- -D warnings: clean
  • cargo build --all-targets --all-features: ok
  • cargo test --all-features: all pass, including 17 new eval unit tests
  • ./scripts/cortexdb-live.sh: all three live suites pass with the new compose file
  • ./scripts/memory-flag-sweep.sh (mock): 16 profiles ran and compared. rerank-cohere was skipped because COHERE_API_KEY is unset. bitemporal-off exposed a 503 from v1/conflicts, which is now handled.
  • MODELS=openrouter: one full baseline run with --llm succeeded. The paid sweep was cut off by the key's daily limit (see above).

Tests

  • kpi_tests.rs covers synthesis gain, planted vs spurious conflicts, cost per correct answer including answerer spend, KPIs left unmeasured on the reference engine, latency percentiles and formatting.
  • compare_tests.rs covers verdict direction, the noise band, the one-probe floor, deltas, repeat averaging, and errors for empty or unreadable reports.
  • Untested: the sweep script's shell logic and the live inspector routes, which were exercised by the runs above.

Documentation

  • New: docs/evals/cortex-flags.md (method, profiles, KPI definitions, results so far, caveats).
  • Updated: docs/evals/README.md (now twelve scenarios, plus the Captured metric and KPIs), integration/cortexdb/README.md (flag profiles) and .env.example (the eval's variables).

Checklist

  • The change is focused on one logical change
  • No new #[allow(...)], #[ignore], or relaxed lints
  • No secrets, tokens, or .env contents in the diff or the description

Summary by CodeRabbit

  • New Features
    • Added tools to run memory evaluations across configurable CortexDB flag profiles and compare repeated runs against a baseline.
    • Evaluation reports now summarize accuracy, learning, conflicts, cost, and latency, with profile settings and model usage included.
  • Documentation
    • Added guidance for configuring flag profiles, running evaluations, and interpreting comparison results, including known limitations and caveats.

senamakel and others added 24 commits October 4, 2026 20:33
Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
…ion/cortexdb/flags/bitemporal-e

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
…ion/cortexdb/cortex.no-learning

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
…val/inspect.rs

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Adds an LLM client and a scoring module to the memory_eval example so the
harness can call a model and grade its answers. The scoring module computes
the metrics the example reports, and the LLM module handles the request and
response plumbing they depend on.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
…val/kpi.rs

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
…val/kpi.rs

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
…val/compare.rs

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
…/tinymemory-integrations/exampl

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
…val/kpi.rs,crates/tinymemory-in

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
The sweep now probes upward from the current port for one nothing is listening on, so runs no longer collide with a leftover process holding the previous port. The run helper also drops its extra arguments after the first three so the port shift does not leak into the profile invocation.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
…val/compare.rs,crates/tinymemor

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
…tegration/cortexdb/flags/no-ver

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
…val/compare.rs,crates/tinymemor

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
…val/inspect.rs,scripts/memory-f

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
…val/compare.rs,crates/tinymemor

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Co-authored-by: Medulla <medulla@tinyhumans.ai>
@tinysweeper

tinysweeper Bot commented Oct 4, 2026 •

Copy link
Copy Markdown

Tiny Sweeper review

⚠️ Review failed for 85482e243578. the review of #197 did not finish within 900s

@coderabbitai

coderabbitai Bot commented Oct 4, 2026 •

Copy link
Copy Markdown

Review in Change Stack →

Navigate logical layers of code changes, visualize relationships, and explore their blast radius.

🧰 Additional context used
📚 Code guidelines (1)
AGENTS.md — auto-discovered

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration
  • Configuration used: Organization UI
  • Review profile: CHILL
  • Plan: Advanced
  • Run ID: 41bdc97a-dc13-4e66-8501-bf0565914db8
📥 Commits

Reviewing files that changed from the base of the PR and between 4b18323 and 85482e2.

📒 Files selected for processing (34)
  • .env.example
  • crates/tinymemory-integrations/Cargo.toml
  • crates/tinymemory-integrations/examples/memory_eval/compare.rs
  • crates/tinymemory-integrations/examples/memory_eval/compare_tests.rs
  • crates/tinymemory-integrations/examples/memory_eval/inspect.rs
  • crates/tinymemory-integrations/examples/memory_eval/kpi.rs
  • crates/tinymemory-integrations/examples/memory_eval/kpi_tests.rs
  • crates/tinymemory-integrations/examples/memory_eval/llm.rs
  • crates/tinymemory-integrations/examples/memory_eval/main.rs
  • crates/tinymemory-integrations/examples/memory_eval/score.rs
  • docs/evals/README.md
  • docs/evals/cortex-flags.md
  • integration/cortexdb/README.md
  • integration/cortexdb/cortex.no-learning.toml
  • integration/cortexdb/docker-compose.yml
  • integration/cortexdb/flags/baseline.env
  • integration/cortexdb/flags/bitemporal-enforce.env
  • integration/cortexdb/flags/bitemporal-off.env
  • integration/cortexdb/flags/cost-optimized.env
  • integration/cortexdb/flags/enrich-batched.env
  • integration/cortexdb/flags/layers-incremental.env
  • integration/cortexdb/flags/max-recall.env
  • integration/cortexdb/flags/no-auto-route.env
  • integration/cortexdb/flags/no-graph.env
  • integration/cortexdb/flags/no-hyde-multihop.env
  • integration/cortexdb/flags/no-learning-loop.env
  • integration/cortexdb/flags/no-polarity-recheck.env
  • integration/cortexdb/flags/no-verifier.env
  • integration/cortexdb/flags/rerank-cohere.env
  • integration/cortexdb/flags/salience-high.env
  • integration/cortexdb/flags/surprise-loose.env
  • integration/cortexdb/flags/surprise-strict.env
  • scripts/memory-eval.sh
  • scripts/memory-flag-sweep.sh

Included review availability: This review used your included allowance. Your plan provides up to 1 included review per hour; 0 remain after this review.


📝 Walkthrough

Walkthrough

The memory evaluation harness now records additional run and usage data, computes and compares KPIs, and supports sweeps across configurable CortexDB flag profiles.

Changes

Memory Evaluation

Layer / File(s) Summary
CortexDB flag profiles
.env.example, integration/cortexdb/docker-compose.yml, integration/cortexdb/flags/*, integration/cortexdb/cortex.no-learning.toml, integration/cortexdb/README.md
Compose loads baseline and selected profile settings. The change adds flag profiles, a no-learning TOML preset, configurable credentials and configuration paths, and setup documentation.
Scenario capture and report data
crates/tinymemory-integrations/examples/memory_eval/inspect.rs, llm.rs, score.rs, main.rs, scripts/memory-eval.sh
The evaluation records profile metadata, server version, and per-scenario usage. Captured beliefs include stance counts and available confidence values. LLM answers include token and optional cost data, and report output supports a configurable directory.
KPI computation and validation
crates/tinymemory-integrations/examples/memory_eval/kpi.rs, kpi_tests.rs, crates/tinymemory-integrations/Cargo.toml, docs/evals/README.md
The new KPI module computes accuracy, learning, surprise, conflict, cost, and latency metrics. Tests cover calculations and formatting. The example target is configured so its unit tests run under cargo test.
Profile comparison and report output
crates/tinymemory-integrations/examples/memory_eval/compare.rs, compare_tests.rs, main.rs
The comparison command groups reports by profile and compares KPI means against a baseline using repeat spreads and unit-specific thresholds. Tests cover verdicts, formatting, aggregation, and invalid inputs.
Profile sweep execution and documentation
scripts/memory-flag-sweep.sh, docs/evals/README.md, docs/evals/cortex-flags.md
The sweep script runs selected or default profiles with configurable repeats and parallelism, then compares available reports. The evaluation documentation describes profile runs, KPI interpretation, and recorded findings.

Priority: ⬇️ Low

Estimated code review effort: 4 (Complex) | ~45 minutes

Change: Feature

Sequence Diagram(s)

sequenceDiagram
  participant Sweep as memory-flag-sweep.sh
  participant Eval as memory-eval.sh
  participant CortexDB
  participant Compare as compare command
  Sweep->>Eval: Run each profile and repetition
  Eval->>CortexDB: Capture scenario results and usage
  Eval-->>Sweep: Write JSON report
  Sweep->>Compare: Compare available reports
  Compare-->>Sweep: Write summary.md
Loading

Merge Risk: ⚪ Minimal · up to 85482

This change adds CortexDB flag profiles, richer eval reports, and a profile comparison and sweep workflow. It affects only local evaluation tooling. No concrete merge-blocking risk was found. The one known edge case is a KPI comparison mix-up that needs an exact role-name match.

Security Architecture Review

Security architecture risk: 🔵 Low · up to 85482

The changes are concentrated in a local evaluation workflow. Overlapping sweeps can share server ownership and tear down another evaluation's services or data. Normal runs have useful isolation and cleanup controls, and no production-wide exposure is established.

Retained concerns

  • Low · reliability · inferred: Overlapping sweeps using the same model mode and port range can pass the availability checks before either server starts and select the same Compose project. Startup failure or completion in either invocation can then remove the other's services and evaluation volume. Sequential ports protect runs within one sweep, but do not establish ownership across sweep processes. The new automated caller broadens exposure to this existing ownership limitation.
Security review details

Security Blast Radius

  • inferred — The demonstrated ownership failure is bounded to evaluations sharing a Docker host, model mode, and port selection. It can affect their containers, credential-bearing process lifetime, and synthetic evaluation data. Exploiting this path requires local execution with Docker access; no tenant-crossing or production privilege gain is established.

Trust Boundaries and Controls

  • observed — Profiles are executable configuration: the sweep sources the selected file while inheriting the operator's environment and credentials. They must therefore be trusted like the scripts themselves, not treated as inert untrusted flag data. Server access remains loopback-bound and uses the configured API key.

Resilience and Maintainability Implications

  • observed — The sweep has no parent-level child-cancellation or recovery handler. Its spending check rejects a reported insufficient balance, but treats an empty parsed balance as unlimited or a failed check. These mechanisms are best-effort operational safeguards, not guaranteed cancellation or spending boundaries.

Hardening Proposals

  • proposed — Give each sweep a unique ownership namespace for Compose projects and output files, coordinate port allocation across processes, and restrict teardown to resources owned by that run. Explicit child cancellation and verifiable cleanup would strengthen containment when a sweep is interrupted.
🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title clearly describes the primary change: adding a CortexDB flag sweep to the memory evaluation.
Docstring Coverage ✅ Passed Docstring coverage is 85.45% which is sufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 55 functions across 27 files. (7 skipped: 7…
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

A rabbit checks each flag at dawn,
Then gathers scores as runs roll on.
With carrots near, it weighs the spread,
And writes a chart before its bed.
The burrow hums; the trials are done!

Comment @coderabbitai help to get the list of available commands.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant