Conversation
Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
…ion/cortexdb/flags/bitemporal-e Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
…ion/cortexdb/cortex.no-learning Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
…val/inspect.rs Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
Adds an LLM client and a scoring module to the memory_eval example so the harness can call a model and grade its answers. The scoring module computes the metrics the example reports, and the LLM module handles the request and response plumbing they depend on. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
…val/kpi.rs Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
…val/kpi.rs Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
…val/compare.rs Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
…/tinymemory-integrations/exampl Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
…val/kpi.rs,crates/tinymemory-in Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
The sweep now probes upward from the current port for one nothing is listening on, so runs no longer collide with a leftover process holding the previous port. The run helper also drops its extra arguments after the first three so the port shift does not leak into the profile invocation. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
…val/compare.rs,crates/tinymemor Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
…tegration/cortexdb/flags/no-ver Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
…val/compare.rs,crates/tinymemor Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
…val/inspect.rs,scripts/memory-f Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
…val/compare.rs,crates/tinymemor Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Tiny Sweeper review
|
|
Navigate logical layers of code changes, visualize relationships, and explore their blast radius. 🧰 Additional context used📚 Code guidelines (1)No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configuration
📒 Files selected for processing (34)
Included review availability: This review used your included allowance. Your plan provides up to 1 included review per hour; 0 remain after this review. 📝 WalkthroughWalkthroughThe memory evaluation harness now records additional run and usage data, computes and compares KPIs, and supports sweeps across configurable CortexDB flag profiles. ChangesMemory Evaluation
Priority: ⬇️ Low Estimated code review effort: 4 (Complex) | ~45 minutes Change: Feature Sequence Diagram(s)sequenceDiagram
participant Sweep as memory-flag-sweep.sh
participant Eval as memory-eval.sh
participant CortexDB
participant Compare as compare command
Sweep->>Eval: Run each profile and repetition
Eval->>CortexDB: Capture scenario results and usage
Eval-->>Sweep: Write JSON report
Sweep->>Compare: Compare available reports
Compare-->>Sweep: Write summary.md
Merge Risk: ⚪ Minimal · up to This change adds CortexDB flag profiles, richer eval reports, and a profile comparison and sweep workflow. It affects only local evaluation tooling. No concrete merge-blocking risk was found. The one known edge case is a KPI comparison mix-up that needs an exact role-name match. Security Architecture ReviewSecurity architecture risk: 🔵 Low · up to The changes are concentrated in a local evaluation workflow. Overlapping sweeps can share server ownership and tear down another evaluation's services or data. Normal runs have useful isolation and cleanup controls, and no production-wide exposure is established. Retained concerns
Security review detailsSecurity Blast Radius
Trust Boundaries and Controls
Resilience and Maintainability Implications
Hardening Proposals
🚥 Pre-merge checks | ✅ 5✅ Passed checks (5 passed)
A rabbit checks each flag at dawn, Comment |
Summary
This adds a harness that measures whether CortexDB's server flags change what memory is worth to an agent. It tracks accuracy, learning, surprise, conflicts, cost and latency.
integration/cortexdb/flags/baseline.env, with a profile fromCORTEX_FLAGS_FILEon top. There are 17 profiles. Most change one flag: graph, auto-route, HyDE/multihop, salience, surprise gate, bitemporal mode, polarity recheck, incremental layers, verifier, enrichment batching, the background learning loop, and Cohere rerank. Two are composites: max-recall and cost-optimized. Each profile's header names its target KPI and links its docs page. Every flag was checked against the v0.10.4 binary.memory_eval. A newkpi.rsturns a run into about 35 named KPIs. Cost is broken down by CortexDB role, and the--llmanswerer's tokens and cost now count too. Per-scenario model usage is reported, and the JSON records the profile, its flags and the server version.memory_eval compare(newcompare.rs) groups runs by profile, averages repeats, and marks each KPI's delta from baseline as ▲, ▼ or ~ against a noise band. That band is the repeats' spread, and never less than one probe's worth.scripts/memory-flag-sweep.shruns the profiles in parallel on free ports, with repeats, and writessummary.md. ForMODELS=openrouterit refuses to start when the key's remaining limit can't cover the planned runs.Findings so far (details in
docs/evals/cortex-flags.md):shadowbitemporal mode raises no conflict for a planted disagreement.CORTEX_VERIFIER_URLdoes not turn the verifier off in v0.10.4.no-verifierclears the keys instead; a boot-log diff confirms the verifier is the only lane that changes.Not done: the full real-model sweep. The OpenRouter key hit its $80/day limit mid-sweep, and every later run failed on 403s. The rerun command and its ~$16 cost are in the doc.
Related issue
None.
API or behavior changes
None to any crate's public API. All changes are to the eval example, scripts and the docker harness.
The harness behaves the same by default: the baseline profile reproduces the previous compose env, and
cortexdb-live.shpasses. Thememory_evalexample now setstest = true, so its unit tests run undercargo test.Validation
cargo fmt --all -- --check: cleancargo clippy --all-targets --all-features -- -D warnings: cleancargo build --all-targets --all-features: okcargo test --all-features: all pass, including 17 new eval unit tests./scripts/cortexdb-live.sh: all three live suites pass with the new compose file./scripts/memory-flag-sweep.sh(mock): 16 profiles ran and compared.rerank-coherewas skipped becauseCOHERE_API_KEYis unset.bitemporal-offexposed a 503 fromv1/conflicts, which is now handled.MODELS=openrouter: one full baseline run with--llmsucceeded. The paid sweep was cut off by the key's daily limit (see above).Tests
kpi_tests.rscovers synthesis gain, planted vs spurious conflicts, cost per correct answer including answerer spend, KPIs left unmeasured on the reference engine, latency percentiles and formatting.compare_tests.rscovers verdict direction, the noise band, the one-probe floor, deltas, repeat averaging, and errors for empty or unreadable reports.Documentation
docs/evals/cortex-flags.md(method, profiles, KPI definitions, results so far, caveats).docs/evals/README.md(now twelve scenarios, plus the Captured metric and KPIs),integration/cortexdb/README.md(flag profiles) and.env.example(the eval's variables).Checklist
#[allow(...)],#[ignore], or relaxed lints.envcontents in the diff or the descriptionSummary by CodeRabbit