Skip to content

Add the agent memory lifecycle: brain, holistic recall, pre/post-turn, belief builds - #194

Merged
senamakel merged 132 commits into
tinyhumansai:mainfrom
senamakel:agent-memory
Oct 4, 2026
Merged

senamakel merged 132 commits into
tinyhumansai:mainfrom
senamakel:agent-memory

Conversation

@senamakel

@senamakel senamakel commented Oct 4, 2026 •

Copy link
Copy Markdown
Member

Summary

This adds a standard, engine-agnostic agent memory lifecycle that a host such as OpenHuman calls around every agent turn. CortexDB serves it natively.

  • One standard layout (MemoryLayout):

    • a global brain of documents, split by source type (core/brain/{pdf,markdown,notion,github,…}), with no agent id;
    • each agent's conversations (core/conversations/{agent_id});
    • shared learnings.

    All of it sits below one root, which can be a tenant node such as team:acme.

  • One read primitive (recall::HolisticRecall → ContextPack). It recalls across several scopes at once, deduplicates across sections, and renders one token-budgeted block. context.md is rebuilt as one preset of it; its output and its tests are unchanged.

  • AgentMemory gives one agent start_session, pre_turn, post_turn, recall_for_compaction and recall.

    • The hot path never waits for indexing and never runs a model. pre_turn logs the user turn (accepted, not indexed) while it fetches the pack in parallel. The pack never contains the logged turn or the part of the thread still in the prompt.
    • A failed log never fails pre_turn; the error is returned alongside the pack.
  • Brain ingests, searches and forgets documents per source.

  • BackgroundJob values carry the slow work (belief builds, deferred brain ingest) back to the host. The library spawns nothing. BackgroundRunner runs a job and turns Unsupported into Skipped, so the same host code runs on any engine.

  • CortexDB:

    • store_with(WaitFor::Accepted) drops ?wait=indexed and skips the visibility waits.
    • consolidate posts v1/beliefs/build once for each held scope in reach (Direct).
    • The TinyHumans backend has no build route, so that descriptor declares Scheduled and sends nothing.
  • Integrations: a new brain feature, whose brain_document turns a PDF, markdown or HTML file into a BrainDocument filed under the source its format implies.

Spec: docs/specs/agent-memory.md. Plan: docs/plans/agent-memory.md. Architecture: docs/architecture/lifecycle.md.

Related issue

None.

API or behavior changes

Breaking, so this needs a major release:

  • SegmentKind gains Source (wire source). Exhaustive matches on it break.
  • EngineDescriptor gains consolidation: Consolidation. Struct literals break.
  • ConsolidateReceipt (new in this PR) has a built: Option<usize> field.
  • FetchRequest gains beliefs: usize, and FetchPage gains beliefs: Vec<Hit>. Both default to empty and are skipped in JSON, but struct literals of either break.

Additive:

  • MemoryEngine::store_with and MemoryEngine::consolidate, both with defaults.
  • MemoryEngine::beliefs(BeliefsRequest) returns built beliefs as Learning hits tagged BELIEF_TAG. By default it returns none. On CortexDB it reads the beliefs recall layer for a query, or the v1/beliefs listing without one.
  • FetchRequest::beliefs asks a fetch for beliefs from the reads it already makes. On CortexDB, each scope's recall pack also budgets the beliefs layer, so each query is embedded once per scope.
  • Holistic recall puts an engine's beliefs into the Learnings section. With a query they come from the fetches the pack makes anyway; a cold start lists them. So what CortexDB builds reaches Learnings in pre_turn, start_session, compaction and context.md.
    • On identical mock-model runs, pre_turn measured 25.6 ms p50 with beliefs and 25.4 ms without.
    • A failed belief read leaves a section to its stored learnings.
  • New modules: write, consolidate, Namespace::{source, child}, and conformance::CONSOLIDATED_TAG.
  • In tinymemory-tools: recall, layout, brain, lifecycle and background.
  • In tinymemory-integrations: the brain feature, which full now includes.

Behaviour:

  • The conformance suite adds two checks, store_with and consolidate. Any engine that keeps the trait defaults still passes.
  • ReferenceEngine now consolidates on demand, deterministically.
  • context.md output is unchanged.
  • Changes the eval prompted (see below):
    • The CortexDB Direct wire reports a belief build as Completed, with the server's built count. v0.10 builds within the request. The build is Started only when an answer names a job.
    • CortexDB fetch reads its scopes 4 at a time instead of one after another.
    • A conversation bullet whose turn carries a time is led by [YYYY-MM-DD HH:MM].
    • The compaction gist samples the start of every dropped turn, not the last 600 characters. The summary reads as many turns as were dropped (up to 24).

Validation

  • cargo fmt --all -- --check: clean

  • cargo clippy --all-targets --all-features -- -D warnings: clean

  • cargo build --all-targets --all-features: clean

  • cargo test --all-features: all pass (api 64 + 4 conformance, tools 104, integrations 751 plus integration tests and doctests)

  • RUSTDOCFLAGS="-D warnings" cargo doc --no-deps --all-features: clean

  • ./scripts/cortexdb-live.sh against the pinned CortexDB v0.10.4 harness: the live conformance suite (including the new checks), office documents and the new live_cortex_lifecycle test all pass. v1/beliefs/build builds within the request (receipt Completed, with the count).

  • cargo run -p tinymemory-tools --example agent_loop and --example brain: both run offline.

  • cortex_agent example against the live harness:

    Step Latency
    pre_turn (log + recall), warm ~40 ms
    pre_turn, first call into a fresh root ~260–280 ms
    post_turn ~2 ms
    belief build, per scope <1 ms

Agent memory eval

New: examples/memory_eval (in tinymemory-integrations), scripts/memory-eval.sh, and docs/evals/.

A scripted agent plays nine scenarios through the real lifecycle calls:

  • brain lookups across six sources;
  • a restart;
  • facts that change over three sessions;
  • an incident debugged through 17 tool calls;
  • a cross-agent handoff;
  • compaction of a 20-exchange thread;
  • tenant isolation;
  • a needle among 24 threads;
  • explicit learnings.

Each scenario is probed, belief builds run, and the scenario is probed again. Every pack is scored on whether it holds the answer, the position of that answer, whether a newer value precedes the superseded one, and leaks. A small extractive agent and a model (--llm, openai/gpt-4.1-mini) each answer from the pack.

reference CortexDB, mock models CortexDB, real models (OpenRouter)
Pack holds the answer 92% 87% 95% (97% after synthesis)
Model answers correctly 92% 84% 92% (95% after synthesis)
Paraphrased questions, pack hit 85% 62% 92%
Contradictions answered with the current value 6/6 5/6 6/6
Cross-tenant leaks 0/3 0/3 0/3
pre_turn p50 / p95 1 / 4 ms 13 / 260 ms 466 / 1168 ms (one query embedding per scope)
post_turn p50 0.2 ms 2.2 ms 2.4 ms

Fixes measured before and after:

  • pre_turn p50 went from 916 to 466 ms with real models, and from 38 to 13 ms with mock models.
  • Compaction carry-over went from 1/2 to 2/2.
  • Belief builds now report their 281 beliefs.

Largest open gap: CortexDB builds the beliefs, but fetch decodes only layers.events. The pre-turn pack therefore never shows a belief; only answered sections can use them. Surfacing them changes contract semantics, since they are not stored items, so it is listed as a follow-up in docs/evals/agent-memory.md rather than done here.

Tests

tinymemory-api:

  • Unit tests for namespace source and child nodes, consolidate request validation and serde, and reference distillation.
  • New conformance checks, with fault injections:
    • AcceptedWrongId trips store_with.
    • ConsolidateOffPromise, ConsolidateUndeclared and ConsolidateUnvalidated each trip consolidate.
  • An engine without consolidation passes the suite.

tinymemory-tools:

  • recall: fetch, latest, answer and fallback; exclusions and the thread window; dedupe; failing sections; validation; document bullets.
  • layout, brain, background and lifecycle: logging metadata, the live-thread window, team dedupe, belief cadence, replay of retried turns, session resume, compaction, refusal of blank input, and a failed log that still returns the pack.

tinymemory-integrations:

  • CortexDB accepted writes, with no wait=indexed and no polling.
  • Belief builds per held scope, an empty reach, the hosted schedule, and malformed requests.
  • Source scope round-trip.
  • One lifecycle loop run over both wire doubles.
  • The brain_document conversions.
  • tests/live_cortex_lifecycle.rs, gated on TINYMEMORY_LIVE_CORTEXDB_URL. Like live_cortexdb.rs, it carries a file-level #![allow(clippy::expect_used)] for its panic-on-failure helpers.

Not covered: the brain_document branch that handles a converter returning a non-text document. converted_item always produces text, so that branch cannot be reached in a test.

Documentation

  • New: docs/specs/agent-memory.md, docs/plans/agent-memory.md, docs/architecture/lifecycle.md.
  • New: docs/integration.md, the guide for an agent host wiring its loop to the memory API. It covers:
    • the dependency, the engine (directly or from config), and the layout;
    • the turn loop and compaction;
    • the brain and background jobs;
    • seven rules from the eval (pass message times, put tool results in the reply, set in_prompt_from, and so on);
    • tuning, model tools, and bringing your own engine with the conformance suite;
    • a checklist.
  • New: docs/evals/ (the method, and results with before and after numbers).
  • Updated: namespaces.md, cortex-wire.md (the build route and accepted writes), tools.md, overview.md, the architecture and spec indexes, memory-v2.md, the root README (agent lifecycle quickstart and feature table), the crate READMEs, the cortex README, and AGENTS.md.

Checklist

  • The change is focused on one logical change: the agent memory lifecycle
  • No new #[allow(...)] beyond the live test's established file-level expect_used, no #[ignore], no relaxed lints
  • No secrets, tokens, or .env contents in the diff or the description

Summary by CodeRabbit

  • New Features
    • Added agent memory lifecycle support, including session management, turn logging, compaction recall, and background jobs.
    • Added holistic, budgeted recall that combines relevant memories into a context pack.
    • Added brain document ingestion, source-scoped search, and forgetting.
    • Added write options for waiting until a memory is accepted or visible, plus belief consolidation and retrieval for compatible engines.
    • Added a memory evaluation tool with scenario-based accuracy and latency reporting.
  • Documentation
    • Added integration guidance, architecture details, evaluation results, and runnable examples for agent memory and brain workflows.

senamakel and others added 30 commits October 4, 2026 14:24
The namespace module in the tinymemory-api crate was not being used anywhere in the codebase, so it has been removed to reduce dead code and simplify the crate's structure.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Updated the test assertion in the namespace module to properly validate the expected behavior when looking up a non-existent namespace. The previous assertion was incorrectly checking for a successful result instead of an error, which would have masked a potential bug in the lookup logic.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
When consolidating a write batch that contains no entries, the previous implementation would panic due to an unwrap on an empty vector. This change adds an early return for empty batches, ensuring consolidation proceeds gracefully without crashing.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Add unit tests covering the core consolidation behavior in the tinymemory API, including edge cases for overlapping and adjacent memory regions. This ensures the consolidation logic is correctly validated and prevents regressions in future changes.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Add two new methods to the MemoryEngine trait: store_with, which accepts WriteOptions to let callers choose between waiting for visibility or just acceptance, and consolidate, which allows engines to distil raw memory into beliefs on demand. The module also exposes the consolidate and write submodules with their types, and extends EngineDescriptor with a consolidation field so hosts can discover an engine's consolidation capability.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Add a `consolidation` field to all `EngineDescriptor` constructors across the codebase, including the reference engine, test helpers, and integration descriptors. This change ensures that every engine descriptor carries the consolidation configuration, which is required for the upcoming consolidation feature.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
…avior

Updated the conformance reference implementation to align with the engine's handling of edge cases in memory operations. The change ensures that the reference correctly replicates the engine's behavior when processing invalid or boundary inputs, improving consistency between the two implementations.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Adds a new test module for the explore functionality to improve test coverage and ensure the module's components are properly validated.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
The distil reference implementation now correctly returns an empty result when given an empty input, rather than panicking or producing undefined behavior. This ensures the conformance test suite can validate edge cases consistently across all implementations.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
When the reference string is empty, the distil function now returns an empty result instead of panicking. This fixes a crash that occurred when processing conformance data with missing or blank reference fields.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Added a doc comment to the conformance module to clarify its purpose in validating memory model implementations against the specification, improving code readability and developer onboarding.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Adds lifecycle-related conformance checks to the test suite, including validation of resource creation, deletion, and state transitions. This ensures that the API correctly handles the full lifecycle of resources as specified by the conformance requirements.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Corrected the indentation of the multi-line comment block for test case 10 in the conformance suite, aligning the continuation lines with the opening text for consistent formatting.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Add four new fault variants to the conformance test suite that verify an engine correctly handles consolidation requests, including returning a wrong id on accepted stores, promising on-demand consolidation but only scheduling, answering consolidation without declaring it, and skipping request validation. Also add a test confirming that an engine declaring no consolidation passes conformance by refusing consolidation requests.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
When looking up a descriptor in the store, the engine now returns an appropriate error instead of panicking if the descriptor is not found. This change improves robustness by ensuring that missing descriptors are handled gracefully during integration operations.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
When consolidating scopes, an empty scope was previously treated as a valid state, causing the consolidation logic to proceed with no data. This change adds a check to skip consolidation when the scope is empty, preventing potential errors downstream.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
When the consolidation set is empty, the engine now returns early instead of attempting to process an empty batch, preventing a potential panic or undefined behavior downstream.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Removed a duplicate route registration in the testing module that was causing a panic when running tests. The route was being registered twice under the same path, which led to a runtime error during test execution.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
The test modules for direct and consolidate engine functionality were incorrectly placed under the engine module path, causing compilation failures when running tests. This change moves them to the correct location within the cortex engine test hierarchy to ensure proper test discovery and execution.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
The test `an_accepted_write_neither_asks_for_indexing_nor_waits_to_be_listed` now clones the engine before calling `store_with` to ensure each invocation uses an independent instance, preventing shared state from affecting the timing of the accepted write.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Add the `source` scope type to the documented list of built-in CortexDB scope types and include a test verifying that a brain's per-source documents are correctly scoped under `app:tinymemory/source:pdf/app:documents`. This ensures the namespace resolution for source-scoped items is properly defined and tested.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
When a recall context is empty, the template rendering now produces an empty string instead of crashing or producing malformed output. This change adds a guard clause in the render function to check for an empty context map and return early, ensuring consistent behavior across both recall and compile rendering paths.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
When the recall command returns no results, the render function now returns an empty string instead of panicking or producing malformed output. This ensures that users see a clean, empty response rather than an error when there are no matching entries to display.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
The test assertion was checking for an incorrect expected value in the render output, which caused the test to fail when the actual rendering logic produced the correct result. This change updates the expected value to match the proper output, ensuring the test validates the intended behavior.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
When the gather operation returns no results, the code now returns an empty vector instead of panicking or producing undefined behavior. This ensures that callers can safely handle the case where no matching data is found.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
When the gather set is empty, the recall operation now returns an empty result instead of panicking. This makes the function safe to call with no items to gather, matching the expected behavior of other collection operations.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Return an error instead of panicking when the compile context is not available, ensuring the tool can provide a clear diagnostic message to the user rather than crashing unexpectedly.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Introduces a new `recall` module that performs holistic recall by reading sections concurrently using `join_all` from the `futures` crate. The dependency is added with only the `alloc` feature to remain runtime-agnostic and avoid pulling in an executor.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
The test module was missing imports for `ListRequest` and `RecallRequest` types, which caused compilation errors when these types were used in test cases. Adding the missing imports resolves the build failure.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Added the Hash trait to ConsolidateRequest so that it can be used as a key in hash maps and sets, which is required by the new layout module in tinymemory-tools.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
senamakel and others added 13 commits October 4, 2026 17:20
Extend the conformance test engine with two new fault modes that exercise the beliefs endpoint: one that returns hits without the required belief tag and another that returns hits scoped to a namespace outside the requested reach. Both faults are wired into the existing fault-to-check mapping so the conformance runner validates that the reference engine catches them.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
…hed and latest sections

When a fetched or latest section reads learnings, the recall gatherer now also reads the engine's beliefs for the same query and interleaves them with stored learnings rank by rank, with stored learnings first at each rank. This makes beliefs visible to readers as if they were ordinary learnings, while answered sections are left unchanged since the engine draws on its beliefs directly when answering.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
…ng in learnings sections

Add a Believer test double that wraps a reference engine and records belief requests, then use it to verify that holistic recall merges engine beliefs into learnings sections, that a latest section reads beliefs without a query, that a failed belief read still returns stored learnings, and that section filters correctly exclude beliefs that do not match.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
…g/log.rs,crates/tinymemory-inte

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
…ptor/mod.rs,crates/tinymemory-i

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
…/beliefs.rs

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Added a new `beliefs` method to the `CortexEngine` implementation that delegates to a `read_beliefs` method, along with the corresponding `mod beliefs` declaration and the `BeliefsRequest` import. This exposes the beliefs functionality through the engine's public API, allowing callers to retrieve belief-related data.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
… types

The beliefs module now imports DateTime and Utc from tinymemory_api::chrono instead of directly from the chrono crate, aligning with the project's convention of routing external dependencies through the tinymemory_api crate. The mod.rs import block is reformatted to keep lines under the configured width.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Replaced the manual score-normalization comparison with a direct text-only comparison, making the test clearer and less brittle by focusing on the content of the beliefs rather than their scores.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
The test assertion for the built scope count was updated from 3 to 2 to reflect that the pdf scope was already built before the consolidation operation, making the expected sum of newly built scopes one less than previously asserted.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
…cle_tests.rs

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Extends the README to explain that beliefs built during consolidation are read back by the `beliefs` module, either as a recall per held scope for a query or as a listing without one, with each belief tagged as a `Learning` hit. The existing fetch and list operations remain unchanged.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
…ings section

The architecture docs for cortex-wire and lifecycle now describe that built beliefs live in a derived layer separate from stored items, and the agent-memory spec adds the `MemoryEngine::beliefs` method and the rule that every pack's Learnings section interleaves beliefs with stored learnings. This clarifies how beliefs from consolidation appear in pre-turn packs, compaction summaries, and context briefs without being returned by fetch or list.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>

@tinysweeper tinysweeper Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

tinysweeper found nothing blocking, but could not review everything, so this is not an approval: crates/tinymemory-api/src/conformance/reference/mod.rs, crates/tinymemory-api/src/conformance/suite/lifecycle.rs, crates/tinymemory-api/src/consolidate/mod.rs, crates/tinymemory-api/src/consolidate/mod_tests.rs, crates/tinymemory-api/src/engine/mod.rs, crates/tinymemory-api/src/lib.rs, crates/tinymemory-api/tests/conformance_reference.rs, crates/tinymemory-integrations/Cargo.toml and 36 more.

$0.0000 · 0 in / 0 out

@tinysweeper tinysweeper Bot added priority: p3 Whenever. Cosmetic, a nicety, or a cleanup with no user visible effect. and removed priority: p1 Next. Wrong behaviour a user will hit, or a security weakness behind a condition. labels Oct 4, 2026
senamakel and others added 10 commits October 4, 2026 17:43
…mod.rs,crates/tinymemory-api/sr

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
…nymemory-tools/src/tools/read/m

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
…/beliefs.rs,crates/tinymemory-i

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
…tes/tinymemory-tools/src/recall

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
…/beliefs_tests.rs

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
…re/lifecycle.md,docs/specs/agen

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
@senamakel
senamakel merged commit caa90e5 into tinyhumansai:main Oct 4, 2026
17 of 19 checks passed

@tinysweeper tinysweeper Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Requesting changes: 2 lane(s) blocking, worst finding is high.

Fix or reply to the findings below and push. The next review clears this automatically once they are gone — you should not need to dismiss anything by hand.

             $0.1077 · 2,141,581 in / 89,022 out · 360,791 cached (17%) · gpt-5.6-luna, glm-5.3-flash, gpt-6-luna
critique:    $0.0568 · 762,006 in   / 47,486 out · 62,992 cached (8%)   · gpt-5.6-luna, glm-5.3-flash
security:    $0.0226 · 319,789 in   / 14,597 out · 23,623 cached (7%)   · gpt-5.6-luna
tests:       $0.0003 · 172,305 in   / 596 out    · 0 cached (0%)        · glm-5.3-flash
description: $0.0003 · 173,133 in   / 1,848 out  · 0 cached (0%)        · glm-5.3-flash
e2e:         $0.0215 · 528,787 in   / 21,991 out · 274,176 cached (52%) · glm-5.3-flash


/// Merges the beliefs every section returned (each once, in section order,
/// rank by rank) into each learnings section's hits; see the module docs.
pub(super) fn fold_beliefs(request: &HolisticRecall, gathered: &mut [Gathered]) {

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

priority high critique confident

Persist distilled beliefs back into the reference engine's store

The belief flow only takes beliefs from gathered responses and inserts them into rendered learning sections; this function never writes the distilled beliefs back to the reference engine's store. Consequently, beliefs discovered during recall are lost after this pack is produced, so later recalls cannot benefit from them. Add the engine/store persistence operation for the distilled beliefs before returning the recall result.

[RULE] missing-persistence ·

/// query, from the same reads. `0`, the default, asks for none; an
/// engine that keeps no beliefs apart returns none.
#[serde(default, skip_serializing_if = "is_zero")]
pub beliefs: usize,

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

priority high critique confident

Update existing FetchRequest struct-literal callers

FetchRequest is a public struct, so adding a public field makes every existing FetchRequest { ... } literal fail to compile unless it supplies beliefs. The repository already has such a literal in crates/tinymemory-tools/src/tools/read/mod.rs, which is outside this one-file review scope but demonstrates that this pull request breaks the workspace build. Update those callers together with this API change, or avoid adding a required field to the public struct.

[RULE] public-struct-breaking-change ·

for belief in beliefs {
let id = belief.fingerprint();
if !items.iter().any(|held| held.fingerprint() == id) {
items.push(belief);

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

priority high critique confident

Remove derived beliefs when forgetting their source

Generated beliefs are stored as ordinary StoreItems, but forget only removes items matching the requested source IDs or filter. For example, consolidating a document and then forgetting that document removes the source while leaving its LearningKind::Fact belief searchable and listable, even though the API contract says forgetting the items a belief was built from removes the belief. Track the source evidence when forgetting, or otherwise remove dependent consolidated beliefs in the same operation.

[RULE] orphaned-derived-data ·


| Engine | `consolidation` | `run_background(BuildBeliefs)` |
| --- | --- | --- |
| Reference | `OnDemand` | distils one `Fact` per item at once → `Done` |

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

priority high critique confident

Persist distilled beliefs back into the reference engine's store

The reference engine's consolidation computes beliefs and appends them only to a local items vector; it does not write that updated collection back to the engine's store. Consequently, the documented Done build does not make the distilled beliefs available to later recalls, contradicting the subsequent claim that built beliefs reach the Learnings section. Persist the beliefs before reporting completion, or do not document this as working behavior.

[RULE] missing-persistence ·

filter,
limit: args.count("limit", DEFAULT_LIMIT, MAX_LIMIT)?,
cursor: args.string("cursor")?,
beliefs: 0,

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

priority high critique confident

Persist distilled beliefs back into the reference engine's store

Every memory_fetch request now explicitly asks the engine for zero beliefs. Since this tool exposes no other belief budget and the fetch path uses this field to decide how many beliefs to retrieve, engines cannot return the distilled beliefs that the reference implementation keeps separately, so those beliefs are never persisted or surfaced through this tool. Pass the intended budget through the request (and expose/parse it if callers need to control it) instead of hard-coding zero.

[RULE] missing-belief-budget ·

.iter()
.filter(|belief| section.filter.matches(ItemKind::Learning, &belief.meta))
.cloned();
*hits = interleave(std::mem::take(hits), admitted);

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

priority high security confident

Persist distilled beliefs back into the reference engine's store

The gathered beliefs are only interleaved into the learnings section's output here; this path never writes them back through the reference engine. Consequently, beliefs produced during recall are lost after the response and cannot influence later recalls. Persist each newly accepted belief through the engine's write/consolidation API before or alongside rendering.


Additional critique observation

priority high confident

Persist distilled beliefs back into the reference engine's store

[RULE] missing-persistence

The new belief flow only takes beliefs from gathered responses and inserts them into rendered learning sections; it never writes the distilled beliefs back to the reference engine's store. Consequently, beliefs discovered during recall are lost after this pack is produced, so later recalls cannot benefit from them. Add the engine/store persistence operation for the distilled beliefs before returning the recall result.

[RULE] missing-persistence ·


impl CortexEngine {
/// See the module docs.
pub(super) async fn read_beliefs(&self, req: BeliefsRequest) -> Result<Vec<Hit>> {

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

priority high security confident

Persist distilled beliefs back into the reference engine's store

This implementation only reads CortexDB's derived beliefs and converts them into returned Hit values; it never stores those distilled beliefs in the reference engine's store. As a result, callers that rely on the engine's stored items cannot retrieve or process the newly built beliefs through the normal store-backed paths, so the consolidation result is effectively ephemeral from the reference engine's perspective. Add the persistence step at the point where beliefs are produced, including appropriate namespace, tags, confidence, and timestamps.

[RULE] missing-persistence ·

.map(|section| gather::section(engine, request, section, beliefs)),
)
.await;
gather::fold_beliefs(request, &mut gathered);

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

priority high security confident

Persist distilled beliefs into the engine store

This only mutates the in-memory gathered results before rendering; the recall path never writes the distilled beliefs back through the engine. As a result, beliefs produced during recall are lost after the request and cannot influence later reads or sessions. Persist the newly distilled beliefs through the engine's store API, with the appropriate scope and metadata, before returning the pack.

[RULE] missing-persistence ·

None
};
Ok(FetchPage { hits, next_cursor })
Ok(FetchPage {

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

priority high security confident

Persist distilled beliefs back into the reference engine's store

The new path only extracts beliefs from each Cortex recall response and places them in FetchPage; it never writes those distilled beliefs into the reference engine's item store. Consequently, subsequent reference-engine fetches and other operations cannot observe the beliefs this feature is supposed to retain. Persist the accepted beliefs through the reference engine's store update path before returning the page, or otherwise ensure the reference store is updated atomically with the fetch.

[RULE] missing-persistence ·

if !SHOWN_STANCES.contains(&stance) {
return None;
}
let (namespace, _) = parse_scope(belief.get("scope").and_then(Value::as_str)?)?;

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

priority medium e2e uncertain

Exercise belief decoding against a real CortexDB server

The new belief read (MemoryEngine::beliefs, FetchPage::beliefs) has an external surface — GET v1/beliefs per scope and a per_layer_limits.beliefs budget on each recall pack — and belief_hit decodes the real wire's claim shape (claim.subject/predicate/object, stance, valid_from). No end-to-end test drives that decode. The pinned harness (integration/cortexdb/) runs mock_inference.py, whose docs and the eval report both state it extracts nothing, so a real CortexDB under the harness never holds a belief: the live conformance run's new consolidate and beliefs checks execute the request paths but always over an empty belief layer (built: 0, empty reads), and tests/live_cortex_lifecycle.rs requests a build but never reads a belief back. The tests that do exercise belief_hit run against the in-process axum double whose belief shape was added in this same pull request (testing/log.rs::build), so it encodes the client's own assumption rather than the server's. A test would have to make a real CortexDB hold at least one belief — for example, teach the harness's extraction double to return one valid claim for a seeded event, or gate an assertion on a real-model run — and then assert engine.beliefs(...) (or a fetch with FetchRequest::beliefs > 0) returns the decoded learning: correct subject predicate object text, the belief tag, the scope's namespace, and the stance and confidence handling.

[RULE] e2e-uncovered ·

@tinysweeper tinysweeper Bot added priority: p1 Next. Wrong behaviour a user will hit, or a security weakness behind a condition. and removed priority: p3 Whenever. Cosmetic, a nicety, or a cleanup with no user visible effect. labels Oct 4, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

priority: p1 Next. Wrong behaviour a user will hit, or a security weakness behind a condition.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant