Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
340 changes: 145 additions & 195 deletions CLAUDE.md

Large diffs are not rendered by default.

37 changes: 36 additions & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -42,7 +42,8 @@ Or paste this into `opencode.json` (project) or

The first time OpenCode talks to the server it opens a browser so you can
click Allow. You can also run `opencode mcp auth memvara`. That grant lasts
90 days, and no API key ships in this repo.
until you revoke it, or ten years, whichever comes first. No API key ships
in this repo.

## What runs on your machine

Expand Down Expand Up @@ -77,6 +78,40 @@ before a retry succeeded — and `capture.log` is where that shows.

To have the endpoint and none of this, install with `--mcp-only`.

### What else the hooks keep and send

The hooks are copied from memvara/memvara v0.15.0, and `hooks.lock` names the
exact commit. Besides the logs, they keep two kinds of small file in
`~/.memvara/.hooks/`:

- `projects/` holds the git project of each directory the hooks ran in, for
one hour. The project is the repository's `origin` remote, written as
`host/owner/repo`, or `path:` and a hash of the repository root when there
is no remote. On the hosted server the hooks send it with every call, in a
`Memvara-Project` header, so that memories are kept per repository.
- `counts/` holds one file per session with three numbers: the memory lines
the recall hook put into prompts, the read-only memory tools the model
called, and the facts capture stored. The Claude Code plugin shows them in
a status line. This plugin has no status line, so here they are only a
record. A file untouched for 14 days is removed.

Every memory line the hooks put in front of the model starts with `⋈`, and
capture ignores lines that start with it, so a recalled memory is not stored
a second time.

You can switch each of these off in `~/.memvara/settings.json`, a JSON object
of `true` and `false` values in which a missing key means on: `project_scope`
for the project header, `status_line` for the counts, and `recall_mark` for
the mark. Setting `MEMVARA_FEATURE_<NAME>` to `0` or `1` in the environment
overrides the file.

The Claude Code plugin also has agentic capture, where capture searches your
memory before it proposes changes. It runs only when `claude -p` is the first
extractor, and here `opencode run` is. Capture here still reads each turn with one
call, and each turn's `capture.log` entry includes a line saying that agentic
capture was skipped. Setting `"agentic_capture": false` in the same file stops
that line.

## Skill

The judgment that spans tools is in `skills/memvara/SKILL.md`. Copy the
Expand Down
2 changes: 1 addition & 1 deletion hooks.lock
Original file line number Diff line number Diff line change
Expand Up @@ -6,6 +6,6 @@
# no hooks.json here and nothing generates one; tools/generate.py refuses this host by
# name for exactly that reason.
repo=memvara/memvara
sha=eb25ea028e9f70372d2da7903215636a40df320a
sha=026be5cfd815419fa4c7eda2cb0ada109ea22cab
path=plugin/hooks
host=opencode
18 changes: 15 additions & 3 deletions hooks/approve.py
Original file line number Diff line number Diff line change
Expand Up @@ -4,6 +4,9 @@
SuperMemory auto-allows search; writes still ask. Same split here. A silent
no-op on any other tool, so this matcher can be wide (`mcp__.*memvara.*`)
without approving a forget.

Each read it approves is also counted as one `searched` for the status line, in
`~/.memvara/.hooks/counts/<session>.json`, unless the `status_line` setting is off.
"""

from __future__ import annotations
Expand All @@ -15,11 +18,15 @@

from core.envelope import read_event, write # noqa: E402
from core.host import Reply, active # noqa: E402
from lib import counts # noqa: E402
from lib.ipc import payload # noqa: E402

#: Every memory_* tool the server marks `readOnlyHint`. A read that prompts is a read the
#: model learns to avoid, and the two graph tools were missing for no reason other than
#: that they were added after this list.
#: that they were added after this list. `memory_standing` and `memory_ask` were missing for
#: the same reason, and the memory-research subagent calls both, so it stopped at a
#: permission prompt on its first search. `memory_profile` is listed before the server
#: ships it so that the subagent can call it the day it does.
READ_ONLY = frozenset({
"memory_recall",
"memory_search",
Expand All @@ -29,6 +36,9 @@
"memory_stats",
"memory_neighborhood",
"memory_paths",
"memory_standing",
"memory_ask",
"memory_profile",
})


Expand All @@ -45,12 +55,14 @@ def main() -> int:
if host.approve is None:
# No pre-tool event on this client: there is no prompt to pre-empt.
return 0
leaf = _tool_leaf(read_event(host, "approve", payload()).tool_name,
host.approve.separators)
event = read_event(host, "approve", payload())
leaf = _tool_leaf(event.tool_name, host.approve.separators)
if leaf not in READ_ONLY:
return 0
write(host, Reply("approve", decision=host.approve.allow,
reason="Memvara recall is read-only."))
if counts.enabled():
counts.bump(event.session, "searched")
return 0


Expand Down
73 changes: 64 additions & 9 deletions hooks/capture.py
Original file line number Diff line number Diff line change
@@ -1,8 +1,18 @@
#!/usr/bin/env python3
"""Stop — mine the turn that just ended for anything worth knowing next week.

This runs once per turn and looks at one turn: the prompt the user typed and the reply it
got. Nothing earlier, because the earlier turns were mined when they happened.
This runs once per turn and mines one turn: the prompt the user typed and the reply it
got. Nothing earlier is mined, because the earlier turns were mined when they happened.

**Agentic capture is the default way to mine it** (`lib/agentic.py`, switch
`agentic_capture`). The headless agent command gets read-only access to the user's memory
for one run, searches it a few times, and returns proposals: a new fact, a replacement of
a stored claim, the end of a stored claim, or a link between two claims. This hook checks
each proposal and applies the ones that pass through its own write paths. The model is
also shown up to 4,000 characters of the turns before this one, marked as already mined,
so that a short reply can be read against the question it answers; nothing is taken from
that window. When the agentic run cannot use the store, fails or times out, the turn gets
the single-call extraction described below instead, and `capture.log` says so.

It mines both halves because they hold different things. The prompt carries standing
instructions, the reply carries what was decided and where it landed, and a fact usually
Expand Down Expand Up @@ -46,6 +56,11 @@
and a refusal raises rather than returning quietly. See `lib/write.py`.
* **It repeats.** `Stop` can fire more than once over one reply, so the size of the
transcript at the last run is recorded and an unchanged size means there is nothing new.

Two smaller jobs ride along. The hosted client sends the project worked out from the
repository's remote (`lib.project.bind`) with every write. And the number of facts a turn
stored is added to the session's `captured` count for the status line (`lib.counts`), after
the write succeeds and never before.
"""

from __future__ import annotations
Expand All @@ -59,9 +74,11 @@

from core.envelope import read_event # noqa: E402
from core.host import active # noqa: E402
from lib import agentic, counts, settings # noqa: E402
from lib.extract import project_subject, triples # noqa: E402
from lib.ipc import payload # noqa: E402
from lib.transcript import last_turn_with_injections # noqa: E402
from lib.project import bind as bind_project # noqa: E402
from lib.transcript import last_turn_with_context # noqa: E402
from lib.write import (EPISODE_ROLE, log, open_writer, store_facts, # noqa: E402
turn_ids)

Expand Down Expand Up @@ -194,23 +211,27 @@ def _write_state(state: dict) -> None:
pass


def _turn(transcript: Path) -> "tuple[str, list[str]]":
"""The turn that just ended, and the memories this plugin injected into it.
def _turn(transcript: Path) -> "tuple[str, list[str], str]":
"""The turn that just ended, the memories this plugin injected into it, and context.

The second half is not decoration. Recall puts stored notes in front of the model
before it replies; if the reply restates one, mining it writes the store's own output
back into the store as though it were something new. Handing them to the extractor is
what lets it tell an observation from an echo.

The third is up to `agentic.CONTEXT_CHARS` of the turns before this one. Only agentic
capture reads it, as reference: those turns were mined when they ended.
"""
try:
size = transcript.stat().st_size
with transcript.open("rb") as fh:
fh.seek(max(0, size - TAIL_BYTES))
raw = fh.read()
except OSError:
return "", []
text, injected = last_turn_with_injections(raw)
return (text[-MAX_TURN_CHARS:] if len(text) > MAX_TURN_CHARS else text), injected
return "", [], ""
text, injected, context = last_turn_with_context(raw, agentic.CONTEXT_CHARS)
text = text[-MAX_TURN_CHARS:] if len(text) > MAX_TURN_CHARS else text
return text, injected, context


def main() -> int:
Expand All @@ -223,6 +244,9 @@ def main() -> int:
if event.reentrant:
# Re-entry from a hook-triggered continuation. Mining here would double-count.
return 0
# Before anything else that can be slow: a config an earlier, killed capture left
# behind holds a credential, and this is the next moment anything can remove it.
agentic.sweep_configs()

if not event.transcript_path:
return 0
Expand All @@ -240,7 +264,7 @@ def main() -> int:
# Stop fired twice over one reply. Nothing has been added since the last run.
return 0

turn, injected = _turn(transcript)
turn, injected, context = _turn(transcript)
if not turn.strip():
log("no turn to mine")
return 0
Expand All @@ -255,13 +279,26 @@ def main() -> int:
log(f"turn={len(turn)}c skipped={why}")
return 0

# Before the store is opened: the hosted client sends this project with every write.
bind_project(event.cwd)
store, close = open_writer()
if store is None:
log(f"turn={len(turn)}c stored=0 failed=no store or login")
return 0

try:
kept, turn_of = _keep_turn(store, turn, event.cwd)
outcome = None
if settings.enabled("agentic_capture"):
# None means the agentic run could not use the store, and it has already
# logged why. The turn then gets today's single-call extraction instead.
outcome = agentic.capture(store, turn, context, event.cwd or None, injected,
hosted=close is not None, sources=turn_of)
if outcome is not None:
_log_agentic(turn, outcome, kept)
if outcome.applied.stored and counts.enabled():
counts.bump(event.session, "captured", outcome.applied.stored)
return 0
facts = triples(turn, event.cwd or None, injected=injected)
if not facts:
log(f"turn={len(turn)}c facts=0 episode={'yes' if kept else 'no'}")
Expand All @@ -275,10 +312,28 @@ def main() -> int:
log(f"turn={len(turn)}c facts={len(facts)} stored={stored} "
f"episode={'yes' if kept else 'no'}"
+ ("; failed=" + "; ".join(failed) if failed else ""))
if stored and counts.enabled():
counts.bump(event.session, "captured", stored)

return 0


def _log_agentic(turn: str, outcome: "agentic.Outcome", kept: bool) -> None:
"""The one line an agentic turn leaves in capture.log, in the single-call line's shape.

`stored` counts facts written, new or replacing an old value, which is also what the
status line's `captured` count adds; `replaced` says how many of those ended a stored
claim by id. Ends and links are counted separately because they write no fact.
"""
applied = outcome.applied
log(f"turn={len(turn)}c agentic searches={outcome.searches} "
f"proposals={outcome.proposed} refused={outcome.refused} "
f"stored={applied.stored} replaced={applied.replaced} ended={applied.ended} "
f"linked={applied.linked} "
f"episode={'yes' if kept else 'no'}"
+ ("; failed=" + "; ".join(applied.failed) if applied.failed else ""))


def _keep_turn(store: object, turn: str, cwd: str) -> "tuple[bool, list[str]]":
"""Store the turn itself as an episode. `(landed, ids of the turn)`.

Expand Down
57 changes: 42 additions & 15 deletions hooks/daemon.py
Original file line number Diff line number Diff line change
Expand Up @@ -47,6 +47,7 @@

sys.path.insert(0, os.path.dirname(os.path.abspath(__file__)))

from lib.fast import read_kinds, text_of # noqa: E402
from lib.ipc import IDLE_TIMEOUT_SEC, socket_path, store_key # noqa: E402
from lib.open import open_store # noqa: E402

Expand Down Expand Up @@ -89,6 +90,11 @@ def __init__(self, path: str, store: object) -> None:
self.last_seen = time.monotonic()
self._lock = threading.Lock()
self.failures = 0
#: The keyword arguments for a plain read and for a rewritten read of this store,
#: decided once here rather than on every request (`lib.fast.read_kinds`). An
#: empty `rewrite_read` means this store cannot rewrite, which is true of the
#: hosted client and of every library released before query rewrite.
self.plain_read, self.rewrite_read = read_kinds(store)

# -- serving ---------------------------------------------------------------

Expand Down Expand Up @@ -134,21 +140,35 @@ def _answer(self, request: dict) -> dict:
types = request.get("memory_types")
if isinstance(types, list) and types:
kwargs["memory_types"] = [str(t) for t in types]
# A plain read unless the client asked for a rewrite and this store can do one --
# the same rule `lib.fast.recall` applies on the direct route, so the two routes
# hand one backend the same call.
rewrite = bool(request.get("query_rewrite")) and bool(self.rewrite_read)
read_kind = self.rewrite_read if rewrite else self.plain_read
try:
# Serialised deliberately. The store is a read handle over SQLite and is not
# documented as thread-safe; a per-prompt hook has no concurrency worth the
# risk of finding out otherwise.
# Both backends answer the same call. The local one is a `Memvara`; the hosted
# one is a `HostedRecall` holding a kept-alive TLS connection, which is the
# whole reason a hosted install wants a daemon: the same request costs 609ms on
# a fresh connection and 177ms on a warm one. Both raise on failure and return
# a result on success, which is what lets one `except` cover both backends
# without knowing which one it holds. `text_of` notes a rejected key.
if rewrite:
# Outside the lock, because this read waits on a model call for up to its
# 10-second deadline, and every other client of this daemon -- a second
# session on the same project -- would wait behind it past its 2-second
# timeout and fall back to the slow route even for a plain read. Only a
# library store is ever handed a rewrite, and `SQLiteStore` gives each
# thread its own reader connection, so this read can run beside another.
text = text_of(self.store.recall(query, **kwargs, **read_kind))
else:
# Serialised deliberately, and still needed: the hosted client holds one
# kept-alive connection, which two threads must not use at once. Nothing
# under this lock waits on a model.
with self._lock:
text = text_of(self.store.recall(query, **kwargs, **read_kind))
with self._lock:
# Both backends answer the same call. The local one is a `Memvara`; the
# hosted one is a `HostedRecall` holding a kept-alive TLS connection,
# which is the whole reason a hosted install wants a daemon: the same
# request costs 609ms on a fresh connection and 177ms on a warm one.
#
# Both raise on failure and return text on success, which is what lets one
# `except` cover both backends without knowing which one it holds.
text = str(self.store.recall(query, **kwargs) or "")
self.failures = 0
return {"ok": True, "text": text}
return {"ok": True, "text": text}
except Exception:
# Still never a raised exception out of here -- but no longer an empty string
# either, because the client cannot act on what it cannot see.
Expand Down Expand Up @@ -268,7 +288,8 @@ def run(self) -> int:

def main() -> int:
store = open_store()
if store is None:
hosted = store is None
if hosted:
# No library, or no local store. On a paste-the-URL hosted install that is the
# normal state, not a broken one, so fall through to the stdlib HTTP client
# rather than exiting.
Expand All @@ -280,14 +301,20 @@ def main() -> int:
# accept connections and answer every one with silence, which is indistinguishable
# from a working daemon over a store that happens to be empty.
return 0
served = Daemon(socket_path(store_key()), store)
# The warm-up is a plain read: it exists to pay connection costs, not a model call.
# The library's store is told so with the plain read the daemon decided on at
# startup; the stdlib hosted client takes no such argument and always asks its server
# for a plain read.
plain_read = {} if hosted else served.plain_read
try:
# Pay the first-query costs -- imports, page cache, TLS handshake -- before any
# prompt is waiting on them. For hosted this is the handshake that turns a 609ms
# first call into a 177ms one.
store.recall("warm", k=1)
store.recall("warm", k=1, **plain_read)
except Exception:
pass
return Daemon(socket_path(store_key()), store).run()
return served.run()


if __name__ == "__main__":
Expand Down
6 changes: 5 additions & 1 deletion hooks/hosts/claude.py
Original file line number Diff line number Diff line change
Expand Up @@ -49,7 +49,11 @@
detach_capture=False,
#: This client imposes no ceiling of its own, so nothing is declared to it.
context_limit_key=0,
timeouts={"session_start": 20, "recall": 10, "capture": 120, "approve": 5},
#: `capture` covers an agentic run (`lib.agentic.TIMEOUT_SEC`, 60s) followed, when that
#: run fails, by the single-call extraction (`lib.extract.TIMEOUT_SEC`, 90s), plus the
#: writes. Only this host runs agentic capture, because only here is `claude` the
#: first extractor. The hook is async, so the longer limit holds no turn open.
timeouts={"session_start": 20, "recall": 10, "capture": 180, "approve": 5},
client_configs=("~/.claude.json", "~/.claude/settings.json"),
config_format="json",
transcript=TranscriptSpec(format="jsonl"),
Expand Down
Loading
Loading