Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
12 changes: 12 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
@@ -1,5 +1,17 @@
# Changelog

## [0.2.14] — 2026-09-09

- Gate and outsider CLI recompute MIRROR-SPEC content hashes and links from a
single snapshot. Missing, empty, malformed, duplicate-key and tampered ledgers fail closed.
- Publish requires a reasoned retraction or explicit `action=result` with
`payload.status`, `payload.summary`, and `payload.prereg_seal` binding the first registration.
Merely starting work no longer resolves a claim. Negative results remain publishable.
- Expose verification depth and unverified content truth, author identity,
external time and independent reproduction. A separately checked OTS proof is not
presented as proof of this ledger's precedence without checking its binding.
- Migration: append a real explicit result; do not rewrite historical sealed actions.

All notable changes to mirror-stack-mcp are documented here.
Format follows [Keep a Changelog](https://keepachangelog.com/en/1.0.0/).

Expand Down
28 changes: 26 additions & 2 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -36,6 +36,30 @@ Windsurf, any MCP client.

## Tools (19)

### Verification depth and gate migration (0.2.14)

`mirror-stack-verify` now recomputes both content hashes and chain links; pointer-only
fixtures do not certify integrity. Legacy 16-hex seals are accepted with a warning.
`stack_verify_all` labels checks `HASH_RECOMPUTED`, `LOCAL_SNAPSHOT`, or `HEAD_WITNESS`.
None certifies content truth, author identity, independent reproduction or external time.
The outsider CLI checks an optional OTS proof separately; ledger-to-proof binding is
not checked, so it does not claim this ledger's external-clock precedence.

The compute/publish gate verifies ledger integrity before reading the decision facts.
Publish requires a reasoned retraction, or an explicit result bound to the original seal:

```python
am_record(ledger_path="actions.jsonl", agent="researcher", action="result", target="claim-id",
payload={"status": "fail", "summary": "Measured value did not meet the preregistered bar",
"prereg_seal": "<seal returned by the first mm_preregister>"})
```

Allowed statuses are `pass`, `fail`, and `inconclusive`. A started action does not
resolve a claim. Append a new result instead of rewriting historical records. GO permits
reporting a resolved result, including a failure; it does not certify scientific success.

### Tool reference

| Tool | Mirror | Does |
|---|---|---|
| `mm_preregister` | 🪞 claims | seal a claim + kill-condition **before** measuring (response carries an auto seal-quality lint) |
Expand Down Expand Up @@ -160,7 +184,7 @@ into the action site:
kill-condition* is sealed. Wire it into your training/experiment launcher so it refuses to
spend compute on an unsealed claim.
- `mm_preflight(ledger, claim_id, gate="publish", am_ledger=…)` → **BLOCK** unless the *result*
is also sealed (a retraction, or `am_record(target=claim_id)`). Wire it into a pre-commit /
is also sealed (a reasoned retraction, or the explicit bound result above). Wire it into a pre-commit /
pre-publish hook so unresolved claims can't ship.

The MCP only *judges* GO/BLOCK — **your** launcher/hook does the actual blocking. The server
Expand All @@ -174,7 +198,7 @@ the part that actually exits non-zero, so a shell can do the blocking the MCP ca
mirror-stack-gate compute --ledger L.jsonl --claim my_claim && python run.py
# exits 1 (run.py never starts) unless a kill-conditioned preregistration is sealed
mirror-stack-gate publish --ledger L.jsonl --claim my_claim --am-ledger A.jsonl
# exits 1 unless the result is sealed too (a retraction or am_record(target=claim))
# exits 1 unless a reasoned retraction or an explicit bound result is sealed too
```

For git, drop in [`hooks/pre-commit.sample`](hooks/pre-commit.sample): it runs the publish gate
Expand Down
6 changes: 6 additions & 0 deletions docs/STACK_CANONICAL.md
Original file line number Diff line number Diff line change
@@ -1,5 +1,11 @@
# Mirror Stack — canonical surfaces & the one shared primitive

Since 0.2.14, `check_chain()` remains the linkage-only compatibility API, but the
outsider CLI and gate use `integrity.read_verified()` for a single snapshot with
both links and MIRROR-SPEC content hashes verified. Signature identity, external
time and content truth remain outside this check. The directory orchestrator
explicitly labels its automatically included ledgers as linkage-only.

Two repos make up the running stack. They are **different surfaces, not
duplicates** — keep them separate. This is the canonical map of who owns what,
so "which one is authoritative?" has a written answer.
Expand Down
2 changes: 1 addition & 1 deletion mirror_stack_mcp/__init__.py
Original file line number Diff line number Diff line change
@@ -1,2 +1,2 @@
"""🪞🔎🪪 Mirror Stack unified MCP server."""
__version__ = "0.2.13"
__version__ = "0.2.14"
105 changes: 63 additions & 42 deletions mirror_stack_mcp/gate.py
Original file line number Diff line number Diff line change
Expand Up @@ -12,17 +12,16 @@

publish (resolution-before-publish), e.g. a git pre-commit hook:
mirror-stack-gate publish --ledger L.jsonl --claim my_claim --am-ledger A.jsonl
# exits 1 unless a retraction or an am_record(target=my_claim) is sealed
# requires a reasoned retraction or action=result bound by payload.prereg_seal

The same `decide()` backs server.mm_preflight, so the agent-facing tool and the
shell enforcer can never drift apart.
"""
import json
import os
import sys
from .integrity import read_verified


def scan_claim(ledger_path, claim_id):
def scan_claim(ledger_path, claim_id, entries=None):
"""Return (prereg_entry_or_None, retracted_bool, leaked_entry_or_None) for claim_id.

leaked_entry is a preregistration that carries NO kill fields but whose `metric`
Expand All @@ -31,38 +30,46 @@ def scan_claim(ledger_path, claim_id):
the gate report the real reason instead of a misleading 'no preregistration'.
"""
prereg, retracted, leaked = None, False, None
if os.path.exists(ledger_path):
with open(ledger_path, encoding="utf-8") as fh:
for line in fh:
line = line.strip()
if not line:
continue
try:
e = json.loads(line)
except json.JSONDecodeError:
continue
if e.get("claim_id") != claim_id:
continue
if e.get("_type") == "retraction":
retracted = True
elif e.get("_type") is None and ("kill_threshold" in e or "kill_condition" in e) \
and e.get("metric") != "protocol_amendment":
prereg = e
elif e.get("_type") is None and e.get("metric") not in (None, "protocol_amendment"):
from measure_mirror import mm
if leaked is None and mm._looks_like_kill_prose(e.get("metric", "")):
leaked = e
seen_registration = False
if entries is None:
entries, error = read_verified(ledger_path)
if error:
return None, False, None
for e in entries:
if e.get("claim_id") != claim_id:
continue
if e.get("_type") == "retraction":
retracted = bool(prereg and isinstance(e.get("reason"), str)
and e["reason"].strip()) or retracted
elif e.get("_type") is None and e.get("metric") != "protocol_amendment":
if seen_registration:
continue
seen_registration = True
if "kill_threshold" in e or "kill_condition" in e:
prereg = e # First-write wins, even when the first registration is invalid.
else:
from measure_mirror import mm
if isinstance(e.get("metric"), str) and mm._looks_like_kill_prose(e["metric"]):
leaked = e
return prereg, retracted, leaked


def decide(ledger_path, claim_id, gate="compute", am_ledger=None, reported_acc=None):
"""Pure GO/BLOCK decision. Returns {decision, gate, claim_id, reasons, checks}."""
prereg, retracted, leaked = scan_claim(ledger_path, claim_id)
entries, error = read_verified(ledger_path)
prereg, retracted, leaked = scan_claim(ledger_path, claim_id, entries)
checks: list[str] = []

def out(decision, reasons):
return {"decision": decision, "gate": gate, "claim_id": claim_id,
"reasons": reasons, "checks": checks}
"reasons": reasons, "checks": checks,
"verification": {"depth": "HASH_RECOMPUTED" if not error else "UNVERIFIED",
"content_truth": "unverified", "external_time": "unverified",
"independent_reproduction": "unverified"}}

if error:
return out("BLOCK", ["claims ledger integrity failed: " + error])
checks.append("claims ledger: linkage and all content hashes verified")

if prereg is None:
if leaked is not None:
Expand All @@ -77,14 +84,19 @@ def out(decision, reasons):
checks.append("preregistration: sealed" + ("" if has_kill else " (NO kill-condition)"))

if gate == "compute":
if retracted:
return out("BLOCK", ["claim is retracted; use a new preregistration"])
if not has_kill:
return out("BLOCK", ["preregistration has no kill-condition (unfalsifiable) — "
"add one before spending compute"])
# Seal-quality lint: a FAIL (e.g. a pass bar at/below chance) means the
# automated checks can't do their job — compute would be spent against a
# meaningless bar. WARN/INFO inform but don't block.
from measure_mirror import mm
lint = mm._preseal_lint(prereg)
try:
lint = mm._preseal_lint(prereg)
except (TypeError, ValueError, AttributeError, KeyError) as exc:
return out("BLOCK", [f"preregistration cannot be evaluated: {exc}"])
fails = [f for f in lint if f.level == "FAIL"]
warns = [f for f in lint if f.level == "WARN"]
if warns:
Expand All @@ -98,25 +110,34 @@ def out(decision, reasons):
resolved = retracted
if retracted:
checks.append("resolution: retraction sealed")
if not resolved and am_ledger and os.path.exists(am_ledger):
with open(am_ledger, encoding="utf-8") as fh:
for line in fh:
try:
a = json.loads(line)
except json.JSONDecodeError:
continue
if a.get("_type") == "action" and a.get("target") == claim_id:
resolved = True
checks.append("resolution: am_record(target) sealed")
break
if not resolved and am_ledger:
actions, action_error = read_verified(am_ledger)
if action_error:
return out("BLOCK", ["action ledger integrity failed: " + action_error])
checks.append("action ledger: linkage and all content hashes verified")
for a in actions:
payload = a.get("payload")
if (a.get("_type") == "action" and a.get("target") == claim_id
and a.get("action") == "result" and isinstance(payload, dict)
and payload.get("status") in ("pass", "fail", "inconclusive")
and isinstance(payload.get("summary"), str) and payload["summary"].strip()
and payload.get("prereg_seal") == prereg["seal"]):
resolved = True
checks.append("resolution: explicit result bound to verified preregistration")
break
if not resolved:
return out("BLOCK", ["no sealed resolution — seal the result "
"(am_record target=claim_id, or mm_retract) before publishing. "
"(am_record action=result, target=claim_id, payload with status, "
"summary, prereg_seal; or mm_retract with reason) before publishing. "
"Prose doesn't count."])
if reported_acc is not None:
from measure_mirror import mm
checks.append(str(mm.falsifiability_check(ledger_path, claim_id, reported_acc=reported_acc)))
return out("GO", ["sealed preregistration + sealed resolution"])
# Evaluate the already-verified snapshot, not a second mutable file read.
finding = mm._falsifiability_eval(prereg, reported_acc)
checks.append(str(finding))
checks.append("publication may report a negative result; GO is NOT claim success")
return out("GO", ["verified preregistration + explicit sealed resolution; "
"publication permitted, content truth not certified"])

return out("BLOCK", [f"unknown gate '{gate}' — use 'compute' or 'publish'"])

Expand Down
49 changes: 49 additions & 0 deletions mirror_stack_mcp/integrity.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,49 @@
"""Read and verify one immutable-in-memory MIRROR-SPEC ledger snapshot.

Hash integrity is not signature identity, external timing, or content truth.
Both the gate and outsider CLI use this path; neither certifies pointer linkage alone.
"""
import hashlib
import json
from pathlib import Path


def _object(pairs):
out = {}
for key, value in pairs:
if key in out:
raise ValueError(f"duplicate JSON key: {key}")
out[key] = value
return out


def read_verified(path):
"""Return (entries, error). Fail closed on missing, empty or corrupt input.

Reads once so decisions inspect exactly the snapshot whose hashes were checked.
Accepts legacy 16-hex seals at their original, weaker assurance level.
"""
try:
raw = Path(path).read_text(encoding="utf-8")
entries = [json.loads(line, object_pairs_hook=_object)
for line in raw.splitlines() if line.strip()]
if not entries:
return [], "ledger is empty; nothing verified"
previous = "genesis"
for i, entry in enumerate(entries, 1):
if not isinstance(entry, dict):
return [], f"entry {i}: JSON object required"
link, seal = entry.get("prev_seal"), entry.get("seal")
if not isinstance(link, str) or (link.lower() != "genesis" if i == 1 else link != previous):
return [], f"entry {i}: chain linkage broken"
if not isinstance(seal, str) or len(seal) not in (16, 64):
return [], f"entry {i}: missing or invalid seal"
body = {k: v for k, v in entry.items() if k not in ("seal", "sig")}
digest = hashlib.sha256(json.dumps(body, sort_keys=True, ensure_ascii=False,
allow_nan=False).encode("utf-8")).hexdigest()
if seal != digest[:len(seal)]:
return [], f"entry {i}: seal mismatch; content modified"
previous = seal
return entries, None
except (OSError, UnicodeError, ValueError, TypeError, RecursionError) as exc:
return [], f"ledger cannot be verified: {exc}"
31 changes: 20 additions & 11 deletions mirror_stack_mcp/server.py
Original file line number Diff line number Diff line change
Expand Up @@ -137,7 +137,8 @@ def _compact(findings):

REMINDERS = {
"mm_preregister": "🪞 Sealed. Not done until the RESULT is sealed too — "
"am_record(target=claim_id) on a verdict, or mm_retract if falsified. Prose doesn't "
"am_record(action=result, target=claim_id, payload={status, summary, prereg_seal}) "
"on a verdict, or mm_retract if falsified. Prose doesn't "
"count. Your kill_condition is the stop-loss; if big compute follows, seal first, then run.",
"mm_verify": _VERIFY,
"mm_audit": _VERIFY,
Expand Down Expand Up @@ -373,7 +374,11 @@ def mm_preflight(ledger_path: str, claim_id: str, gate: str = "compute",
gate="compute": GO only if a sealed preregistration WITH a kill-condition exists for
claim_id (enforces seal-before-compute).
gate="publish": additionally GO only if a RESOLUTION is sealed — a retraction in
ledger_path, or an am_record(target=claim_id) in am_ledger.
ledger_path (with reason), or am_record(action=result, target=claim_id)
with payload status=pass/fail/inconclusive, nonempty summary,
and prereg_seal matching the first verified registration.
Both ledgers must pass hash and linkage verification. GO authorizes publication
of a resolved result, including failures; it does NOT certify claim success.
This is a PRIMITIVE: the MCP returns GO/BLOCK; YOUR script must do the actual blocking
(the MCP cannot intercept external compute or commits — that is by design). The shell
enforcer that DOES exit non-zero is `mirror-stack-gate` (mirror_stack_mcp.gate); both
Expand Down Expand Up @@ -442,15 +447,15 @@ def stack_verify_all(mm_ledger: str, anchor_dir: str | None = None,
def add(level, layer, name, msg):
nonlocal ok
ok = ok and level
out.append({"ok": level, "layer": layer, "name": name, "msg": msg})
depth = {"L1 chain": "HASH_RECOMPUTED", "L3 anchor": "LOCAL_SNAPSHOT",
"L2 witness": "HEAD_WITNESS"}[layer]
out.append({"ok": level, "layer": layer, "name": name, "msg": msg,
"depth": depth})

findings = mm.verify_chain(mm_ledger)
bad = [str(f) for f in findings if getattr(f, "level", "OK") not in ("OK", "INFO")]
# Say how many seals were checked. "seals valid" is also true of an empty ledger.
n_entries = sum(1 for l in Path(mm_ledger).read_text(encoding="utf-8",
errors="replace").splitlines() if l.strip())
add(not bad, "L1 chain", Path(mm_ledger).name,
f"seals valid ({n_entries} entries checked)" if not bad else str(bad))
from .integrity import read_verified
entries, error = read_verified(mm_ledger)
add(error is None, "L1 chain", Path(mm_ledger).name,
error or f"seals valid ({len(entries)} entries checked)")

if anchor_dir:
for af in sorted(Path(anchor_dir).glob("anchor_*.json")):
Expand Down Expand Up @@ -486,7 +491,11 @@ def add(level, layer, name, msg):
"scope": {"mm_ledger": Path(mm_ledger).name,
"layers_run": sorted({c["layer"] for c in out}),
"layers_not_requested": [] if anchor_dir else ["L3 anchor"],
"layers_skipped": skipped}}
"layers_skipped": skipped,
"external_time": "unverified",
"author_identity": "unverified",
"independent_reproduction": "unverified",
"content_truth": "unverified"}}
if skipped:
# `skipped` means REQUESTED-but-did-not-run, which is the defect this fixes.
# Not passing `anchor_dir` at all is not that — you did not ask for L3, so it is
Expand Down
Loading
Loading