Skip to content

Fix empty-project index readiness, shell-search nudge, context() budget, and global rules workspace stamp - #145

Merged
escott- merged 5 commits into
mainfrom
claude/exciting-heisenberg-c2mthp
Oct 3, 2026
Merged

escott- merged 5 commits into
mainfrom
claude/exciting-heisenberg-c2mthp

Conversation

@escott-

@escott- escott- commented Oct 3, 2026 •

Copy link
Copy Markdown
Member

Summary

Four agent-facing fixes found while investigating why agents mostly skip ContextStream search, plus a metric to measure the effect. No tool names, descriptions, input schemas or HTTP routes change.

  1. Empty projects no longer report "index ready". The hosted index status labels a project with no index row project_index_state: "ready" / status: "completed" with 0 files. project_index_status_reports_canonical_ready accepted that label (or a stale generation number), so init returned canonical_index_ready: true for a project that was never indexed and sent agents to an empty search. A reported file count of 0 is now authoritative; state/generation still stand in when the response omits the count, and an explicit indexed: true still wins.
  2. Shell code search is nudged in a checkout with no recorded index. Bash rg / grep -r / find -name only got a nudge when the checkout was in indexed-projects.json. Otherwise the hook took the initial-index-wait path, which has no Bash arm, and fell through to a silent allow. It now emits a non-blocking nudge at most once per 10 minutes per checkout. It does not tell the agent to search (that would hit an empty index); it says no index is recorded and to run project(action="index"). Piped filters, log greps and single-file greps are still left alone.
  3. context() keeps its actionable fields under the wire budget. The budget stripped every structured field, including instructions, matched_skills, coordination_inbox and grounding_hits, before touching the text that was actually over budget, then removed whole text blocks, leaving a 4k budget under a third used. The large duplicates (items, summary, context) now go first, coordination_inbox is listed explicitly (it used to fall into the lexical pass where it sorts first), low-priority text shrinks before the small fields are dropped when that fits, and a truncated head of the largest block is kept rather than dropping it whole. Under a budget too tight to hold both copies, the old order is kept so the text's instructions outlive their structured duplicates.
  4. Global rules no longer carry a workspace. setup and rules refresh stamped the workspace of whichever directory they ran in into the global rules file, which applies to every project, so the file named a different workspace after each run and disagreed with the project rules beside it. write_editor_rules now writes the neutral block, so a repair also clears an old stamp while keeping the user's own text around the managed block. A global-only refresh skips the API workspace lookup.
  5. testing/adoption/search_adoption.py (+ tests, wired into CI like testing/grounding) counts per agent how often code search goes through ContextStream versus shell rg/grep/find or the native Grep/Glob tools, using a port of the hook's shell classifier. It reads local Claude Code / Codex transcripts and prints counts only; --json writes a report for before/after comparison. The Claude Code reader is checked against a real transcript; the Codex reader is best-effort.

Public-boundary impact: none. No private implementation, operator material or deployment detail is added.

Validation

  • cargo fmt --all --check
  • cargo clippy --locked --workspace --all-targets -- -D warnings
  • cargo test --locked --workspace --all-targets
  • npm test
  • python3 .github/scripts/public_boundary.py .

Not ticked because the exact command was not run. What was run instead:

  • cargo clippy --locked -p mcp-client -p mcp-tools -p mcp-server --all-targets -- -D warnings: clean.
  • cargo test --locked -p mcp-client -p mcp-tools -p mcp-server --lib: mcp-server 961, mcp-tools 1469 and mcp-client 283 pass. 3 mcp-client tests fail (ingest_guard::tests::rejects_home_directory, rejects_sensitive_directory, rejects_path_inside_sensitive_directory); they fail identically on unmodified main when run as root with HOME=/root, so they are environmental. 2 mcp-server watch:: tests also fail when CONTEXTSTREAM_WATCH=0 is set in the environment and pass with it unset.
  • New tests fail on the old code and pass now: the two context() budget tests (oversized_context_pack_does_not_cost_the_small_actionable_fields, budget_is_filled_by_truncating_a_block_rather_than_dropping_it_whole), plus tests for the empty-project readiness, the shell nudge and its throttle, and the global rules rewrite.
  • testing/adoption unit tests (7) and the .github/scripts workflow, public-boundary, release-contract and DCO policy tests pass.

The context() fixture is a hand-built, production-sized payload, not captured production output, so it demonstrates the mechanism rather than reproducing exact production numbers.

Data handling and compatibility

  • The change does not add credentials, customer data, raw local paths, or private deployment topology.
  • Any new collection or transmission of user data is documented in docs/data-handling.md. Nothing new is collected or transmitted: the hook records a local timestamp in ~/.contextstream/prompt-state.json, and the adoption script reads local transcripts and prints counts.
  • Existing npm executable aliases and MCP wire compatibility are preserved. No tool schema changes; the context() change only affects which fields survive when a response exceeds its requested token budget.

Contribution certification

  • Every commit includes a Signed-off-by: trailer (git commit -s).

The commits do not carry Signed-off-by: yet; the author needs to add it (for example git rebase --signoff main) before the DCO check can pass.

🤖 Generated with Claude Code

https://claude.ai/code/session_018gMJ8CEW7U8zTm2GifkwEH


Generated by Claude Code


View with [code]smith Autofix with [code]smith
Need help on this PR? Tag @codesmith-bot with what you need. Autofix is disabled.

escott- and others added 5 commits October 3, 2026 05:26
The hosted index status labels a project with no index row
`project_index_state: "ready"` / `status: "completed"`, with zero files.
project_index_status_reports_canonical_ready accepted the label (or a stale
generation number) even when the response said 0 files, so init set
`canonical_index_ready: true` for projects that were never indexed and agents
were sent to an empty search.

A reported file count of zero is now authoritative. State and generation still
stand in when the response omits the count, and an explicit `indexed: true`
still wins.

Co-Authored-By: Claude <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018gMJ8CEW7U8zTm2GifkwEH
Bash `rg`/`grep -r`/`find -name` was only nudged when the checkout was in the
local indexed-projects registry. For any other checkout the hook took the
initial-index-wait path, whose discovery-tool list has no Bash arm, and fell
through to a silent Allow, so agents never heard anything.

Emit a non-blocking nudge there, at most once per 10 minutes per checkout. It
does not send the agent to search (that would hit an empty index); it says no
index is recorded and to build one with project(action="index"). Commands that
are not code discovery (piped filters, log greps, single targeted files) are
still left alone.

Co-Authored-By: Claude <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018gMJ8CEW7U8zTm2GifkwEH
The whole-wire budget stripped every structured field, including the small
actionable ones (instructions, matched_skills, coordination_inbox,
grounding_hits), before it touched the text that was actually over budget, and
then removed whole text blocks, leaving a 4k budget under a third used.

- The large duplicates (items, summary, context) now go before the small
  actionable fields, and coordination_inbox is listed explicitly instead of
  falling into the lexical pass where it sorts first.
- When the high-priority text plus the remaining structured fields fit, shrink
  the low-priority text first. Under a budget too tight for both copies the old
  order is kept, so the text's instructions still outlive their duplicates.
- Keep a truncated head of the largest block instead of dropping it whole when
  that leaves a useful amount of room.

Also adds a test that init does not report an empty project as ready.

Co-Authored-By: Claude <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018gMJ8CEW7U8zTm2GifkwEH
Global rules apply to every project the editor opens, but setup and rules
refresh stamped them with the workspace of whichever directory they ran in, so
the one global file named a different workspace after each run and disagreed
with the project rules loaded beside it.

write_editor_rules no longer takes a workspace and writes the neutral block, so
a repair also removes an old stamp while keeping the user's own text around the
managed block. A global-only refresh skips the API workspace lookup.

Co-Authored-By: Claude <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018gMJ8CEW7U8zTm2GifkwEH
…nscripts

Counts, per agent, how often code search goes through ContextStream versus shell
rg/grep/find or the native Grep/Glob tools, using a port of the PreToolUse
hook's shell classifier so "shell code search" means what the hook would nudge.
Reads Claude Code and Codex transcripts locally and reports counts only; --json
writes a report for before/after comparison. Wired into CI like the grounding
tests.

Co-Authored-By: Claude <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018gMJ8CEW7U8zTm2GifkwEH
escott- pushed a commit that referenced this pull request Oct 3, 2026
Bumps hyper-util 0.1.21, tiktoken-rs 0.12.1, ignore 0.4.33, tokio-test 0.4.6 and console 0.16.6, with their lockfile updates (bstr, fancy-regex, regex, regex-syntax). The workspace version is unchanged.

Checked on ovh-desktop with #139, #140, #144 and #145 applied together on main 086e8f5: cargo fmt --check, clippy --locked --workspace --all-targets -D warnings, cargo test --locked --workspace --all-targets (2940 passed, 0 failed), account_connection_smoke (9 tests), plus the public-boundary, release-contract, DCO, workflow, grounding and adoption script tests and npm test.

Signed-off-by: escott- <escott05@gmail.com>
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
escott- pushed a commit that referenced this pull request Oct 3, 2026
Bumps actions/checkout 7.0.1, actions/setup-node 7.0.0, actions-rust-lang/setup-rust-toolchain 2.0.0, Swatinem/rust-cache 2.9.2, actions/upload-artifact 7.0.1, actions/download-artifact 8.0.1 and actions/attest-build-provenance 3.0.0. Every pinned SHA was checked against the upstream tag or commit it names. The old rust-cache and attest-build-provenance pins were annotated tag object ids, not commit ids; the new pins are commit ids.

Checked on ovh-desktop with #139, #140, #144 and #145 applied together on main 086e8f5: cargo fmt --check, clippy --locked --workspace --all-targets -D warnings, cargo test --locked --workspace --all-targets (2940 passed, 0 failed), account_connection_smoke (9 tests), plus the public-boundary, release-contract, DCO, workflow, grounding and adoption script tests and npm test.

Signed-off-by: escott- <escott05@gmail.com>
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
escott- added a commit that referenced this pull request Oct 3, 2026
Starting or resuming a second session in the same checkout could reset the shared initialization gate and block an already initialized session with "First call required". Initialization is now tracked by the top-level host session id in a separate, locked state file with atomic replacement. Replayed session-start events keep a completed init, and each fresh session still has to initialize. A successful init or context result is recorded in PostToolUse, and the managed post-tool matcher now includes context and session so quick-start grounding is observed.

Checked on ovh-desktop with #139, #140, #144 and #145 applied together on main 086e8f5: cargo fmt --check, clippy --locked --workspace --all-targets -D warnings, cargo test --locked --workspace --all-targets (2940 passed, 0 failed), account_connection_smoke (9 tests), plus the public-boundary, release-contract, DCO, workflow, grounding and adoption script tests and npm test.

Signed-off-by: escott- <escott05@gmail.com>
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
@escott-
escott- merged commit 7918768 into main Oct 3, 2026
1 of 10 checks passed
escott- added a commit that referenced this pull request Oct 3, 2026
Bumps the workspace, npm package, MCP Registry metadata, and the boundary self-test to 1.0.12, and adds the 1.0.12 changelog covering #139, #140, #144, and #145.

Updates rustls 0.23.40 -> 0.23.45 and rustls-webpki 0.103.13 -> 0.103.15 for RUSTSEC-2026-0285. The advisory predates 1.0.10 and was present in 1.0.10 and 1.0.11. No other lockfile entries change.

Checked on ovh-desktop against main 7918768 plus this change: cargo fmt --check, clippy -D warnings, cargo test (2940 passed, 0 failed), account_connection_smoke (9 tests), cargo audit and cargo deny (clean), plus the public-boundary, release-contract, DCO, workflow, grounding and adoption script tests and npm test.

Signed-off-by: escott- <escott05@gmail.com>
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant