Skip to content

Agent workspace: native /model and /effort, Claude mods, context panel - #347

Draft
raiseCatError wants to merge 15 commits into
release/v0.18.0from
feature/agent-workspace
Draft

raiseCatError wants to merge 15 commits into
release/v0.18.0from
feature/agent-workspace

Conversation

@raiseCatError

Copy link
Copy Markdown
Owner

Native controls for NMSh-managed Claude Code sessions. Draft: the live plugin-toggle probe still needs the maintainer's disposable login.

Base: release/v0.18.0 (#345). Merge #345 first; this PR then retargets to master. It is independent of the input-awareness PR, and the two merge cleanly in either order.

What it adds

  • /model and /effort as native pickers over Claude's documented control requests (set_model, apply_flag_settings). A change shows as "requested" until Claude reports it.
  • Claude mods: real plugin types and states per launch identity, from claude plugin list --json. Enable and disable go through Claude's own command with the installed scope, then reload_plugins for running managed sessions of that account. NMSh never edits Claude's settings files.
  • A right-hand context panel showing what the target is doing to its context and the project (/context: Claude's own breakdown).
  • A control plane and runtime telemetry with provenance: every fact says whether it was provided, requested, observed or estimated.
  • Untrusted-text Markdown for agent output: blocks, inline styles, code highlighting, real link destinations revealed, smuggled characters refused, linear parsing.
  • Results: a result is read only as the tool that produced it; offers are exact and last only for the session.

Design: docs/design/agent-workspace-controls.md.

Safety

  • Probes in scripts/probes/ run only against disposable state. The plugin probe refuses every real identity, resolved through symlinks: ~/.claude, every ~/.claude-* sibling and an inherited CLAUDE_CONFIG_DIR. It accepts only /tmp/nmsh-cc-probe, and hashes the real accounts' settings before and after.
  • The one test that reads real Claude accounts is opt-in (NMSH_REAL_CLAUDE=1) and read-only.

Verification

  • npm run verify on this head: see the CI checks below. A local run is noted in the comments once complete.
  • tests/intelligenceHardening.test.ts "native fuzzy filtering…" asks the real zsh for completions under a time budget, and it can time out under full-suite load. It also does this on release/v0.18.0, unrelated to this PR, and passes 5/5 alone.

Still needs physical validation

  • Live plugin enable/disable inside a managed session (scripts/probes/managed-plugins-live.ts), after CLAUDE_CONFIG_DIR=/tmp/nmsh-cc-probe claude auth login.
  • /model, /effort, /context and the right-hand panel in Ghostty and an integrated terminal, at normal and narrow widths, with NO_COLOR and Safe glyphs.

…son transport

Records the initialize metadata (model catalog, commands, agents), get_context_usage,
mcp_status, set_model, apply_flag_settings (effort), set_permission_mode and the
runtime facts that confirm each change, in a disposable workspace with customizations
disabled and a budget cap. Account identity values are never printed.
… and bounded rows

Agent replies, provider command output and GitHub bodies are Markdown from outside
NMSh. This renders them as terminal rows that always fit their width: headings,
nested and task lists, quotes, GFM tables that wrap cells (records when far too
narrow), fenced code with a gutter and language-aware highlighting from the
person's syntax roles, and rules. Links become OSC 8 hyperlinks only for http(s)
targets and always show where they go; refused targets stay visible as text.
Safe glyphs and NO_COLOR keep the structure with ASCII markers and backticks.
The authored-only /help renderer is unchanged.
…keep parsing linear

A security review of the untrusted-Markdown renderer found three problems, fixed here:

- Link text could hide a different destination when the target's host was only a
  substring of the text (example.com -> le.com). Links are now judged on their whole
  text: no destination is revealed only when the text names exactly the target's host
  as a whole hostname and no other; internationalized (punycode) hosts are always
  shown.
- Numeric entities decoded after the caller's scrubber, so ‮ and similar could
  reintroduce bidi overrides or invisible characters. Entities now refuse controls,
  format characters, separators and private use, and the renderer scrubs its own
  input instead of trusting callers.
- Several recognizers backtracked quadratically on long lines (heading, fence, rule,
  hard and soft breaks, YAML keys, code spans, emphasis, link brackets, long-word
  breaking). They are linear scans or bounded searches now, with timing tests on
  pathological inputs.
The adapter now speaks the documented Agent SDK control requests over its existing
stream-json channel, verified against Claude Code 2.1.295 by the control probe:

- set_model, apply_flag_settings (session-only effortLevel), set_permission_mode,
  get_context_usage, mcp_status/toggle/reconnect and stop_task, each correlated by
  request id with a timeout, waiting for the initialize handshake. Bypassing every
  permission check is never sent; refusals are reported as the provider worded them.
- The initialize response is kept: the model catalog with per-model effort levels,
  the provider's commands (built-in or not), agents, output styles and account plan.
- Telemetry with provenance (provider, requested, observed, estimate): runtime facts
  from each turn's init, usage and cost, rate limits (shared only with targets of
  the same launch profile, labelled), compaction, retries, tasks, the TodoWrite plan
  and the files tools read and edited, with the provider's own patch hunks.
- Structured tool results (patches, reads, command output, searches, plans,
  subagent reports), bounded, never keeping whole file contents.
- Permission requests carry the provider's reason and its own rule offers; an
  accepted offer is returned verbatim, never composed by NMSh.
- Streaming text (--include-partial-messages) is presentation only and is replaced
  by the completed message; a profile's effort is a launch flag.
…s are exact and session-only

An MCP tool returning a result shaped like an edit could have appeared as a diff and a
file activity. Structured results are now bound to the built-in tool that produced them.
Permission rule offers are rebuilt from exactly the fields shown, limited to this session,
allow-only, and never switch to a mode that bypasses checks.
In a managed Claude composer, /model and /effort open NMSh's own picker in the
composer's place instead of going to Claude's prompt. Rows are what Claude published:
its model catalog, and only the effort levels the running model supports (none for a
model without levels, which the picker says). Up/Down choose, typing filters, Enter
applies through set_model or apply_flag_settings, Esc backs out. /model sonnet and
/effort high apply directly on an exact match; a leading space sends text verbatim.

Nothing is claimed without evidence: a model change reads "requested" until Claude
reports it; effort is shown as acknowledged, set at launch, read from the account's
settings, or the model default (unreported) - never invented. The header shows model
and effort with that provenance. A refusal keeps the picker open with the reason.
A FIFO or very large file at a project's .claude/settings.json could block or flood
the effort lookup. The path is checked with lstat first.
…and, live reload

/claude mods and /mods now show each account's Claude plugins as they are:
- Type from the plugin's own manifest files, read as data (nothing is imported or
  run): Mod (hook modules), Plugin, Settings hook or Skill, with its components.
- The effective state for this directory. The old inventory read projectEnabled,
  which is false for nearly every installed plugin, so everything showed disabled;
  project and local enabledPlugins overrides are now read and named instead.
- Conflicts (also installed at another scope, other version), dependencies and
  options, in words; project installations of other directories are not shown.
- What it does in an NMSh-managed session, per Claude's documented behavior for
  claude -p: a mod's hooks run, nothing it draws appears; skills, agents, commands,
  hooks, MCP and LSP load. "Loaded in" names targets whose own report lists it.

Space proposes a change and Enter confirms it: claude plugin enable|disable at the
plugin's installed scope and account (CLAUDE_CONFIG_DIR), --json. Only ids from the
current listing are sent, since Claude's command writes an unknown id into settings.
The listing is read again as evidence, then running managed sessions on that account
reload plugins (reload_plugins) and the result is reported, errors included.
Standalone skills and settings hooks are listed without a switch, with the reason.
i inspects through claude plugin details and validate (static). No sandbox is implied.
…and the project

From telemetry already collected (nothing read on the render path): context pressure
with its source and age (Claude's own count with headroom to auto-compact and the
largest categories, or an estimate from the last turn, never 0% when unreported),
model, effort and permission mode with provenance, the plan when Claude emits one,
running subagents and tasks, files read and edited with their line changes,
compactions, usage limits only once close, and the session's estimated cost.

It sits beside the conversation from 120 columns; /panel shows or hides it (72
columns at least), and /context asks Claude for its breakdown. Safe glyphs draw
ASCII marks and bars; NO_COLOR keeps every distinction in words.
Mechanisms chosen for each change, what is operational versus a representation of
provider facts versus unsupported, the research conclusions that shaped the panel,
and the next steps (left sidebar, per-turn diffs, MCP status, prompt navigation).
… real Claude

Runs the picker logic and the documented control requests against a logged-in
Claude Code in safe mode (no plugins, hooks or mods run, no session persistence),
and proves the account's settings.json is unchanged. Verified with 2.1.295: the
catalog opens the picker, set_model reads as requested until the next turn reports
the model, effort is acknowledged (never claimed as confirmed), and an unknown model
is refused in Claude's own words.
…able configuration

Installs a code-free probe plugin from a local marketplace into a disposable,
separately logged-in Claude configuration, starts a managed session that loads it,
then disables and re-enables it through NMSh's own path (listing, setPluginEnabled,
reload_plugins) and checks Claude's own reports: plugin list, offered commands, load
errors. Refuses any non-temp or account-looking directory and hashes the real
accounts' settings before and after.
…on through any path

The plugin probe now accepts only /tmp/nmsh-cc-probe, resolved through symlinks, and
refuses a directory that equals, contains or lies inside ~/.claude, ~/.claude-account1,
~/.claude-account2 or an inherited CLAUDE_CONFIG_DIR. Checked before any claude
command runs, since even auth status writes into a fresh directory.
…ounts

The plugin probe's guard and settings hashes listed two personal account
directories by name; it now discovers the default ~/.claude, every
~/.claude-* sibling and an inherited CLAUDE_CONFIG_DIR, so it refuses any
real identity on any machine. The picker probe's usage line no longer
points at a real account.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant