Skip to content

Keep quoted CLI arguments exact, share one CLI parser, and fix vault gym path scoring - #168

Merged
ProfSynapse merged 4 commits into
feat/submodule-cloud-api-v1from
claude/cli-multiline-writes
Oct 2, 2026
Merged

ProfSynapse merged 4 commits into
feat/submodule-cloud-api-v1from
claude/cli-multiline-writes

Conversation

@ProfSynapse

Copy link
Copy Markdown
Owner

Summary

Stacked on #166.

Root cause: multi-line CLI writes were broken

shared/environments/tool_executor.py had three problems:

  • It normalized all whitespace before tokenizing, which flattened newlines inside quoted values. Curly quotes in content were also rewritten as straight ones.
  • It never decoded \n, even though the CLI schema tells models to use it for multi-line content.
  • It treated any value starting with -- as a flag. YAML front matter (---) was therefore looked up as an unknown option and dropped.

As a result, content write couldn't produce a multi-line file, and vault gym cases like vault_create_daily_note couldn't pass.

Fix

  • One shared parser: shared/validation/parsing/cli_commands.py holds the tokenizer, option detection, command matching and argument binding.
    • Quoting follows POSIX shlex, the same tokenizer the repo already used everywhere.
    • Commands split on commas outside quotes, braces and brackets.
    • A token is a flag only if it is a declared flag or has the form --name.
    • -- keeps its catalog meaning: prompt sub declares it as a literal flag.
  • Escapes come from config: tool_call_formats.yaml gains command_escapes: {n, t, r}, decoded only inside double quotes. A format without the key gets plain POSIX quoting.
  • The three copied parsers are deleted, with no re-exports: Evaluator/config_validator.py (_expand_cli_wrapper), Tools/migrations/cli_schema_utils.py and tools/analyze_tool_coverage.py. Each copy had the same --- bug.
  • The multistep tests write directly again. The copy-then-replace workaround only existed because of this bug.

Vault gym path scoring

Evaluator/runner.py::_run_path_scoring compared the CLI command names in vault_gym.yaml against wrapper call names. Those paths could never match, even when the environment passed.

Observed calls are now matched at the level each path is written in:

  • Catalog commands (content write) are matched against the CLI commands a wrapper call expands to.
  • Tool names (contentManager_write) are matched against the tools those commands run.
  • Anything else, including the wrapper paths in the existing scoring tests, is matched against calls as written.

vault_create_daily_note now scores 1.0 on its preferred path.

Effect on existing output

  • All 60 catalog examples and 16 probe commands expand identically before and after.
  • Tool-coverage output is unchanged across all 379k checked-in dataset rows.
  • Migration output changes in 934 rows. 920 of them are content write calls whose --- content the old parser dropped; the rest gain decoded \n or inline --flag=value handling.

No hash-locked file is touched.

Test plan

  • New tests:
    • executor CLI tests: multi-line front matter, values starting with ---, embedded quotes, \n decoding, single quotes not decoded;
    • config validator CLI tests;
    • dataset-tool tests;
    • path scoring at each level;
    • vault gym end to end (environment passes, template-driven-daily-note, score 1.0).
  • Targeted suites: 1169 passed. All 14 failures are unrelated to this change and come from the sandbox: torch and pytest_asyncio aren't installed, and the launcher-lock test rejects the sandbox's ambient environment.
  • Ruff clean on changed files.

🤖 Generated with Claude Code

https://claude.ai/code/session_013etFNjLuo22CCoQmUGYqCT


Generated by Claude Code

claude added 4 commits October 2, 2026 17:55
The local executor that expands useTools CLI command strings rewrote every
whitespace character, including newlines inside quoted arguments, to a
space before splitting, never decoded the \n escape the CLI schema tells
models to use for multi-line content, and treated any token starting with
"--" as an option, so a "---" front-matter value was silently dropped. A
CLI write could not produce a multi-line file.

Replace the comma splitter plus shlex.split with one tokenizer that keeps
POSIX shell quoting (the shlex semantics the executor already used, and
the shape of the CLI-first training data: real newlines and \" inside
double quotes), leaves quoted text untouched, and splits commands on
top-level commas as before. Backslash escapes decoded inside double
quotes come from the tool-call format config (command_escapes on the
default format: n, t, r); formats without it keep plain POSIX quoting.
A token is an option only when it is a declared flag or option-shaped
(--name), so values that merely start with dashes stay arguments, while
the catalog's declared "--" flag (prompt sub) behaves as before.

Every catalog example and the existing command probes expand
identically; new tests cover multi-line, front-matter, embedded-quote and
escaped-newline writes and run vault_create_daily_note end to end
through the local executor with a direct multi-line write.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013etFNjLuo22CCoQmUGYqCT
The multistep loop tests built their multi-line daily notes by copying the
template and replacing single lines only because the local executor
collapsed whitespace in CLI commands. With quoted arguments now kept
exactly, both tests write the note in one content write command again,
restoring their original fixtures and step limits while keeping the
search/read/write sequence, environment assertions, and recovery checks.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013etFNjLuo22CCoQmUGYqCT
The environment executor, the Evaluator config validator, the CLI-schema
migration utilities and the tool-coverage script each carried their own
copy of the comma splitter, shlex tokenizing, command matching and
argument binding. The three copies still treated any token starting
with "--" as an option, so "---" front-matter values were dropped; over
the checked-in datasets that lost the content of 920 content write calls
in the migration inventory.

Move the tokenizer, option detection, command matching and argument
binding into shared/validation/parsing/cli_commands.py and delete the
copies. Each caller builds its catalog from its own schema source and
keeps its own policy for unknown commands: execution and validation
fall back to the original call, the dataset tools skip the command.
Escapes come from the configured wrapper's command_escapes, as in the
executor.

Over the checked-in datasets, tool-coverage results are unchanged for
every assistant message. Migration parsing differs only where the old
copy dropped dash-leading values or inline --flag=value options, or
left a configured \n escape undecoded. Catalog-example expansion in the
executor is unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013etFNjLuo22CCoQmUGYqCT
Path scoring compared configured tool names against the names of the
calls as made. The vault gym paths name CLI commands ("content write"),
but a CLI-wrapper model makes only wrapper calls, so none of those paths
could ever match.

Name the observed calls at three levels: as called, as the CLI commands
a configured wrapper call expands to (the executor's own expansion), and
as the catalog tools those commands run. A path naming catalog commands
is matched against the commands, one naming catalog tools against the
tools, and any other path, including count-only paths, against the calls
as before, so existing wrapper-name and direct-tool paths keep their
results.

The vault gym end-to-end test now also asserts that
vault_create_daily_note scores 1.0 on its preferred path.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013etFNjLuo22CCoQmUGYqCT
@ProfSynapse
ProfSynapse changed the base branch from claude/ci-and-test-repair to feat/submodule-cloud-api-v1 October 2, 2026 23:29
@ProfSynapse
ProfSynapse merged commit 49120db into feat/submodule-cloud-api-v1 Oct 2, 2026
9 checks passed
@ProfSynapse
ProfSynapse deleted the claude/cli-multiline-writes branch October 6, 2026 19:17
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants