Repository navigation
Keep quoted CLI arguments exact, share one CLI parser, and fix vault gym path scoring - #168
Merged
ProfSynapse merged 4 commits intoOct 2, 2026
Conversation
The local executor that expands useTools CLI command strings rewrote every whitespace character, including newlines inside quoted arguments, to a space before splitting, never decoded the \n escape the CLI schema tells models to use for multi-line content, and treated any token starting with "--" as an option, so a "---" front-matter value was silently dropped. A CLI write could not produce a multi-line file. Replace the comma splitter plus shlex.split with one tokenizer that keeps POSIX shell quoting (the shlex semantics the executor already used, and the shape of the CLI-first training data: real newlines and \" inside double quotes), leaves quoted text untouched, and splits commands on top-level commas as before. Backslash escapes decoded inside double quotes come from the tool-call format config (command_escapes on the default format: n, t, r); formats without it keep plain POSIX quoting. A token is an option only when it is a declared flag or option-shaped (--name), so values that merely start with dashes stay arguments, while the catalog's declared "--" flag (prompt sub) behaves as before. Every catalog example and the existing command probes expand identically; new tests cover multi-line, front-matter, embedded-quote and escaped-newline writes and run vault_create_daily_note end to end through the local executor with a direct multi-line write. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013etFNjLuo22CCoQmUGYqCT
The multistep loop tests built their multi-line daily notes by copying the template and replacing single lines only because the local executor collapsed whitespace in CLI commands. With quoted arguments now kept exactly, both tests write the note in one content write command again, restoring their original fixtures and step limits while keeping the search/read/write sequence, environment assertions, and recovery checks. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013etFNjLuo22CCoQmUGYqCT
The environment executor, the Evaluator config validator, the CLI-schema migration utilities and the tool-coverage script each carried their own copy of the comma splitter, shlex tokenizing, command matching and argument binding. The three copies still treated any token starting with "--" as an option, so "---" front-matter values were dropped; over the checked-in datasets that lost the content of 920 content write calls in the migration inventory. Move the tokenizer, option detection, command matching and argument binding into shared/validation/parsing/cli_commands.py and delete the copies. Each caller builds its catalog from its own schema source and keeps its own policy for unknown commands: execution and validation fall back to the original call, the dataset tools skip the command. Escapes come from the configured wrapper's command_escapes, as in the executor. Over the checked-in datasets, tool-coverage results are unchanged for every assistant message. Migration parsing differs only where the old copy dropped dash-leading values or inline --flag=value options, or left a configured \n escape undecoded. Catalog-example expansion in the executor is unchanged. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013etFNjLuo22CCoQmUGYqCT
Path scoring compared configured tool names against the names of the
calls as made. The vault gym paths name CLI commands ("content write"),
but a CLI-wrapper model makes only wrapper calls, so none of those paths
could ever match.
Name the observed calls at three levels: as called, as the CLI commands
a configured wrapper call expands to (the executor's own expansion), and
as the catalog tools those commands run. A path naming catalog commands
is matched against the commands, one naming catalog tools against the
tools, and any other path, including count-only paths, against the calls
as before, so existing wrapper-name and direct-tool paths keep their
results.
The vault gym end-to-end test now also asserts that
vault_create_daily_note scores 1.0 on its preferred path.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013etFNjLuo22CCoQmUGYqCT
4 tasks done
ProfSynapse
changed the base branch from
claude/ci-and-test-repair
to
feat/submodule-cloud-api-v1
October 2, 2026 23:29
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Stacked on #166.
Root cause: multi-line CLI writes were broken
shared/environments/tool_executor.pyhad three problems:\n, even though the CLI schema tells models to use it for multi-line content.--as a flag. YAML front matter (---) was therefore looked up as an unknown option and dropped.As a result,
content writecouldn't produce a multi-line file, and vault gym cases likevault_create_daily_notecouldn't pass.Fix
shared/validation/parsing/cli_commands.pyholds the tokenizer, option detection, command matching and argument binding.shlex, the same tokenizer the repo already used everywhere.--name.--keeps its catalog meaning:prompt subdeclares it as a literal flag.tool_call_formats.yamlgainscommand_escapes: {n, t, r}, decoded only inside double quotes. A format without the key gets plain POSIX quoting.Evaluator/config_validator.py(_expand_cli_wrapper),Tools/migrations/cli_schema_utils.pyandtools/analyze_tool_coverage.py. Each copy had the same---bug.Vault gym path scoring
Evaluator/runner.py::_run_path_scoringcompared the CLI command names invault_gym.yamlagainst wrapper call names. Those paths could never match, even when the environment passed.Observed calls are now matched at the level each path is written in:
content write) are matched against the CLI commands a wrapper call expands to.contentManager_write) are matched against the tools those commands run.vault_create_daily_notenow scores 1.0 on its preferred path.Effect on existing output
content writecalls whose---content the old parser dropped; the rest gain decoded\nor inline--flag=valuehandling.No hash-locked file is touched.
Test plan
---, embedded quotes,\ndecoding, single quotes not decoded;template-driven-daily-note, score 1.0).pytest_asyncioaren't installed, and the launcher-lock test rejects the sandbox's ambient environment.🤖 Generated with Claude Code
https://claude.ai/code/session_013etFNjLuo22CCoQmUGYqCT
Generated by Claude Code