Repository navigation
Drive dataset and structure validators from tool-call format config - #169
Merged
ProfSynapse merged 8 commits intoOct 2, 2026
Merged
Conversation
dataset_validator.validate_context required a fixed field list (sessionId, workspaceId, memory, goal, tool) of every tool call and checked sessionId/workspaceId against hardcoded generator ID regexes. Every direct call of a wrapper-less format therefore failed, and the toy useTools wrapper was treated as the canonical format. validate_wrapper_arguments now resolves the wrapper a call uses with the shared match_configured_wrapper helper, over specs built from the same tool-call format registry SynthChat and the Evaluator parsers use (SynthChat/config/tool_call_formats.yaml via load_tool_call_formats). A matched wrapper must carry its configured argument_required fields, and each field must satisfy its configured property schema (type, minLength, enum). A call that matches no configured wrapper is a direct call and needs no wrapper fields. The hardcoded ID regexes are gone; ID formats belong in the format's property schema if a host wants them. configured_formats gains build_wrapper_specs(formats), and match_configured_wrapper an optional specs argument, so callers and tests can validate against another registry; the validator CLI takes --tool-call-formats PATH for a host registry. SynthChat's agentic loop validator only existed to drop the removed "does not match generator format" errors; it now returns the structural result directly. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013etFNjLuo22CCoQmUGYqCT
StructureValidator._validate_tools read the rule's `error` template and
_validate_against_schema read a schema-level `_item_schema`, then ignored
both (the dead reads were dropped by the pyflakes cleanup, leaving the
behaviour unchanged). Every checked-in tool rubric and the SFT fitness config
set `error: "useTools validation failed: {details}"`, and the generator that
emits tool rules (SynthChat/scripts/embed_tool_schemas.py) writes
`error: "Tool '{tool_name}': {details}"`, so the intended contract is clear.
- Each failure of a tools rule (schema violations, unknown tool, invalid
argument JSON, subtool params) is now formatted with the rule's `error`
template, with {tool_name} and {details}. Without `error` the default
"Tool '{tool_name}': {details}" reproduces the previous messages.
- `_item_schema` is handled in one place: any schema carrying it describes an
array whose items must match the item schema, which may be a nested
schema or a type name. This keeps field-level `_item_schema` working and
adds scalar item types and arrays of arrays, which previously crashed or
were skipped.
Docs: rubric README gains a Tool Manifest Validation section; the shared
validation README, rubric example and rubric-authoring skill describe the
template and _item_schema (the skill example now declares the fields its
_required list names, since _required only applies to declared fields).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013etFNjLuo22CCoQmUGYqCT
score_legacy_fields read comparison.call_fields and never used it (the dead read was dropped by the pyflakes cleanup; the docstring still listed it). History shows no live intent: call_fields was introduced in 165f344 with the legacy GRPO scorer, which even then compared hardcoded agent/tool/params instead, and the only config that set it (Trainers/rtx3090_grpo/configs/rewards/args_match.yaml) replaced it with explicit `mappings` in 98098e4. No checked-in config sets it now; the only remaining reference was an inert `call_fields: []` in a characterization test fixture, removed here. The docstring now points at the `mappings` scheme, which is how a rubric selects the call fields to compare; new tests drive that through a verifier spec fixture. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013etFNjLuo22CCoQmUGYqCT
…om config
dataset_validator looked for tools/tool_schemas.json; the catalog lives at
Tools/tool_schemas.json, so schema validation never ran on a case-sensitive
filesystem. SCHEMAS_FILE now resolves from the engine root
(shared.utilities.paths.get_engine_root), independent of the CWD.
With schema validation live, the remaining hardcoding is replaced by config:
- validate_tool_against_schema checks the catalog entry's required_params and
warns about arguments not declared in its parameters. The sessionId/
workspaceId/context/workspaceContext allow-list, the 'context' skip and
nested context_schema check (the old flat-context toy format) and the
get_tools special case (no catalog entry) are removed.
- validate_ids_match_system_prompt is replaced by prompt_bound_fields, a new
key on a tool-call format: field -> list of {pattern, in_tag} sources that
state the field's allowed values in the system prompt. The default format
declares sessionId and workspaceId with the patterns the validator used to
hardcode. The memoryManager_/agentManager_ target-ID warnings had no config
basis (the configured format has no such direct tools) and are removed,
along with the dead toy-format validate_tool_call_structure, MIN_TOOL_CALLS
and the hardcoded system-prompt regexes.
Wrapper format and tool catalog checks share one message for a missing
required field and ExampleReport reports each finding once. The validator
takes a ValidatorConfig (wrapper specs + tool schemas; engine defaults when
omitted) and the CLI gains --tool-schemas PATH.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013etFNjLuo22CCoQmUGYqCT
Two helpers still special-cased toy-format field names:
- configured_formats.sanitize_* and extract_lenient_wrapper_arguments treated
an argument literally named `tool` as the free-form command string. A
tool-call format now declares it as `command_field` (the default format sets
`command_field: tool`; build_wrapper_specs rejects a command_field that is
not a declared argument field). Formats without one get no command
recovery. The CLI expansion paths read the same spec key instead of
`tool`: the environment executor's expand_cli_wrapper_commands (which feeds
the shared cli_commands parser together with the spec's command_escapes),
the Evaluator config validator's _expand_cli_wrapper, and the response
parser's context extraction.
sanitize_wrapper_string_fields and extract_lenient_wrapper_arguments take an
optional specs argument like match_configured_wrapper;
sanitize_extracted_tool_value is renamed sanitize_command_value.
- StructureValidator._validate_subtool_params read each array item's
`agent`, `tool` and `params`. An array schema with `_subtools` must now also
declare `_subtool_keys: {group, tool, params}` naming those item fields;
a `_subtools` without it is a ValueError. Every config and doc example that
sets `_subtools` (the SFT fitness config, the archived rubrics, the
evolutionary fine-tuning and subtool docs) declares
`{group: agent, tool: tool, params: params}`.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013etFNjLuo22CCoQmUGYqCT
.skills/synthetic-data-generation/scripts/validate_syngen.py only star- imported shared.validation.dataset_validator and called main(), a re-export shim CLAUDE.md forbids. It is deleted and every caller now runs `python3 -m shared.validation.dataset_validator` from the engine root: skill references (synced to .agents/.claude), docs, Tools/run_synth_chat.*, Tools/README.md, Datasets helper scripts and specs (including the stale tools/validate_syngen.py and misspelt .skills/synethetic-data-generation paths), Evaluator/requirements.txt notes and a docs pseudo-code example. tests/scripts/test_skill_script_references.py bans the old name and checks the shim stays gone. The synthetic-data-generation skill also documents the config-first checks from the previous commits: --tool-schemas, prompt_bound_fields, the tool schema catalog checks and the command_field / prompt_bound_fields format keys. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013etFNjLuo22CCoQmUGYqCT
SynthChat/workspace/sections.py built the available_tools "Required wrapper
fields" line from `context_fields.required`, which the default format does not
set, so the configured default rendered "Required wrapper fields: .".
The line now lists the format's required argument fields, resolved by a new
format_resolver.required_argument_fields (argument_required, falling back to
argument_fields.required); build_wrapper_specs uses the same helper, so the
prompt and the validators agree on one definition. The placeholder is renamed
{required_fields_csv}, and a template naming it is omitted for a format that
requires no fields. The example formats' `context_fields` blocks, which
nothing reads any more, are removed.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013etFNjLuo22CCoQmUGYqCT
Tools/tool_schemas.json declared the wrapper's command field as `wrapper.cli_field: "tool"`, but nothing reads the catalog's wrapper block. The tool-call format's `command_field` is now the single source. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013etFNjLuo22CCoQmUGYqCT
ProfSynapse
changed the base branch from
claude/cli-multiline-writes
to
feat/submodule-cloud-api-v1
October 2, 2026 23:29
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Stacked on #168.
This removes
useTools-specific hardcoding from the validators, as CLAUDE.md requires: tool-call formats come from config.Dataset validator (
shared/validation/dataset_validator.py)SynthChat/config/tool_call_formats.yaml, read through the existing loader andconfigured_formats.build_wrapper_specs. Wrapper-less formats now validate.--tool-call-formats PATHlets a host validate against its own config.SCHEMAS_FILEpointed attools/tool_schemas.json, but the file is atTools/, so schema validation was always skipped on Linux. It now resolves from the engine root, and--tool-schemas PATHlets a host override it.prompt_bound_fields({pattern, in_tag}sources per field). This replaces the hardcodedsessionId/workspaceIdregexes.memoryManager_/agentManager_ID warnings;get_toolsspecial case;parametersalready declares every accepted argument);validate_tool_call_structure.validate_syngen.pystar-import shim is deleted. Every caller now runspython3 -m shared.validation.dataset_validator, and a test keeps the shim from coming back.Wrapper command field
command_field, names the CLI command argument. The default format sets it totool.tooleverywhere it was hardcoded: the sanitizers inconfigured_formats,response_parser, the CLI expansion intool_executor, and Keep quoted CLI arguments exact, share one CLI parser, and fix vault gym path scoring #168'sEvaluator/config_validator.py.wrapper.cli_fieldis removed fromTools/tool_schemas.json.Structure validator (
structure_validator.py)errortemplate is now applied, with{tool_name}and{details}. Every active tool rubric sets it, but it was read and then ignored._item_schemais now enforced. This also fixes a crash on scalar item schemas, and arrays of arrays that were silently skipped._subtool_keys: {group, tool, params}instead of hardcodedagent/tool/params. All configs and docs that use_subtoolsare updated.Other changes
args_match.py: removedcall_fields. It has been unused since it was replaced bymappings(98098e4), and no config sets it.argument_required. It used to render an empty list because the default format never setscontext_fields.What schema validation now finds in the checked-in datasets
These are reported, not suppressed. No datasets were changed.
useTools{context, calls}arguments, which gives about 178k "missing required parameter" errors. Examples:nonthinking_tools_sft_01.02.26.jsonl,sft_train_67pct_12.25.jsonl, and the older manager tool sets.nonthinking_tools_sft_04.22.26.jsonlandpromptManager/tools_v2.6–v2.8.No hash-locked file is touched.
Test plan
test_dataset_validator.py(20)test_configured_formats.py(6)useToolsruff check .andsync_skill_trees.py --checkare clean.🤖 Generated with Claude Code
https://claude.ai/code/session_013etFNjLuo22CCoQmUGYqCT
Generated by Claude Code