Skip to content

Drive dataset and structure validators from tool-call format config - #169

Merged
ProfSynapse merged 8 commits into
feat/submodule-cloud-api-v1from
claude/validator-config-first
Oct 2, 2026
Merged

ProfSynapse merged 8 commits into
feat/submodule-cloud-api-v1from
claude/validator-config-first

Conversation

@ProfSynapse

Copy link
Copy Markdown
Owner

Summary

Stacked on #168.

This removes useTools-specific hardcoding from the validators, as CLAUDE.md requires: tool-call formats come from config.

Dataset validator (shared/validation/dataset_validator.py)

  • Wrapper checks come from config. The wrapper name, required argument fields and per-field constraints come from SynthChat/config/tool_call_formats.yaml, read through the existing loader and configured_formats.build_wrapper_specs. Wrapper-less formats now validate. --tool-call-formats PATH lets a host validate against its own config.
  • Schema validation now actually runs. SCHEMAS_FILE pointed at tools/tool_schemas.json, but the file is at Tools/, so schema validation was always skipped on Linux. It now resolves from the engine root, and --tool-schemas PATH lets a host override it.
  • ID checks come from config. They use a new format key, prompt_bound_fields ({pattern, in_tag} sources per field). This replaces the hardcoded sessionId/workspaceId regexes.
  • Removed, with no config basis:
    • the memoryManager_/agentManager_ ID warnings;
    • the get_tools special case;
    • the argument allow-list (the catalog's parameters already declares every accepted argument);
    • the old flat-context check;
    • the dead validate_tool_call_structure.
  • The validate_syngen.py star-import shim is deleted. Every caller now runs python3 -m shared.validation.dataset_validator, and a test keeps the shim from coming back.

Wrapper command field

  • A new format key, command_field, names the CLI command argument. The default format sets it to tool.
  • It replaces the literal tool everywhere it was hardcoded: the sanitizers in configured_formats, response_parser, the CLI expansion in tool_executor, and Keep quoted CLI arguments exact, share one CLI parser, and fix vault gym path scoring #168's Evaluator/config_validator.py.
  • The duplicate, unread wrapper.cli_field is removed from Tools/tool_schemas.json.

Structure validator (structure_validator.py)

  • The tools rule's error template is now applied, with {tool_name} and {details}. Every active tool rubric sets it, but it was read and then ignored.
  • Schema-level _item_schema is now enforced. This also fixes a crash on scalar item schemas, and arrays of arrays that were silently skipped.
  • Subtool keys come from config. They are declared through a new required _subtool_keys: {group, tool, params} instead of hardcoded agent/tool/params. All configs and docs that use _subtools are updated.

Other changes

  • args_match.py: removed call_fields. It has been unused since it was replaced by mappings (98098e4), and no config sets it.
  • SynthChat workspace prompt: "Required wrapper fields" now renders from argument_required. It used to render an empty list because the default format never sets context_fields.

What schema validation now finds in the checked-in datasets

These are reported, not suppressed. No datasets were changed.

  • Legacy argument shape in about 49 files. They use the old useTools {context, calls} arguments, which gives about 178k "missing required parameter" errors. Examples: nonthinking_tools_sft_01.02.26.jsonl, sft_train_67pct_12.25.jsonl, and the older manager tool sets.
  • Field-constraint violations in 10 newer CLI-format files: 56 empty required fields and 18 values outside their allowed list. Includes nonthinking_tools_sft_04.22.26.jsonl and promptManager/tools_v2.6–v2.8.
  • 44 IDs that don't match their system prompt, across 16 files.

No hash-locked file is touched.

Test plan

  • New tests:
    • test_dataset_validator.py (20)
    • test_configured_formats.py (6)
    • structure validator: error template, item schemas, subtool keys
    • verifier and workspace tests
    • fixtures use a format whose names have nothing to do with useTools
  • Keep quoted CLI arguments exact, share one CLI parser, and fix vault gym path scoring #168's suites (tool-executor CLI, vault gym, multistep, evaluator scoring, config-validator CLI, dataset tools): 202 passed.
  • Broader sweep: 3800 passed. It has the same 39 failures and 8 errors as the untouched base, all from missing packages in the sandbox.
  • ruff check . and sync_skill_trees.py --check are clean.

🤖 Generated with Claude Code

https://claude.ai/code/session_013etFNjLuo22CCoQmUGYqCT


Generated by Claude Code

claude added 8 commits October 2, 2026 18:29
dataset_validator.validate_context required a fixed field list
(sessionId, workspaceId, memory, goal, tool) of every tool call and checked
sessionId/workspaceId against hardcoded generator ID regexes. Every direct
call of a wrapper-less format therefore failed, and the toy useTools wrapper
was treated as the canonical format.

validate_wrapper_arguments now resolves the wrapper a call uses with the
shared match_configured_wrapper helper, over specs built from the same
tool-call format registry SynthChat and the Evaluator parsers use
(SynthChat/config/tool_call_formats.yaml via load_tool_call_formats). A
matched wrapper must carry its configured argument_required fields, and each
field must satisfy its configured property schema (type, minLength, enum).
A call that matches no configured wrapper is a direct call and needs no
wrapper fields. The hardcoded ID regexes are gone; ID formats belong in the
format's property schema if a host wants them.

configured_formats gains build_wrapper_specs(formats), and
match_configured_wrapper an optional specs argument, so callers and tests can
validate against another registry; the validator CLI takes
--tool-call-formats PATH for a host registry.

SynthChat's agentic loop validator only existed to drop the removed
"does not match generator format" errors; it now returns the structural
result directly.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013etFNjLuo22CCoQmUGYqCT
StructureValidator._validate_tools read the rule's `error` template and
_validate_against_schema read a schema-level `_item_schema`, then ignored
both (the dead reads were dropped by the pyflakes cleanup, leaving the
behaviour unchanged). Every checked-in tool rubric and the SFT fitness config
set `error: "useTools validation failed: {details}"`, and the generator that
emits tool rules (SynthChat/scripts/embed_tool_schemas.py) writes
`error: "Tool '{tool_name}': {details}"`, so the intended contract is clear.

- Each failure of a tools rule (schema violations, unknown tool, invalid
  argument JSON, subtool params) is now formatted with the rule's `error`
  template, with {tool_name} and {details}. Without `error` the default
  "Tool '{tool_name}': {details}" reproduces the previous messages.
- `_item_schema` is handled in one place: any schema carrying it describes an
  array whose items must match the item schema, which may be a nested
  schema or a type name. This keeps field-level `_item_schema` working and
  adds scalar item types and arrays of arrays, which previously crashed or
  were skipped.

Docs: rubric README gains a Tool Manifest Validation section; the shared
validation README, rubric example and rubric-authoring skill describe the
template and _item_schema (the skill example now declares the fields its
_required list names, since _required only applies to declared fields).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013etFNjLuo22CCoQmUGYqCT
score_legacy_fields read comparison.call_fields and never used it (the
dead read was dropped by the pyflakes cleanup; the docstring still listed
it). History shows no live intent: call_fields was introduced in 165f344
with the legacy GRPO scorer, which even then compared hardcoded
agent/tool/params instead, and the only config that set it
(Trainers/rtx3090_grpo/configs/rewards/args_match.yaml) replaced it with
explicit `mappings` in 98098e4. No checked-in config sets it now; the only
remaining reference was an inert `call_fields: []` in a characterization
test fixture, removed here.

The docstring now points at the `mappings` scheme, which is how a rubric
selects the call fields to compare; new tests drive that through a verifier
spec fixture.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013etFNjLuo22CCoQmUGYqCT
…om config

dataset_validator looked for tools/tool_schemas.json; the catalog lives at
Tools/tool_schemas.json, so schema validation never ran on a case-sensitive
filesystem. SCHEMAS_FILE now resolves from the engine root
(shared.utilities.paths.get_engine_root), independent of the CWD.

With schema validation live, the remaining hardcoding is replaced by config:

- validate_tool_against_schema checks the catalog entry's required_params and
  warns about arguments not declared in its parameters. The sessionId/
  workspaceId/context/workspaceContext allow-list, the 'context' skip and
  nested context_schema check (the old flat-context toy format) and the
  get_tools special case (no catalog entry) are removed.
- validate_ids_match_system_prompt is replaced by prompt_bound_fields, a new
  key on a tool-call format: field -> list of {pattern, in_tag} sources that
  state the field's allowed values in the system prompt. The default format
  declares sessionId and workspaceId with the patterns the validator used to
  hardcode. The memoryManager_/agentManager_ target-ID warnings had no config
  basis (the configured format has no such direct tools) and are removed,
  along with the dead toy-format validate_tool_call_structure, MIN_TOOL_CALLS
  and the hardcoded system-prompt regexes.

Wrapper format and tool catalog checks share one message for a missing
required field and ExampleReport reports each finding once. The validator
takes a ValidatorConfig (wrapper specs + tool schemas; engine defaults when
omitted) and the CLI gains --tool-schemas PATH.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013etFNjLuo22CCoQmUGYqCT
Two helpers still special-cased toy-format field names:

- configured_formats.sanitize_* and extract_lenient_wrapper_arguments treated
  an argument literally named `tool` as the free-form command string. A
  tool-call format now declares it as `command_field` (the default format sets
  `command_field: tool`; build_wrapper_specs rejects a command_field that is
  not a declared argument field). Formats without one get no command
  recovery. The CLI expansion paths read the same spec key instead of
  `tool`: the environment executor's expand_cli_wrapper_commands (which feeds
  the shared cli_commands parser together with the spec's command_escapes),
  the Evaluator config validator's _expand_cli_wrapper, and the response
  parser's context extraction.
  sanitize_wrapper_string_fields and extract_lenient_wrapper_arguments take an
  optional specs argument like match_configured_wrapper;
  sanitize_extracted_tool_value is renamed sanitize_command_value.
- StructureValidator._validate_subtool_params read each array item's
  `agent`, `tool` and `params`. An array schema with `_subtools` must now also
  declare `_subtool_keys: {group, tool, params}` naming those item fields;
  a `_subtools` without it is a ValueError. Every config and doc example that
  sets `_subtools` (the SFT fitness config, the archived rubrics, the
  evolutionary fine-tuning and subtool docs) declares
  `{group: agent, tool: tool, params: params}`.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013etFNjLuo22CCoQmUGYqCT
.skills/synthetic-data-generation/scripts/validate_syngen.py only star-
imported shared.validation.dataset_validator and called main(), a
re-export shim CLAUDE.md forbids. It is deleted and every caller now runs
`python3 -m shared.validation.dataset_validator` from the engine root:
skill references (synced to .agents/.claude), docs, Tools/run_synth_chat.*,
Tools/README.md, Datasets helper scripts and specs (including the stale
tools/validate_syngen.py and misspelt .skills/synethetic-data-generation
paths), Evaluator/requirements.txt notes and a docs pseudo-code example.
tests/scripts/test_skill_script_references.py bans the old name and checks
the shim stays gone.

The synthetic-data-generation skill also documents the config-first checks
from the previous commits: --tool-schemas, prompt_bound_fields, the tool
schema catalog checks and the command_field / prompt_bound_fields format keys.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013etFNjLuo22CCoQmUGYqCT
SynthChat/workspace/sections.py built the available_tools "Required wrapper
fields" line from `context_fields.required`, which the default format does not
set, so the configured default rendered "Required wrapper fields: .".

The line now lists the format's required argument fields, resolved by a new
format_resolver.required_argument_fields (argument_required, falling back to
argument_fields.required); build_wrapper_specs uses the same helper, so the
prompt and the validators agree on one definition. The placeholder is renamed
{required_fields_csv}, and a template naming it is omitted for a format that
requires no fields. The example formats' `context_fields` blocks, which
nothing reads any more, are removed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013etFNjLuo22CCoQmUGYqCT
Tools/tool_schemas.json declared the wrapper's command field as
`wrapper.cli_field: "tool"`, but nothing reads the catalog's wrapper block.
The tool-call format's `command_field` is now the single source.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013etFNjLuo22CCoQmUGYqCT
@ProfSynapse
ProfSynapse changed the base branch from claude/cli-multiline-writes to feat/submodule-cloud-api-v1 October 2, 2026 23:29
@ProfSynapse
ProfSynapse merged commit 18af454 into feat/submodule-cloud-api-v1 Oct 2, 2026
9 checks passed
@ProfSynapse
ProfSynapse deleted the claude/validator-config-first branch October 6, 2026 19:17
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants