Skip to content

Export replace_text_blocks, canonical_value and the JSON schema validator - #64

Merged
senamakel merged 2 commits into
mainfrom
dedupe-clones
Oct 7, 2026
Merged

senamakel merged 2 commits into
mainfrom
dedupe-clones

Conversation

@senamakel

@senamakel senamakel commented Oct 7, 2026 •

Copy link
Copy Markdown
Member

Summary

Makes three tinyinference-llm helpers public so tinyagents can stop carrying private copies of them. tinyanalyzer's clone detection found the tinyagents harness duplicating each one; its copies have identical bodies apart from visibility and comments.

  • prompt_tools::replace_text_blocks: rewrites a response's visible text after tool-call recovery while keeping non-text blocks (reasoning) in place.
  • cache::canonical_value: sorts JSON object keys recursively so cache keys are canonical. Sharing it also means the two crates can't drift apart and hash equal requests differently.
  • tool::validate_json_value (formerly the private validate_schema_value): the structural JSON-Schema subset tool arguments are held to.

Behavior change

One error message changes. When a schema has required or properties but no type, and the value is not an object, the message was … must be an object with declared fields. It is now … must be an object with the declared fields, got <kind>, which matches the tinyagents wording and names the actual kind. Nothing in this repo asserts the old text.

Everything else is additive: three new public functions, no signature changes.

Validation

  • cargo fmt --all -- --check passes
  • cargo clippy --all-targets --all-features -- -D warnings passes
  • cargo test --all-features passes

Tests

tests/tool_validation.rs covers the three public functions through the public API:

  • every validator keyword, including union types, the untyped-object message and unknown types;
  • key sorting at depth;
  • text replacement around non-text blocks, with empty inputs.

Related

The tinyagents PR that deletes its copies depends on this one and will follow.

Summary by CodeRabbit

  • New Features
    • Added access to utilities for canonicalizing JSON objects and replacing text blocks.
    • Exposed JSON value validation for callers, with clearer errors that identify the value’s type.

senamakel and others added 2 commits October 7, 2026 07:34
Tool definitions are now checked for structural problems such as missing
names, duplicate parameters, and malformed schemas, so invalid tools fail
early with a clear error instead of surfacing later during inference.
Validation results are cached alongside the parsed definitions.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
Extend the replace_text_blocks test to assert that JSON blocks
surrounding text are preserved in place, and add cases for empty
content and empty replacement text. The remaining changes are
rustfmt reflowing existing assertions.

Auto-committed-on: dragonfly
Co-authored-by: Medulla <medulla@tinyhumans.ai>
@tinysweeper

tinysweeper Bot commented Oct 7, 2026 •

Copy link
Copy Markdown

Tiny Sweeper review

Tiny Sweeper completed its review; deterministic results follow.

State: Ready for maintainer review
Priority: none
Reviewed head: 7d0396f41a6f
Updated: 1791348684 (Unix time)

Review snapshot

Change surface Files Review signal Count
Production 3 Active findings 0
Tests 1 Noted findings 0
Documentation 0 Resolved findings 0
Configuration 0 Pending checks/questions 0

Completeness: Complete
Test assessment: Test coverage is assessed from changed tests and lane evidence; execution is not claimed without trusted check data.

What changed

In `crates/tinyinference-llm/src/cache/mod.rs`, `canonical_value` is made public with rustdoc explaining that host crates should canonicalize cache keys the same way so equal requests do not hash apart (`#[must_use]` added). In `crates/tinyinference-llm/src/prompt_tools/mod.rs`, `replace_text_blocks` is made public with rustdoc explaining hosts with their own tool-call recovery can rewrite visible text identically. In `crates/tinyinference-llm/src/tool.rs`, `validate_schema_value` is renamed to the public `validate__value` with rustdoc describing the structural JSON Schema subset supported (type including unions, properties, required, additionalProperties: false, items, enum; unknown keywords ignored; empty/null schema imposes nothing) and its `Error::Validation` behavior; `ToolSchema::validate_call` now calls the renamed function. Two error messages for missing `type` were improved to include the actual value kind via `_value_kind`. In `crates/tinyinference-llm/tests/tool_validation.rs`, new integration tests cover the validator's error paths and permissive-schema behavior, canonical key sorting at every depth, and text-block replacement semantics.

Features

  • Modified — Export canonical_value: Host crates can canonicalize JSON for cache keys using the same recursive key-sorting logic as this crate, preventing drift between two copies that would make equal requests hash apart. (crates/tinyinference-llm/src/cache/mod.rs#fn fnv1a_hex(data: &[u8]) -> String {, crates/tinyinference-llm/src/cache/mod.rs)
  • Modified — Export replace_text_blocks: Hosts with their own tool-call recovery can rewrite visible text content exactly as recover_tool_calls does instead of carrying a duplicated copy. (crates/tinyinference-llm/src/prompt_tools/mod.rs#pub fn recover_tool_calls(mut response: ModelResponse, tools: &[ToolSchema]) ->, crates/tinyinference-llm/src/prompt_tools/mod.rs)
  • Modified — Export validate__value (renamed from validate_schema_value): External callers can validate values against the structural subset of JSON Schema held for tool arguments; two object-missing-type error messages now include the value's kind (e.g., 'got integer'), improving diagnostics. (crates/tinyinference-llm/src/tool.rs#fn validate_schema_value(schema: &Value, value: &Value, path: &str) -> crate::Re, crates/tinyinference-llm/src/tool.rs#impl ToolSchema {, crates/tinyinference-llm/src/tool.rs#pub struct ToolDelta {, crates/tinyinference-llm/src/tool.rs)

Tests

  • integration — canonical_value recursively sorts keys at every depth, asserted via exact serialized JSON output.: Exact-output assertion is a real behavioral check for the newly exported canonicalization helper. (crates/tinyinference-llm/tests/tool_validation.rs)
  • integration — replace_text_blocks keeps non-text blocks in place, substitutes one cleaned text at the first text block, appends the text when input is empty, and emits no text block when cleaned is empty.: Covers the newly exported helper including empty-text and empty-content edge cases. (crates/tinyinference-llm/tests/tool_validation.rs)

Findings

No active actionable findings.

Before merge

None.

How this fits together

flowchart LR
  n0["recover_tool_calls<br/>changed"]:::changed
  n1["ToolDelta<br/>changed"]:::changed
  n2["ToolSchema<br/>changed"]:::changed
  n3["...rguments_fail_even_with_permissive_schema<br/>changed"]:::changed
  n4["schema"]:::impacted
  n5["scrub_prompt_guided_item"]:::impacted
  n6["MessageDelta"]:::impacted
  n7["stream"]:::impacted
  n8["invalid"]:::impacted
  n3 -->|calls| n8
  n3 -->|tests| n8
  n4 -->|uses| n2
  n5 -->|calls| n0
  n5 -->|uses| n2
  n5 -->|calls| n6
  n5 -->|uses| n6
  n6 -->|uses| n1
  n7 -->|calls| n0
  n7 -->|calls| n5
  n7 -->|uses| n6
  classDef changed fill:#0d4429,stroke:#238636,color:#e6edf3
  classDef impacted fill:#161b22,stroke:#6e7681,color:#c9d1d9
  classDef flagged fill:#5a1e02,stroke:#d93f0b,color:#ffffff
  classDef blocking fill:#67060c,stroke:#f85149,color:#ffffff
Loading
Agent review details

critique

  • Conclusion: Success
  • Scope reviewed: all assigned evidence
  • Lane summary: The change makes the JSON validator and existing canonicalization/text-replacement helpers public, updates validation diagnostics, and adds focused coverage. I found no correctness issues introduced by this pull request; it looks safe to merge. _Code retrieval was unavailable (model: ladder embeddings returned 400 Bad Request: {"error":{"message":"unknown ladder vectors; known ladders are flash (also chat-v1, flash-v1), instant (also no-think, instant-v1), reasoning (also deepseek), max-reasoning (also max-reasoning-v1), deepseek-flash (also reasoning-v1, agentic-v1), deep (also luna), scribe, uncensored, vectors-oai3 (also embeddings-oai3-v1), vision (also vision-v1, multimodal-v1), image (also images-v1, image-v1), vi), so this review saw the diff alone._ _Memory was unavailable (model: cortex: v1/recall: error sending request for url (http://cortexdb:3141/v1/recall\)\), so this review ran without it._

security

  • Conclusion: Success
  • Scope reviewed: all assigned evidence
  • Positive: Exposing existing helpers and updating validation naming/diagnostics introduces no security or authorization issues per the security lane.
  • Lane summary: The change exposes existing helpers and updates tool validation naming and diagnostics without introducing security or authorization issues. It looks safe to merge. _Code retrieval was unavailable (model: ladder embeddings returned 400 Bad Request: {"error":{"message":"unknown ladder vectors; known ladders are flash (also chat-v1, flash-v1), instant (also no-think, instant-v1), reasoning (also deepseek), max-reasoning (also max-reasoning-v1), deepseek-flash (also reasoning-v1, agentic-v1), deep (also luna), scribe, uncensored, vectors-oai3 (also embeddings-oai3-v1), vision (also vision-v1, multimodal-v1), image (also images-v1, image-v1), vi), so this review saw the diff alone._ _Memory was unavailable (model: cortex: v1/recall: error sending request for url (http://cortexdb:3141/v1/recall\)\), so this review ran without it._

tests

  • Conclusion: Success
  • Scope reviewed: all assigned evidence
  • Positive: Every newly exposed function has at least one test with genuine behavioral assertions (error paths, exact canonical serialization, block replacement edge cases) that would fail on regression.
  • Positive: Tests are placed in the appropriate integration test file, exercising the new public API from outside the crate.
  • Lane summary: The change makes three previously private helpers public with rustdoc, improves two error messages to include the value's kind, and adds integration tests. The tests are genuine behavioral assertions (error paths, exact canonical serialization, block replacement with empty-text edge cases) that would fail on regression, and each newly exposed function has at least one. The change looks sound and safe to merge. _Code retrieval was unavailable (model: ladder embeddings returned 400 Bad Request: {"error":{"message":"unknown ladder vectors; known ladders are flash (also chat-v1, flash-v1), instant (also no-think, instant-v1), reasoning (also deepseek), max-reasoning (also max-reasoning-v1), deepseek-flash (also reasoning-v1, agentic-v1), deep (also luna), scribe, uncensored, vectors-oai3 (also embeddings-oai3-v1), vision (also vision-v1, multimodal-v1), image (also images-v1, image-v1), vi), so this review saw the diff alone._ _Memory was unavailable (model: cortex: v1/recall: error sending request for url (http://cortexdb:3141/v1/recall\)\), so this review ran without it._

commits

  • Conclusion: Neutral
  • Scope reviewed: all assigned evidence
  • Positive: Nothing sensitive found in what this pull request commits.
  • Lane summary: Nothing sensitive found in what this pull request commits.

description

  • Conclusion: Success
  • Scope reviewed: all assigned evidence
  • Positive: The PR description matches the diff: three private helpers made public as described, the documented error-message change included, and tests added in the right location; nothing else changes.
  • Lane summary: The PR makes three private helpers public exactly as described, includes the one documented error-message change, and adds tests in the right location covering all three functions. The description matches the diff; nothing else changes. Safe to merge. _Code retrieval was unavailable (model: ladder embeddings returned 400 Bad Request: {"error":{"message":"unknown ladder vectors; known ladders are flash (also chat-v1, flash-v1), instant (also no-think, instant-v1), reasoning (also deepseek), max-reasoning (also max-reasoning-v1), deepseek-flash (also reasoning-v1, agentic-v1), deep (also luna), scribe, uncensored, vectors-oai3 (also embeddings-oai3-v1), vision (also vision-v1, multimodal-v1), image (also images-v1, image-v1), vi), so this review saw the diff alone._ _Memory was unavailable (model: cortex: v1/recall: error sending request for url (http://cortexdb:3141/v1/recall\)\), so this review ran without it._

e2e

  • Conclusion: Neutral
  • Scope reviewed: all assigned evidence
  • Lane summary: No end-to-end harness in this repository: no e2e test files and no e2e workflow.
Evidence and run details
  • Models: gpt-5.6-luna, glm-5.3-flash
  • Spend: $0.002926
  • Tokens: 70112 input · 3884 output · 8272 cached · 0 embedding
Head State Pass summary
7d0396f41a6f ready for maintainer review 0 active finding(s), 0 resolved finding(s) (at 1791348684)

tinysweeper 0.1.0

@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Oct 7, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-10-07T04:53:11.732316Z 7d0396f PR opened
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@coderabbitai

coderabbitai Bot commented Oct 7, 2026 •

Copy link
Copy Markdown

Review in Change Stack →

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration
  • Configuration used: Organization UI
  • Review profile: CHILL
  • Plan: Advanced
  • Run ID: 0debb268-78b9-4bbd-b8fc-778a63cc05e1
📥 Commits

Reviewing files that changed from the base of the PR and between c71e6b5 and 7d0396f.

📒 Files selected for processing (4)
  • crates/tinyinference-llm/src/cache/mod.rs
  • crates/tinyinference-llm/src/prompt_tools/mod.rs
  • crates/tinyinference-llm/src/tool.rs
  • crates/tinyinference-llm/tests/tool_validation.rs

Included review availability: This review used your included allowance. Your plan provides up to 1 included review per hour; 0 remain after this review.


📝 Walkthrough

Walkthrough

The change makes JSON canonicalization, text-block replacement, and JSON schema validation functions public. It also renames the validator, updates its recursive calls, adds JSON value kinds to some validation errors, and adds tests for these behaviors.

Changes

JSON canonicalization

Layer / File(s) Summary
Expose JSON canonicalization
crates/tinyinference-llm/src/cache/mod.rs, crates/tinyinference-llm/tests/tool_validation.rs
canonical_value is now public. Tests check that object keys are sorted at the top level and in nested objects within arrays.

Text block replacement

Layer / File(s) Summary
Expose text block replacement
crates/tinyinference-llm/src/prompt_tools/mod.rs, crates/tinyinference-llm/tests/tool_validation.rs
replace_text_blocks is now public. Tests cover preserving other blocks, replacing or inserting text, and removing text when the replacement is empty.

JSON schema validation

Layer / File(s) Summary
Expose and test JSON validation
crates/tinyinference-llm/src/tool.rs, crates/tinyinference-llm/tests/tool_validation.rs
The public validator is named validate_json_value; its call sites and recursive calls use the new name. Some errors include the received JSON value kind. Tests cover required fields, types, properties, arrays, enums, union types, and permissive handling of empty or unsupported schemas.

Priority: ⬇️ Low

Estimated code review effort: 2 (Simple) | ~10 minutes

Change: Feature

Merge Risk: ⚪ Minimal · up to 7d039

This change exposes three existing helpers for reuse and slightly improves one validation error message. No actionable merge risk remains.

Security Architecture Review

Security architecture risk: ⚪ Minimal · up to 7d039

The new APIs expose existing data-processing behavior without adding tool-execution authority or weakening the existing tool-call checks. No material security risk was found in the reviewed change. Downstream integrations remain responsible for authorization and correct use of the structural validator.

Retained concerns
No architecture-level concerns identified.

Security review details

Security Blast Radius

  • inferred — The newly exposed entrypoints operate on JSON values and content supplied by an in-process caller. Their implementations add no network access, credential access, persistent writes or tool-execution authority. Any downstream expansion of authority depends on host integration outside the reviewed repository.

Trust Boundaries and Controls

  • observed — ToolSchema::validate_call still rejects malformed provider arguments and mismatched tool names before structural validation. Direct use of validate_json_value omits those call-specific checks by design, but does not invoke a tool or confer authorization.

Resilience and Maintainability Implications

  • observed — The transformation helpers consume owned inputs and construct returned values; the validator borrows its inputs without modifying them. These exported functions do not introduce shared-state transitions, reservations or cleanup obligations, so interruption does not leave an externally persisted intermediate state within these functions.
🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title clearly identifies the three helpers made public, which is the pull request’s main change.
Docstring Coverage ✅ Passed Docstring coverage is 100.00% which is sufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 8 functions across 4 files.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
✨ Finishing Touches
📝 Generate docstrings
  • Commit to this branch
  • Create a new PR
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

A rabbit checks the keys in a row
Then helps fresh words replace the old
It tests each shape, each branch, each line
The public paths now stand in view
And hops away, its work complete

Comment @coderabbitai help to get the list of available commands.

@tinysweeper tinysweeper Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

tinysweeper found nothing blocking. Approving.

             $0.0029 · 70,112 in / 3,884 out · 8,272 cached (12%) · gpt-5.6-luna, glm-5.3-flash
critique:    $0.0015 · 27,487 in / 1,644 out · 4,284 cached (16%) · gpt-5.6-luna
security:    $0.0012 · 25,644 in / 638 out   · 3,796 cached (15%) · gpt-5.6-luna
tests:       $0.0001 · 6,153 in  / 199 out   · 64 cached (1%)     · glm-5.3-flash
description: $0.0000 · 5,886 in  / 54 out    · 64 cached (1%)     · glm-5.3-flash

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 7d0396f41a

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

///
/// Returns [`crate::Error::Validation`] naming the first failing instance
/// path.
pub fn validate_json_value(schema: &Value, value: &Value, path: &str) -> crate::Result<()> {

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Expose shared helpers from a supported public crate

This public export, together with the new cache::canonical_value and prompt_tools::replace_text_blocks exports, is explicitly intended for consumption by TinyAgents, but repository policy designates only tinyinference-core and tinyinference-local as public crates and assigns provider-neutral inference, cache, message, and tool-call APIs to core. Depending on these helpers through tinyinference-llm creates a new unsupported cross-repository API boundary; place the normalized helpers/types behind the core public surface instead (or update the repository architecture policy as part of the change).

AGENTS.md reference: AGENTS.md:L12-L18

Useful? React with 👍 / 👎.

@senamakel
senamakel merged commit a32eec8 into main Oct 7, 2026
14 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant