Skip to content

Fixes #8511 : Add a Structured extract transform to read typed fields… - #8514

Merged
mattcasters merged 2 commits into
apache:mainfrom
bamaer:8511
Sep 24, 2026
Merged

mattcasters merged 2 commits into
apache:mainfrom
bamaer:8511

Conversation

@bamaer

@bamaer bamaer commented Sep 21, 2026

Copy link
Copy Markdown
Contributor

… out of text

Hop can chunk, embed, store and search text. All of that is retrieval. This is the opposite direction and the one closest to what Hop is for: a contract becomes a renewal date and a value, a ticket becomes a severity and a product.

The field grid is the whole configuration. Each row becomes a property in the JSON schema the model is constrained by, a column on the stream, and the type the answer is read back into.

  • where the provider supports it the grid is sent as a real JSON schema, checked per model through supportedCapabilities(); where it is not, the same fields go in the prompt and the answer is checked on the way back
  • allowed values become a JSON enum, which constrains the model rather than asking it, and is the reliable way to classify
  • optional fields are offered null as a schema branch. Without it a model answers 0 for an amount the text never mentions, which reads downstream as "nothing was lost" rather than "the text does not say"
  • a value that will not convert names the field and what came back, and is never written as a silent null, because downstream an absent value and a misread one look identical
  • an error hop diverts the offending row and the run continues

Also adds AiChatModelFactory, the chat counterpart of AiEmbeddingFactory, and lifts the endpoint, credentials and timeout both need into AiProviderSettings. A later rerank or agent transform needs the same chat model. That shared code keeps the embedding transform's existing timeout parsing exactly as it was.

Includes a sample pipeline, a user manual page, and an integration test in the shared ai project against a real Ollama model.

Also corrects the abort message in the Embed text integration test, which named the Language Model Chat transform.

Please add a meaningful description for your change here


Thank you for your contribution! Follow this checklist to help us incorporate your contribution quickly and easily:

  • Run mvn clean install apache-rat:check to make sure basic checks pass. A more thorough check will be performed on your pull request automatically.
  • If you have a group of commits related to the same change, please squash your commits into one and force push your branch using git rebase -i.
  • Mention the appropriate issue in your description (for example: addresses #123), if applicable.

To make clear that you license your contribution under the Apache License Version 2.0, January 2004
you have to acknowledge this by using the following check-box.

…fields out of text

Hop can chunk, embed, store and search text. All of that is retrieval. This is
the opposite direction and the one closest to what Hop is for: a contract
becomes a renewal date and a value, a ticket becomes a severity and a product.

The field grid is the whole configuration. Each row becomes a property in the
JSON schema the model is constrained by, a column on the stream, and the type
the answer is read back into.

- where the provider supports it the grid is sent as a real JSON schema, checked
  per model through supportedCapabilities(); where it is not, the same fields go
  in the prompt and the answer is checked on the way back
- allowed values become a JSON enum, which constrains the model rather than
  asking it, and is the reliable way to classify
- optional fields are offered null as a schema branch. Without it a model
  answers 0 for an amount the text never mentions, which reads downstream as
  "nothing was lost" rather than "the text does not say"
- a value that will not convert names the field and what came back, and is never
  written as a silent null, because downstream an absent value and a misread one
  look identical
- an error hop diverts the offending row and the run continues

Also adds AiChatModelFactory, the chat counterpart of AiEmbeddingFactory, and
lifts the endpoint, credentials and timeout both need into AiProviderSettings.
A later rerank or agent transform needs the same chat model. That shared code
keeps the embedding transform's existing timeout parsing exactly as it was.

Includes a sample pipeline, a user manual page, and an integration test in the
shared ai project against a real Ollama model.

Also corrects the abort message in the Embed text integration test, which named
the Language Model Chat transform.
+ "'.");
}

return switch (settings.type()) {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[bug] supportedCapabilities() is not a live probe of the provider. In langchain4j 1.20 both OllamaChatModel and OpenAiChatModel return only the set passed to the builder, and an unset set is empty. StructuredExtract.supportsJsonSchema() therefore always takes the prompt-only branch, and the schema built in openModel() is never attached as responseFormat. The integration test comment says Ollama constrains qwen2.5:0.5b to the schema; with this factory it does not, so enum and optional-null behavior depend on an unconstrained small model. additionalProperties(false) is also dropped on the wire unless OpenAI strictJsonSchema is true, which this builder never sets.

Suggestion: Call .supportedCapabilities(Capability.RESPONSE_FORMAT_JSON_SCHEMA) on both builders, and .strictJsonSchema(true) on the OpenAI builder. Add a test that a request built the way StructuredExtract.ask builds it carries a JSON-schema response format.

this,
getMetadataProvider());

JsonSchema schema = ExtractionSchema.build(data.fields, getTransformName());

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[bug] The schema name sent to the provider is the transform name. OpenAI requires response_format.json_schema.name to be a-z, A-Z, 0-9, underscores or dashes, at most 64 characters, and rejects anything else with HTTP 400. The sample pipeline and the integration test both name this transform "Structured extract", so the first OpenAI row fails as soon as the schema is actually attached. Ollama only forwards the root element and ignores the name, which hides it.

Suggestion: Pass a sanitized token (replace characters outside [A-Za-z0-9_-], truncate to 64, and fall back to "extraction" when nothing remains), not getTransformName().

case IValueMeta.TYPE_BIGNUMBER ->
node.isNumber() ? node.decimalValue() : new BigDecimal(text.trim());
case IValueMeta.TYPE_BOOLEAN -> toBoolean(text.trim(), field);
case IValueMeta.TYPE_DATE, IValueMeta.TYPE_TIMESTAMP -> toDate(text.trim(), field);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[bug] TYPE_TIMESTAMP is coerced through toDate, which returns java.util.Date. ValueMetaTimestamp.getTimestamp casts the native value to java.sql.Timestamp, so the first preview, clone, or getString throws ClassCastException downstream of this transform's error hop. SimpleDateFormat.parse(String) also accepts a yyyy-MM-dd prefix and ignores the rest, so 2026-03-01T10:30:00 becomes midnight and 2026-03-01 oops is treated as a valid date. ExtractionSchema asks for both Date and Timestamp as yyyy-MM-dd only, so a Timestamp column cannot carry a time even when the model returns one.

Suggestion: Return new Timestamp(parsed.getTime()) for TYPE_TIMESTAMP, parse with a ParsePosition that requires the whole string, try a date-time pattern before the date-only one, and describe Timestamp fields to the model as a date-time rather than yyyy-MM-dd.

Object[] values = new Object[fields.size()];
for (int i = 0; i < fields.size(); i++) {
StructuredExtractField field = fields.get(i);
JsonNode node = root.get(field.getName());

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[bug] The schema, the prompt, and getFields all key the field by trimmedName(), but the answer is looked up with field.getName(). A grid value with surrounding whitespace is requested as severity and read back as severity, so the output column is always null. That is the silent null this transform is written to avoid, and it is not reported as a conversion error.

Suggestion: Look up field.trimmedName(). Trim the name in StructuredExtractDialog.readFields as well, so stored metadata matches the schema key.

for (int i = 0; i < fields.size(); i++) {
StructuredExtractField field = fields.get(i);
JsonNode node = root.get(field.getName());
values[i] = node == null || node.isNull() ? null : coerce(node, field);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[suggestion] Allowed values become a JSON enum, but parse never checks them. On the prompt-only path, and on any provider that does not enforce the schema, "critical" is written into a field constrained to low,medium,high and the row looks successful.

Suggestion: After coercion, if ExtractionSchema.allowedValues(field) is non-empty and the returned text is not in that list, throw a HopException that names the field, the value, and the allowed list so the error hop can divert the row.

Set<String> seen = new HashSet<>();
for (StructuredExtractField field : named) {
String name = field.trimmedName();
if (!seen.add(name)) {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[suggestion] Duplicate detection uses a case-sensitive HashSet, while IRowMeta.indexOfValue and addValueMeta are case-insensitive. Total and total both pass check, then getFields silently renames the second column (for example total_1) and the extracted values no longer sit under the names that were configured.

Suggestion: Compare trimmed names with equalsIgnoreCase here and in ExtractionSchema.build.

- declare JSON schema support on the Ollama and OpenAI chat models; OpenAI uses strict mode. Other OpenAI compatible providers keep the prompt-only path
- sanitize the schema name to [A-Za-z0-9_-], at most 64 characters
- Timestamp fields return java.sql.Timestamp, dates must match in full, and Timestamps are asked for as a date-time
- read answers by the trimmed field name; the dialog stores names trimmed
- reject answers outside the allowed values, including an empty answer for a required field
- detect duplicate field names case-insensitively
- pin the integration test chat model to four Ollama threads
@bamaer

bamaer commented Sep 24, 2026

Copy link
Copy Markdown
Contributor Author

Thanks for the review. All six points are addressed in the latest commit.

  1. supportedCapabilities(): the Ollama and OpenAI chat models now declare RESPONSE_FORMAT_JSON_SCHEMA, and OpenAI sets strictJsonSchema(true). Captured requests confirm Ollama receives the schema as format, and OpenAI receives response_format with json_schema, strict: true, all properties in required and additionalProperties: false. Gemini, Grok and custom OpenAI compatible providers are not sent a schema, because they differ in which schema keywords they accept; they keep the prompt-only path. A test builds each provider's model through the factory and checks the response format on the request.
  2. Schema name: characters outside [A-Za-z0-9_-] are replaced, the name is truncated to 64 characters, and extraction is used when nothing is left. "Structured extract" is sent as Structured_extract.
  3. Timestamp: Timestamp fields return java.sql.Timestamp. Parsing uses java.time, so the whole value must match: 2026-03-01 oops and 2026-02-30 are rejected. A date-time is tried before a date, so a returned time is kept. Timestamp fields are asked for as yyyy-MM-ddTHH:mm:ss, with an example.
  4. Trimmed names: the answer is read by trimmedName(), and the dialog stores names trimmed.
  5. Allowed values: an answer outside the list throws a HopException naming the field, the value and the list. An empty answer counts as outside the list for a required field.
  6. Duplicates: check and ExtractionSchema.build compare names case-insensitively.

The integration test now also runs against the schema path. It timed out on a 16 CPU Docker host because Ollama used all 16 threads. The test model is now a derived qwen2.5-hop-it with num_thread 4, and 0001 and 0002 pass. The user manual page is updated to match.

@mattcasters
mattcasters merged commit a3e49e5 into apache:main Sep 24, 2026
5 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants