Repository navigation
Fixes #8511 : Add a Structured extract transform to read typed fields… - #8514
Conversation
…fields out of text Hop can chunk, embed, store and search text. All of that is retrieval. This is the opposite direction and the one closest to what Hop is for: a contract becomes a renewal date and a value, a ticket becomes a severity and a product. The field grid is the whole configuration. Each row becomes a property in the JSON schema the model is constrained by, a column on the stream, and the type the answer is read back into. - where the provider supports it the grid is sent as a real JSON schema, checked per model through supportedCapabilities(); where it is not, the same fields go in the prompt and the answer is checked on the way back - allowed values become a JSON enum, which constrains the model rather than asking it, and is the reliable way to classify - optional fields are offered null as a schema branch. Without it a model answers 0 for an amount the text never mentions, which reads downstream as "nothing was lost" rather than "the text does not say" - a value that will not convert names the field and what came back, and is never written as a silent null, because downstream an absent value and a misread one look identical - an error hop diverts the offending row and the run continues Also adds AiChatModelFactory, the chat counterpart of AiEmbeddingFactory, and lifts the endpoint, credentials and timeout both need into AiProviderSettings. A later rerank or agent transform needs the same chat model. That shared code keeps the embedding transform's existing timeout parsing exactly as it was. Includes a sample pipeline, a user manual page, and an integration test in the shared ai project against a real Ollama model. Also corrects the abort message in the Embed text integration test, which named the Language Model Chat transform.
| + "'."); | ||
| } | ||
|
|
||
| return switch (settings.type()) { |
There was a problem hiding this comment.
[bug] supportedCapabilities() is not a live probe of the provider. In langchain4j 1.20 both OllamaChatModel and OpenAiChatModel return only the set passed to the builder, and an unset set is empty. StructuredExtract.supportsJsonSchema() therefore always takes the prompt-only branch, and the schema built in openModel() is never attached as responseFormat. The integration test comment says Ollama constrains qwen2.5:0.5b to the schema; with this factory it does not, so enum and optional-null behavior depend on an unconstrained small model. additionalProperties(false) is also dropped on the wire unless OpenAI strictJsonSchema is true, which this builder never sets.
Suggestion: Call .supportedCapabilities(Capability.RESPONSE_FORMAT_JSON_SCHEMA) on both builders, and .strictJsonSchema(true) on the OpenAI builder. Add a test that a request built the way StructuredExtract.ask builds it carries a JSON-schema response format.
| this, | ||
| getMetadataProvider()); | ||
|
|
||
| JsonSchema schema = ExtractionSchema.build(data.fields, getTransformName()); |
There was a problem hiding this comment.
[bug] The schema name sent to the provider is the transform name. OpenAI requires response_format.json_schema.name to be a-z, A-Z, 0-9, underscores or dashes, at most 64 characters, and rejects anything else with HTTP 400. The sample pipeline and the integration test both name this transform "Structured extract", so the first OpenAI row fails as soon as the schema is actually attached. Ollama only forwards the root element and ignores the name, which hides it.
Suggestion: Pass a sanitized token (replace characters outside [A-Za-z0-9_-], truncate to 64, and fall back to "extraction" when nothing remains), not getTransformName().
| case IValueMeta.TYPE_BIGNUMBER -> | ||
| node.isNumber() ? node.decimalValue() : new BigDecimal(text.trim()); | ||
| case IValueMeta.TYPE_BOOLEAN -> toBoolean(text.trim(), field); | ||
| case IValueMeta.TYPE_DATE, IValueMeta.TYPE_TIMESTAMP -> toDate(text.trim(), field); |
There was a problem hiding this comment.
[bug] TYPE_TIMESTAMP is coerced through toDate, which returns java.util.Date. ValueMetaTimestamp.getTimestamp casts the native value to java.sql.Timestamp, so the first preview, clone, or getString throws ClassCastException downstream of this transform's error hop. SimpleDateFormat.parse(String) also accepts a yyyy-MM-dd prefix and ignores the rest, so 2026-03-01T10:30:00 becomes midnight and 2026-03-01 oops is treated as a valid date. ExtractionSchema asks for both Date and Timestamp as yyyy-MM-dd only, so a Timestamp column cannot carry a time even when the model returns one.
Suggestion: Return new Timestamp(parsed.getTime()) for TYPE_TIMESTAMP, parse with a ParsePosition that requires the whole string, try a date-time pattern before the date-only one, and describe Timestamp fields to the model as a date-time rather than yyyy-MM-dd.
| Object[] values = new Object[fields.size()]; | ||
| for (int i = 0; i < fields.size(); i++) { | ||
| StructuredExtractField field = fields.get(i); | ||
| JsonNode node = root.get(field.getName()); |
There was a problem hiding this comment.
[bug] The schema, the prompt, and getFields all key the field by trimmedName(), but the answer is looked up with field.getName(). A grid value with surrounding whitespace is requested as severity and read back as severity, so the output column is always null. That is the silent null this transform is written to avoid, and it is not reported as a conversion error.
Suggestion: Look up field.trimmedName(). Trim the name in StructuredExtractDialog.readFields as well, so stored metadata matches the schema key.
| for (int i = 0; i < fields.size(); i++) { | ||
| StructuredExtractField field = fields.get(i); | ||
| JsonNode node = root.get(field.getName()); | ||
| values[i] = node == null || node.isNull() ? null : coerce(node, field); |
There was a problem hiding this comment.
[suggestion] Allowed values become a JSON enum, but parse never checks them. On the prompt-only path, and on any provider that does not enforce the schema, "critical" is written into a field constrained to low,medium,high and the row looks successful.
Suggestion: After coercion, if ExtractionSchema.allowedValues(field) is non-empty and the returned text is not in that list, throw a HopException that names the field, the value, and the allowed list so the error hop can divert the row.
| Set<String> seen = new HashSet<>(); | ||
| for (StructuredExtractField field : named) { | ||
| String name = field.trimmedName(); | ||
| if (!seen.add(name)) { |
There was a problem hiding this comment.
[suggestion] Duplicate detection uses a case-sensitive HashSet, while IRowMeta.indexOfValue and addValueMeta are case-insensitive. Total and total both pass check, then getFields silently renames the second column (for example total_1) and the extracted values no longer sit under the names that were configured.
Suggestion: Compare trimmed names with equalsIgnoreCase here and in ExtractionSchema.build.
- declare JSON schema support on the Ollama and OpenAI chat models; OpenAI uses strict mode. Other OpenAI compatible providers keep the prompt-only path - sanitize the schema name to [A-Za-z0-9_-], at most 64 characters - Timestamp fields return java.sql.Timestamp, dates must match in full, and Timestamps are asked for as a date-time - read answers by the trimmed field name; the dialog stores names trimmed - reject answers outside the allowed values, including an empty answer for a required field - detect duplicate field names case-insensitively - pin the integration test chat model to four Ollama threads
|
Thanks for the review. All six points are addressed in the latest commit.
The integration test now also runs against the schema path. It timed out on a 16 CPU Docker host because Ollama used all 16 threads. The test model is now a derived |
… out of text
Hop can chunk, embed, store and search text. All of that is retrieval. This is the opposite direction and the one closest to what Hop is for: a contract becomes a renewal date and a value, a ticket becomes a severity and a product.
The field grid is the whole configuration. Each row becomes a property in the JSON schema the model is constrained by, a column on the stream, and the type the answer is read back into.
Also adds AiChatModelFactory, the chat counterpart of AiEmbeddingFactory, and lifts the endpoint, credentials and timeout both need into AiProviderSettings. A later rerank or agent transform needs the same chat model. That shared code keeps the embedding transform's existing timeout parsing exactly as it was.
Includes a sample pipeline, a user manual page, and an integration test in the shared ai project against a real Ollama model.
Also corrects the abort message in the Embed text integration test, which named the Language Model Chat transform.
Please add a meaningful description for your change here
Thank you for your contribution! Follow this checklist to help us incorporate your contribution quickly and easily:
mvn clean install apache-rat:checkto make sure basic checks pass. A more thorough check will be performed on your pull request automatically.git rebase -i.addresses #123), if applicable.To make clear that you license your contribution under the Apache License Version 2.0, January 2004
you have to acknowledge this by using the following check-box.