Skip to content

OPENNLP-547: Add a dependency parser component - #1236

Draft
krickert wants to merge 145 commits into
mainfrom
OPENNLP-547-dependency-parser
Draft

OPENNLP-547: Add a dependency parser component#1236
krickert wants to merge 145 commits into
mainfrom
OPENNLP-547-dependency-parser

Conversation

@krickert

@krickert krickert commented Aug 22, 2026

Copy link
Copy Markdown
Contributor

Summary

  • add a dependency parser API and immutable dependency graph model
  • provide classical transition parsing and a pure Java feedforward training and inference tier
  • add CoNLL-U reading, evaluation, reproducible model persistence, tests, and user documentation

Base

Rebased on main. The Document annotation container (#1182) is merged, so the branch adds only the parser: 34 files, all new, with no change to existing code.

Review round of 2026-09-04

The branch carries the review conventions from the sibling PRs: Javadoc on every method including private helpers and constructors, @throws on each validating path, validation at the public boundary with IllegalArgumentException, named constants, the final-sigma rule in vocabulary normalization, bounded model dimensions, and failing-test commits ahead of each fix.

Verification

opennlp-api, opennlp-runtime, and opennlp-formats build with checkstyle and forbidden APIs; the parser and CoNLL-U test classes pass (128 tests), and the manual chapter's examples are mirrored by ConlluDependencyParserUsageTest.

krickert added 30 commits August 5, 2026 22:10
…yers over the original text

Adds opennlp.tools.document to opennlp-api: Document (immutable, copy-on-add layer
container over the original text), Annotation (a typed value on a Span), LayerKey
(open, typed layer identity), and DocumentAnnotator (pipeline step declaring the
layers it requires and provides). DocumentAnalyzer assembles annotators into a
pipeline validated at build time. Standard keys in Layers cover sentences, tokens,
part-of-speech tags, and entities, populated through thin adapters over the existing
SentenceDetector, Tokenizer, POSTagger, and TokenNameFinder interfaces, which stay
the primary API for single-task use and are unchanged.

All spans refer to the text as supplied. No new dependencies.
…rom missing layers, validate providers at build time
… adaptive data on failure

The lemmatizer adapter now slices tokens and tags per sentence like its POS and
name-finder siblings, so lemmatization decisions never cross a sentence boundary,
and it declares the sentence layer as required. The POS adapter rejects a tagger
that returns a wrong tag count. The name-finder adapter rejects mentions whose
token indices lie outside their sentence instead of silently reading the next
sentence's tokens, clears adaptive data even when annotation fails, and derives
UNTYPED from NameSample.DEFAULT_TYPE instead of re-declaring the literal.
… definition

A blank check under the toolkit's whitespace definition, which unlike
String.isBlank covers the no-break spaces, so annotators validating labels and
identifiers share one predicate instead of each carrying a private copy. Reads
whole code points; tests pin the no-break and figure spaces, the empty string,
and a supplementary-plane letter.
…nt rule

Adds the Document Annotation Container chapter to the manual, with every code
example and every stated span and value mirroring the passing pipeline example
test. The review pass aligns the branch with the project's conventions: layer
key ids validate through StringUtil.isBlank, the annotator interface leaves
thread safety implementation specific, the sentence and tokenizer adapters
document annotate like their siblings, repeated rejection-message literals
become per-class constants, and the name finder test's nine anonymous fixtures
fold into one helper. Layers now states the key placement rule: core layer
keys live there, capability layer keys on their providing annotator.
Every key the toolkit defines now carries the opennlp: id prefix
(opennlp:sentences, opennlp:tokens, opennlp:pos, opennlp:entities,
opennlp:lemmas, opennlp:stems). An extension defines its keys under its own
prefix, and a bare id stays legal for an application-local layer, so ids from
independent producers cannot collide. The rule is stated on Layers, LayerKey,
and in the manual chapter.
A layer key now declares whether its layer is positional or document-scoped.
A positional key, the default, guarantees a span on every annotation, so
consumers never null-check one. A document-scoped key, created through
LayerKey.document, carries whole-document values without spans, the home for
a language id, a category distribution, or provenance. The scope is declared
per key, never per annotation: the container rejects a span-less annotation
under a positional key and a spanned annotation under a document-scoped key,
naming the layer either way. Scope participates in key equality.
…on text

The three invariants the contract tests already enforce are now stated on the
Document interface and in the manual chapter: layers preserve insertion order
and are never reordered, layers are immutable once added and detached from
the caller's input list, and adding a layer is once-only with the rejection
naming the key. Together they keep index-based references between layers
valid for the lifetime of the document.
A corpus may carry a hand-annotated version of a layer beside a produced one.
The convention is a gold: id prefix on the same key scheme, for example
gold:opennlp:tokens beside opennlp:tokens. Because adding a layer is
once-only, competing versions of a layer always live under distinct keys and
never replace each other. Stated on Layers and in the manual chapter, with a
contract test pinning the coexistence.
…le test

Add {@inheritdoc} to the Document, LayerKey, and adapter overrides, and note in
the manual that DocumentPipelineExampleTest asserts the pipeline round-trip.
…ainer contract

- Reject zero-length finder mentions in NameFinderAnnotator and pin the
  second-sentence case, which was previously mapped silently wrong, with a test
- Add DocumentAnnotators with requireLayers and the per-sentence token walk,
  replacing three copies of the walk loop and four spellings of the
  missing-layer rejection; direct tests pin the helpers as public API
- Capture the document text as a String at construction so ImmutableDocument's
  immutability and thread-safety claims hold for mutable CharSequence inputs
- Move the copy-on-add and threading narrative from the Document interface
  Javadoc to ImmutableDocument; the interface now states that thread safety is
  implementation specific
- Carry the entity type as the annotation value only; entity spans are untyped,
  and the Javadoc names the value as the single source of the type
- Make all six adapter annotators final before the types freeze
- Align the TokenLengthAnnotator example with the documented required-layer
  contract in both the manual and the example test via requireLayers
- Housekeeping per review: docbook CDATA placement, imports over qualified
  names, a ParameterizedTest for the blank-input matrix, shared deterministic
  test components, static assertion imports, inheritDoc on the runtime
  adapters, Layers constructor comment, and the documented NPE of
  StringUtil.isBlank
…ll rejection

- Fold the three verbatim copies of the "Ana runs. Bob sits." document into a
  single twoSentenceDocument() helper in NameFinderAnnotatorTest, so the
  sentence and token layers of the shared fixture are declared once instead of
  drifting between the over-long mention, zero-length mention, and
  per-sentence offset tests
- Hoist the no-op TokenNameFinder out of the blank-input test into a NO_NAMES
  constant in DocumentAnalyzerTest, since a finder that returns no spans is
  pipeline plumbing rather than part of any one test case, and document what it
  is for
- Add testAnnotatorAdaptersRejectNullDocuments to pin that all four adapters
  reject a null document with the same "document must not be null" message,
  whether they validate directly or through DocumentAnnotators.requireLayers;
  the shared message was previously unpinned and free to drift per adapter
- Trim the stale "person-free" qualifier from the New York comment in
  testTokenIndexSpansBecomeCharacterSpans; the finder emits a location mention
  and the extra negation described a distinction the test no longer draws
…, pin blank and span edge cases

- ImmutableDocument: wrap the layer map unmodifiable at construction and expose its cached key set; split the combined null check so the message names the offending argument
- StringUtil.isBlank javadoc: state how it differs from isUnicodeBlank
- Tests: parameterize the isBlank accept and reject sides, pin the null NPE, and pin char-indexed spans over a supplementary-plane character
Addresses rzo1's review comments on the manual:

- Open with a plain-language definition and an inline typed-layer
  example instead of a bolded three-item list.
- Show a stacked-layers figure for the running example up front and
  reference it from the pipeline section, so readers see the shape of
  a document before the API detail.
- Explain span offsets, key identity, and the opennlp: prefix
  convention in prose a first-time reader can follow.
- Reword the single-task API aside; drop the redundant custom
  annotator opener; cut the repeated statically-typed phrasing.
Two contract tests fail red against the default method stub:

  java.lang.UnsupportedOperationException: merge is not implemented yet

merge joins two documents grown independently over the same text, the
parallel fan-out join rzo1 asked for on the pull request: disjoint
layers stack into one document, the sources stay untouched, and a null
argument, a different text, or a duplicate layer key is rejected with
the offending key named.
The default body validates the argument and the shared text, then adds
each of the other document's layers through with(), so every layer is
re-validated against this document's contract and a duplicate key is
rejected by the same once-only rule a direct add follows. The pinned
contract tests now pass; opennlp-api is 386 tests, 0 failures.
Two contract tests fail red against the stubbed two-arg merge:

  java.lang.UnsupportedOperationException: merge with policy is not
  implemented yet

The strict single-arg merge stays the default; the policy variant opts
into keeping one copy of a layer both documents rebuilt identically,
and still rejects differing copies with the key named.
merge(other) stays strict and now delegates to merge(other, REJECT).
The KEEP_EQUAL policy keeps one copy of a layer both documents rebuilt
identically, the shared sentence/tokenizer prefix of two parallel
branches, while differing copies are still rejected with the key named.
The pinned contract tests now pass; opennlp-api is 388 tests,
0 failures. The manual's fan-out paragraph documents the option.
- literallayout class=monospaced makes the docbkx toolchain emit a pre
  block, so the figure's character ruler and layer rows column-align;
  plain literallayout renders in the proportional body font.
- Correct the pipeline section: the figure shows three of the four
  layers; the custom token-lengths layer is the fourth.
- Trim restated clauses in the merge javadoc, the contract test
  javadoc, and the introduction; align the layersEqual and addLayer
  helper javadoc with what the helpers do.
The repinned test fails red: a KEEP_EQUAL merge that rejects a layer
whose copies differ still reports 'layer is already present', which
reads as if the policy was ignored. The caller opted into duplicates;
the reason worth naming is that the contents differ.
…ayer

When the policy is KEEP_EQUAL and a layer key is present on both
documents, a failed equality check now throws directly instead of
falling through to with(), so the message states the actual reason:
the copies differ, not merely that the key is a duplicate. The pinned
contract test passes; opennlp-api is 388 tests, 0 failures.
One javadoc sentence on the constant: equality is Annotation equality,
so spans compare by offsets and type, never by probability, and values
by their own equals. Two branches running different models over the
same text can therefore agree; the kept copy is this document's.
Two guards ahead of an ImmutableDocument merge override: the interface
default serves implementations that do not override merge with the
same join, KEEP_EQUAL, and rejection messages, and merge re-validates
the layers it takes from a foreign document instead of trusting them,
rejecting an out-of-bounds span by name.
The interface default adds the other document's layers through with(),
building one intermediate document and one map copy per layer. The
override validates each incoming layer with the same checks with()
runs, then copies the layer map once and allocates one document; when
nothing was added it returns this, matching the default. The layer
validation moves from with() into a shared helper unchanged. Pinned by
the cross-implementation contract tests; opennlp-api is 390 tests,
0 failures, runtime annotator suites green.
The chapter said offsets count characters; the pinned contract test
shows a supplementary-plane character counts as two. Say Java chars
(UTF-16 units) so the claim matches the tested behavior.
The utility constructors, the package-private model constructor, and the
contribution cache constructor get Javadoc, and the two private trainer
helpers that reject a mismatched pretrained vector or an oversized model
state that in their contracts.
@krickert
krickert force-pushed the OPENNLP-547-dependency-parser branch from 83bc298 to 5cc4058 Compare September 4, 2026 21:59
@krickert

krickert commented Sep 4, 2026

Copy link
Copy Markdown
Contributor Author

Rebased on main .. good to go

#1182 merged the diff is the parser alone, 34 new files and no change to existing code.

Also a run of the review conventions.

All published parser patches are present by patch ID. Preserve later local review fixes and the Document API already merged upstream.
Affected reactor tests and the manual package build pass. Preserve the merged resource installer and existing feature fixes.
krickert added a commit to ai-pipestream/opennlp that referenced this pull request Sep 5, 2026
# Conflicts:
#	opennlp-api/src/main/java/opennlp/tools/depparse/DependencyGraph.java
#	opennlp-api/src/main/java/opennlp/tools/depparse/DependencyParser.java
#	opennlp-api/src/main/java/opennlp/tools/depparse/DependencySample.java
#	opennlp-api/src/test/java/opennlp/tools/depparse/DependencyGraphTest.java
#	opennlp-api/src/test/java/opennlp/tools/depparse/DependencySampleTest.java
#	opennlp-core/opennlp-formats/dev/README-ud-treebanks.md
#	opennlp-core/opennlp-formats/dev/download-ud-treebank.sh
#	opennlp-core/opennlp-formats/src/main/java/opennlp/tools/formats/conllu/ConlluDependencySampleStream.java
#	opennlp-core/opennlp-formats/src/test/java/opennlp/tools/formats/conllu/ConlluDependencyParserEvalTest.java
#	opennlp-core/opennlp-formats/src/test/java/opennlp/tools/formats/conllu/ConlluDependencyParserUsageTest.java
#	opennlp-core/opennlp-formats/src/test/java/opennlp/tools/formats/conllu/ConlluDependencySampleStreamTest.java
#	opennlp-core/opennlp-runtime/src/main/java/opennlp/tools/depparse/ArcStandardOracle.java
#	opennlp-core/opennlp-runtime/src/main/java/opennlp/tools/depparse/ArcStandardState.java
#	opennlp-core/opennlp-runtime/src/main/java/opennlp/tools/depparse/DependencyContextGenerator.java
#	opennlp-core/opennlp-runtime/src/main/java/opennlp/tools/depparse/DependencyModel.java
#	opennlp-core/opennlp-runtime/src/main/java/opennlp/tools/depparse/DependencyParserME.java
#	opennlp-core/opennlp-runtime/src/main/java/opennlp/tools/depparse/FeedforwardContext.java
#	opennlp-core/opennlp-runtime/src/main/java/opennlp/tools/depparse/FeedforwardDependencyModel.java
#	opennlp-core/opennlp-runtime/src/main/java/opennlp/tools/depparse/FeedforwardDependencyParser.java
#	opennlp-core/opennlp-runtime/src/main/java/opennlp/tools/depparse/FeedforwardDependencyTrainer.java
#	opennlp-core/opennlp-runtime/src/main/java/opennlp/tools/depparse/Transition.java
#	opennlp-core/opennlp-runtime/src/test/java/opennlp/tools/depparse/ArcStandardOracleTest.java
#	opennlp-core/opennlp-runtime/src/test/java/opennlp/tools/depparse/ArcStandardStateTest.java
#	opennlp-core/opennlp-runtime/src/test/java/opennlp/tools/depparse/DependencyParserEdgeCaseTest.java
#	opennlp-core/opennlp-runtime/src/test/java/opennlp/tools/depparse/DependencyParserMETest.java
#	opennlp-core/opennlp-runtime/src/test/java/opennlp/tools/depparse/DependencyTestSamples.java
#	opennlp-core/opennlp-runtime/src/test/java/opennlp/tools/depparse/FeedforwardDependencyParserTest.java
#	opennlp-core/opennlp-runtime/src/test/java/opennlp/tools/depparse/TransitionTest.java
#	opennlp-docs/src/docbkx/dependency.xml
@jzonthemtn

Copy link
Copy Markdown
Contributor

Thanks @krickert. A long open JIRA ticket for sure!

The API and graph validation look solid. I found two issues worth checking into:

  • The CoNLL-U reader accepts DEPREL=_ as a label instead of skipping the incomplete annotation.
  • Training with a tag such as UNK overwrites a reserved vocabulary entry, producing a model that serializes successfully but fails to reload.

What do you think about held-out accuracy tests for the neural parser, similar to the treebank evaluation for the classical parser?

A CoNLL-U word with a numeric head and the placeholder relation is accepted
instead of making the sentence incomplete.
Training with a tag or label equal to a reserved vocabulary symbol displaces
the reserved row; the model serializes but fails to load.
A relation column that contains the underscore placeholder now makes the tree
incomplete, as an underscore head does, so the reader skips it instead of
reading the placeholder as a label.
…s them

A tag or label equal to a reserved symbol maps to the reserved entry, as words
already did and as unknown symbols do, so the trained model reloads.
UniversalDependencyParserEval extends AbstractEvalTest, loads the Universal
Dependencies 2.0 English splits from OPENNLP_DATA_DIR with MD5 checks, trains
the transition parser and the feedforward parser, and pins the attachment
scores on the development split. It replaces the property-gated test in
opennlp-formats.
@krickert

krickert commented Sep 6, 2026

Copy link
Copy Markdown
Contributor Author

Thanks Jeff. Both are fixed, each with a failing test committed ahead of the fix. The CoNLL-U reader now skips a sentence whose relation column is the underscore placeholder, the same way it already skipped an underscore head. The feedforward trainer no longer lets a tag or label spelled like a reserved symbol take over that symbol's embedding row; it shares the row, as words already did, so the model reloads.

On held-out accuracy: I added UniversalDependencyParserEval in opennlp-eval-tests, following the AbstractEvalTest and OPENNLP_DATA_DIR pattern. It trains both parsers on the UD 2.0 English train split and pins the dev-split scores: the transition parser at UAS 0.8182 and LAS 0.7861, the feedforward parser at UAS 0.8413 and LAS 0.8174, on 25,148 words. The numbers reproduce exactly across runs. It replaces the property-gated evaluation test that lived in opennlp-formats.

Treebank evaluations usually leave punctuation out of the attachment scores. DependencyEvaluator now keeps a second pair of UAS and LAS totals over the tokens whose gold tag is not punctuation, with the universal PUNCT tag as the default and a tag predicate for other tag sets. The existing getters are unchanged.
DependencyCrossValidator splits a DependencySample stream into k folds, trains a parser on the other folds through a caller-supplied trainer, scores the held-out fold with DependencyEvaluator, and sums UAS and LAS over all tokens, so both the transition parser and the feedforward parser can be cross validated. Fold counts below two are rejected.
…dency parser

Every method, constructor, and nested type of the dependency parser packages now has Javadoc with @throws on the paths that validate or fail. Null checks report the offending argument by name, also for the DependencyModel constructors, whose arguments used to reach the superclass unchecked. Repeated literals such as the CoNLL-U column indices, the greedy beam size, the vocabulary names, the training defaults, and the initialization factors are named constants; the values are unchanged.
…fixtures

Tests that looped over cases (all trees up to five tokens, the beam-of-one comparison, the sigma normalization rules, the settings validation) are parameterized tests with named cases. Seeds, sizes, and repeated sample arrays are named constants, and the shared fixture exposes its three distinct sentences. New tests cover the DependencyModel null-argument paths.
…a test

The usage test now runs the manual's feedforward example (train from the CoNLL-U reader with the default settings, save, load, parse with a beam of four) and its loop over the arcs of a parse, on the six-token example sentence with its final period. The manual and the download helper point at UniversalDependencyParserEval as the accuracy evaluation.
…anks

UniversalDependencyParserEval now pins, per run, the attachment scores over all tokens and over the tokens that are not punctuation: the transition parser on the English, German, Spanish AnCora, and French development splits and in a five-fold cross validation of the English training split, the feedforward parser on the English and Spanish AnCora development splits. Each split file is checked against its MD5 digest before training. Every score was measured twice on the same data and reproduced to the last digit.
The treebank README lists the seven evaluation runs, the pinned scores rounded to the asserted precision, the running time, and the difference between the download helper and the shared evaluation data. The manual describes the punctuation-free scores, the cross validator, and the coverage of the treebank evaluation.
…ord lines

The loader and the CoNLL-U reader report format faults as IOException; the
tests now expect the toolkit's InvalidFormatException. Additional tests pin
the messages of training and refinement on samples without an arc-standard
derivation.
The feedforward model loader and the CoNLL-U reader throw the toolkit's
InvalidFormatException for content that is not the expected format, and keep
IOException for read failures. The model's score and featureIds methods become
package-private; only the parser and trainer call them.
ArcStandardOracle.isProjective applies the arc-crossing definition. The event
stream and the feedforward trainer skip and count only non-projective graphs;
any other failure in the oracle now propagates instead of being skipped.
… in the module

ArcStandardState, ArcStandardOracle, and DependencyContextGenerator have no
callers outside the package and lose their public modifiers. ParserInput
validates tokens and tags for the parsers and the context generator, so the
runtime module no longer calls a package-private member of opennlp-api.
@krickert

krickert commented Sep 7, 2026

Copy link
Copy Markdown
Contributor Author

Eval build for this head (afd68fb) is running on Jenkins: https://ci-builds.apache.org/job/OpenNLP/job/eval-tests-configurable/74/ (eval-tests-configurable, branch OPENNLP-547-dependency-parser). It uses the opennlp-data.zip published on nightlies today and includes the new UniversalDependencyParserEval, so expect it to take longer than a regular eval run.

… agree

Keep hidden-layer contributions in double precision, reserve cache
capacity before allocating so the entry budget is not exceeded under
concurrent initialization, reject non-finite transition scores, and
compute the log-softmax in the shifted domain so large scores do not
overflow.
@krickert

krickert commented Sep 7, 2026

Copy link
Copy Markdown
Contributor Author

Evals passed for this, going to re-run it as I added double precision scoring.

@rzo1
rzo1 marked this pull request as draft September 8, 2026 08:17
@rzo1

rzo1 commented Sep 8, 2026

Copy link
Copy Markdown
Contributor

Evals passed for this, going to re-run it as I added double precision scoring.

Moved to draft for now. Please re-open as ready for review once all your additional checks and stuff are done.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants