Skip to content

OPENNLP-1930: Parse Arvores Deitadas markup with cursor scans - #1276

Draft
krickert wants to merge 20 commits into
apache:mainfrom
ai-pipestream:OPENNLP-1930-ad-markup
Draft

OPENNLP-1930: Parse Arvores Deitadas markup with cursor scans#1276
krickert wants to merge 20 commits into
apache:mainfrom
ai-pipestream:OPENNLP-1930-ad-markup

Conversation

@krickert

@krickert krickert commented Sep 7, 2026

Copy link
Copy Markdown
Contributor

Based on #1275, and independent of the other parts of the epic. It can be reviewed and merged in any order relative to them. This diff carries the #1275 commits until that one merges; the commits of this change alone: ai-pipestream/opennlp@OPENNLP-1928-regex-removal-trivial...OPENNLP-1930-ad-markup

Replaces the remaining regular expressions in the AD corpus readers with cursor parsers that accept the same lines as the patterns they replace.

Class Patterns removed
ADSentenceStream.SentenceParser node, leaf, and bizarre-leaf line patterns; the \w.*?[\.<>].* lexeme check
ADSentenceStream the nine sentence, title, box, paragraph, and text tag patterns matched per line
ADNameSampleStream the three corpus-specific metadata patterns (text id, paragraph, source)

A package-private ADMetadata holds the metadata scan shared by ADNameSampleStream and ADSentenceSampleStream; the copy from the trivial batch in ADSentenceSampleStream is removed, and the line-terminator rule of the original patterns (a dot does not cross a line terminator) is restored there.

Verification: a side-by-side harness ran the original patterns and the new parsers over the ad.sample corpus, hand-written edge lines, random mutations, and random strings, about 911,000 lines for the tree parser, 300,000 for the tag checks, and 500,000 for the metadata, with zero differences. ADSentenceStreamTest grows from 36 to 83 cases, ADNameSampleStreamTest from 58 to 64, ADMetadataTest adds 61. opennlp-formats: 607 tests with checkstyle, offline, -Dopennlp.forkCount=1.

OPENNLP-1930

…in DefaultPOSContextGenerator

Add pinning tests for the accept and reject sides of both predicates.
… checks in FeatureGeneratorUtil

Add pinning tests for the capPeriod accept and reject sides.
…enPatternFeatureGenerator

Add a pinning test that non-letter sub-tokens do not produce st= features.
… char scan

Matches (.+)-\w+ semantics: group(1) is everything before the last hyphen,
the hyphen must not be at index 0, and the suffix must be non-empty word
chars. Add pinning tests for outcomes without hyphen, hyphen at index 0,
empty suffix, non-word suffix, and the normal accept case.
…n BrownCluster

Replicates String.split(\t) semantics, including dropped trailing empty fields.
…explicit char scans in TokenSampleStream

splitOnWhitespace replicates String.split(\\s+): a leading whitespace run
yields one empty leading field, runs collapse, and trailing empty fields are
dropped.
…mojiCharSequenceNormalizer

The replaced pattern contains a high surrogate range, so the regex engine
matches whole code points in the flattened range [U+D83C, U+10FC00]. The
replacement scans code points, collapses each maximal matching run into a
single space, and copies non-matching code points verbatim. Add pinning
tests for unpaired surrogates, BMP chars above U+D83C, and supplementary
code points beyond U+10FC00.
…xplicit scans in ConlluStream

splitOnHyphen replicates String.split("-"): every hyphen is a boundary,
empty fields between consecutive hyphens are kept, and trailing empty
fields are dropped.

extractTextLang replicates find() of text_([a-z]{2,3}): the first
occurrence of "text_" followed by two to three ASCII lowercase letters,
preferring three.
…ss scans in ParserTool

The two replaceAll passes are replicated by two cursor passes with the
same leftmost-first resume-after-match semantics, which matters for
overlapping pairs such as "x((" or "((a)(b))": a pair starting at the
second char of a match is only reconsidered by the second pass.
…cans in DownloadUtil

parseChecksum now scans to the first ASCII whitespace character,
replicating split(\s)[0] on the trimmed content.

extractLinks replicates find() of the <a href="(.*?)">(.*?)</a> pattern
with CASE_INSENSITIVE and DOTALL flags: the href value ends at the first
"> and the first case-insensitive </a> closes the match, so nested link
markup is swallowed by the outer match.
…meric patterns with explicit scans in ADNameSampleStream

splitOnWhitespace and splitOnUnderscores replicate run-based splitting: a
leading separator run yields one empty leading field, trailing empty
fields are dropped, and an all-separator input yields no fields.

matchHyphenatedToken replicates the three-branch hyphen pattern at code
point granularity, isAlphaNumeric replicates ^[\p{L}\p{Nd}]+$ via
Character.isLetter and Character.isDigit, and tagContent replicates
matches() of <(NER:)?(.*?)> including its optional NER: prefix.
…OSSampleStream

replaceWhitespaceWithEquals replicates replaceAll("=") of the \s+
pattern: every run of ASCII whitespace, including leading and trailing
runs, is replaced by a single equals sign.
…tenceSampleStream

parseTextAndParagraph replicates matches() of the
^(?:[a-zA-Z\-]*(\d+)).*?p=(\d+).* pattern: after the optional ASCII
letters and hyphens, the text id is the first ASCII digit run and the
paragraph id is the digit run after the first "p=" that is followed by
at least one digit.
…entenceStream

replaceGuillemetPunctuation replicates replaceAll of the »\s+ punct
patterns: every run of ASCII whitespace between » and the punctuation
character is removed.

parsePunctuationLine replicates matches() of the ^(=*)(\W+)$ pattern:
the line consists of leading equals signs followed by one or more
non-word characters, where a word character is an ASCII letter, digit,
or underscore. A line of only equals signs matches, with the last
equals sign as lexeme.
Adds StringUtil.isAsciiWhitespace, splitOnAsciiWhitespace,
containsAsciiUpperCase, and containsAsciiDigit and removes the copies
from the AD streams, the English TokenSampleStream, DownloadUtil, and
the POS and lemmatizer context generators. NameFinderME.extractNameType
delegates to BioCodec.

Cases the new tests found first: the TokenSampleStream split returned
one empty token for a whitespace-only line where the original split
returned none, and matchHyphenatedToken accepted a single hyphen.
BrownCluster.splitTabs now removes all trailing empty fields, as
String.split does.

Each helper has a test, parameterized where the inputs are a table,
with the reject side and the edge cases: empty input, leading and
trailing separators, non-ASCII spaces and digits, and
supplementary-plane characters. Helpers only called from instance
methods are no longer static; the block comments on the helpers are
now Javadoc that states the behavior.
The tag patterns in the read loop, applied to a full line, accepted an
opening tag as the tag name right after the angle bracket, any
characters other than a closing angle bracket, and the closing angle
bracket as the last character; a closing tag was the exact text of
that tag. isOpeningTag and isClosingTag check the same for a given tag
name, and the tag names are constants.
The node, leaf, and bizarre leaf patterns in getElement read a line as a
level prefix of equals signs and hyphens, a tag with a colon or an
equals sign, and then, for leaves, a quoted lemma, secondary tags in
angle brackets, a morphological tag, the closing parenthesis, and the
lexeme after whitespace. The patterns backtracked in three places: the
prefix could give up hyphens to the tag, the lemma took the last quote
after which the rest of the line still parsed, and the secondary tags
took the longest run of angle-bracket tags that still left a rest. The
scan methods keep each of those preferences: scanLevelAndTag hands out
the prefix candidates longest first, parseLeafAfterTag tries the lemma
ends from the right, and scanSecondaryTags searches tag ends from the
right with a note of the positions after which no rest exists. Line
terminators end the lemma, the tags, and the lexeme as the dot did. The
lexeme check in the fallback branch is isWordWithMarkup. A side by side
run of the old patterns and the new scan over the sample corpus, hand
written edge lines, mutations of both, and random lines showed no
difference; the tests cover the sample lines, the fall through cases,
quotes inside lemma and lexeme, the tag groups, and the bizarre form.
ADNameSampleStream compiled a metadata pattern per call, chosen by the
text collection: for literary texts the ASCII letters and hyphens
before the text id, for CIE the value of the first source attribute,
and otherwise the digits of the text id, each only when a paragraph
number follows a p= later in the line. ADSentenceSampleStream had a
second copy of the last scan. The package-private ADMetadata now
contains the one scan for text and paragraph ids with accessors for
the parsed values, the text digits, and the letter prefix, plus the
source lookup, and both streams call it. As the dot in the patterns
did not cross a line terminator, metadata with one is invalid in all
cases. The unused Type enum and the commented-out patterns are removed.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant