Repository navigation
Add tense-matched forms of formal verbs and a few phrases to the plain-language word list - #95
Merged
Conversation
… list
The table only matches exact words, so "obtain" was flagged but "obtained",
"submitted" and "presented" were not, and suggesting "get" for "obtained"
would read wrong anyway. Add the inflected forms as their own entries with
matching suggestions, plus phrases the table lacked ("set forth",
"aforementioned", "in the event that", "whereby"). "employed" is added only
inside "is/are/currently employed", since "self-employed" is plain already.
Checked against the 180 question files of published Massachusetts
interviews in ~/all_interviews: the additions raise findings from 1555 to
1632. By hand review, 72 of the 78 new findings are real; the other 6 are
quoted statutes, a quoted affidavit, a form section title, and a term the
author already glossed. A sample of 40 existing findings had about 8 such
false positives.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
nonprofittechy
added a commit
to SuffolkLITLab/docassemble-ALWeaver
that referenced
this pull request
Oct 6, 2026
Matching every ending of every table word flagged ordinary words like "required", "completed" and "self-employed" when checked against 180 published interviews. Use exact matches, as the linter does, plus the tense-matched entries proposed in SuffolkLITLab/DAYamlChecker#95 until a release includes them. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What this changes
plain_language_replacements.ymlonly matches exact words. So "obtain" is flagged, but "obtained", "submitted" and "presented" are not, and suggesting "get" in place of "obtained" would read wrong anyway. This adds:obtained: [got, received],submitting: sending,expired: [ran out, ended],granted: [approved, given],provided: [gave, given],incurred: [had, owed],purchased: bought,terminated: [ended, stopped].set forth,aforementioned,in the event that,whereby.is employed,are employedandcurrently employed. The bare word would also match "self-employed", which is plain already.Forms that are usually ordinary words or names were left out on purpose: "required", "benefits", "completed", "options", "exhibits", "deemed" (as in "Petition … to be Deemed Satisfied"), and "authorized" (as in "authorized representative").
The ALWeaver now runs this table over AI-drafted labels and screen text, so the two tools agree on which words to avoid.
Validation against published interviews
I ran the linter's style check (
RuntimeOptions(style_enabled=True)) over all 180data/questions/*.ymlfiles from 50+ published Massachusetts AssemblyLine interviews, before and after this change, and compared the findings (scan script).mainThe one removed finding is in "Are you currently employed?". The new
currently employed → working nowreplaces the oldcurrently → now.All 78 new findings were reviewed by hand. 72 are formal wording in text the litigant reads, where the suggestion fits. For example:
The other 6 are false positives:
That's about 8%. For comparison, a random sample of 40 existing findings had about 8 false positives (~20%) of the same kinds: names and defined terms ("housing authority", "Assistance animal"), quoted affidavit and lease text, and suggestions that don't fit ("correct → exact"). So the new entries are no noisier than the table already is.
While testing I also found that a duplicate key makes the table fail to load (ruamel rejects duplicates), and the plain-language check then silently reports nothing. This change adds no duplicates (
herebyandin the event ofwere already there, so I left them out), but a test that loads the table and fails on duplicates might be worth adding separately.Tests
pytest tests: 655 passed.🤖 Generated with Claude Code