Skip to content

Add offline spellchecking for visible interview text - #96

Merged
nonprofittechy merged 5 commits into
mainfrom
feature/spellchecking
Oct 7, 2026
Merged

nonprofittechy merged 5 commits into
mainfrom
feature/spellchecking

Conversation

@nonprofittechy

@nonprofittechy nonprofittechy commented Oct 7, 2026 •

Copy link
Copy Markdown
Member

Summary

Adds a local spellcheck pass for the visible text interview users see: questions, subquestions, field labels, help, choice labels, and template content. It skips code, stored values, Mako, Markdown/HTML markup, and URLs.

This brings DAYamlChecker closer to parity with the original linter script, which previously ran spelling checks.

  • Stable spelling rule codes: SP701 (spelling_possible_typo) and SP702 (spelling_common_legal_typo, such as judgement → judgment and HIPPA → HIPAA) under a new spelling finding class. The codes remain SP701/SP702 at every configured severity, so suppressions stay stable.
  • On by default as warnings, using US English. Turn spellchecking off with --no-spellcheck or RuntimeOptions(spellcheck=None).
  • Configurable severity: --spellcheck-severity info|warning|error changes the reported level and exit behavior without changing the rule code. This lets new lint rules begin as warnings and become enforceable when a project is ready.
  • Languages: en, es, ru, and sv are built in. Custom Hunspell dictionaries can be added with --spellcheck-dictionary LANG=PATH. Blocks with a declared language outside the selected languages are skipped.
  • Configuration: supports --spellcheck-language, --spellcheck-wordlist, --spellcheck-ignore-word, and --spellcheck-severity, plus SpellcheckOptions in Python. Findings can be suppressed with # no-dayc comments by code, message ID, or class.
  • Designed for few false positives: skips acronyms, mixed-case words, short words, unfamiliar proper names, author names declared in metadata, names in credits and copyright sections, and geographic choice lists.
  • New public API: exports find_spelling_findings_from_string and SpellcheckOptions. scripts/evaluate_spelling.py compares results on a corpus of interviews.

Licensing and dictionaries

Uses spylls (MIT; pure-Python Hunspell) as a dependency. Its English, Russian, and Swedish dictionaries require no network access.

The Spanish dictionary (RLA-ES, via LibreOffice) is not shipped, so the package remains plain MIT. The first run selecting es downloads es_US.aff and .dic (about 850 KB) from a pinned LibreOffice commit and verifies their SHA-256 hashes. Files are cached under $DAYAMLCHECKER_CACHE_DIR when configured; otherwise the platform user-cache directory is used ($XDG_CACHE_HOME/dayamlchecker, normally ~/.cache/dayamlchecker, or %LOCALAPPDATA%\dayamlchecker\cache on Windows). Download failures point users to --spellcheck-dictionary es=PATH for offline use.

Notes for reviewers

  • Existing CI that uses --max-warnings 0 may begin failing on spelling warnings because spellchecking is enabled by default.
  • Per-run cost is mostly loading the pure-Python dictionary: about 0.4 seconds for English, plus about 0.9 seconds when Spanish is enabled.

Testing

  • uv run pytest -q tests: 748 passed, 56 subtests passed.
  • Download, cache reuse, corrupted-cache recovery, hash mismatch, and offline failures are tested with a fake remote. Spanish tests use the real dictionary through pytest's cache and skip when offline except under CI.
  • A clean wheel was checked to confirm that it contains no Spanish dictionary files and declares License-Expression: MIT.

nonprofittechy and others added 5 commits October 7, 2026 11:45
Check questions, labels, help and template content against bundled
Hunspell dictionaries (US English by default; Spanish, Russian and
Swedish bundled; custom dictionaries supported) and report possible
typos (WG701) and common legal misspellings (WG702) under a new
"spelling" finding class. Enabled by default in the CLI and API;
disable with --no-spellcheck or RuntimeOptions(spellcheck=None).

Configure with --spellcheck-language, --spellcheck-dictionary,
--spellcheck-wordlist, --spellcheck-ignore-word and
--spellcheck-severity, or SpellcheckOptions in Python. Adds
find_spelling_findings_from_string and scripts/evaluate_spelling.py
for corpus evaluation.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
RLA-ES offers its dictionary under GPLv3+, LGPLv3+ or MPL 1.1+.
Elect MPL 1.1 and drop the GPL/LGPL license texts and the upstream
README that described those options, so the package does not imply
GPL licensing. Declare the license as "MIT AND MPL-1.1" and ship the
dictionary's license files in the wheel metadata.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Fetch es_US.aff/.dic from a pinned LibreOffice dictionaries commit the
first time Spanish is selected, verify their SHA-256 hashes, and cache
them in the user cache directory ($DAYAMLCHECKER_CACHE_DIR, else
$XDG_CACHE_HOME/dayamlchecker or %LOCALAPPDATA%). Later local runs reuse
the cache; a corrupted cached file is fetched again. Download, hash and
cache failures are configuration errors that suggest
--spellcheck-dictionary es=PATH for offline use.

The package no longer redistributes the dictionary, so it is plain MIT
again. Tests keep the dictionary in pytest's cache and skip Spanish
tests when offline, except under CI.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@nonprofittechy
nonprofittechy merged commit 3c2b739 into main Oct 7, 2026
4 checks passed
@nonprofittechy
nonprofittechy deleted the feature/spellchecking branch October 7, 2026 20:17
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant