diff --git a/README.md b/README.md index 9965fc3..b46b17e 100644 --- a/README.md +++ b/README.md @@ -26,6 +26,167 @@ cannot safely rewrite is not itself an error, so it does not fail the run. python3 -m dayamlchecker --fix path/to/interview.yml ``` +## Spelling checks + +Spelling checks run by default on visible questions, labels, choices, help, +and template content, locally. The default is US English, using the Hunspell +dictionary bundled with [Spylls](https://spylls.readthedocs.io/en/latest/hunspell/dictionary.html) +plus a small reviewed English legal and interface vocabulary. Possible mistakes produce +`SP701` (`spelling_possible_typo`). Spelling is independent of `--style` +and never rewrites text. Use `--no-spellcheck` to disable it. + +`--spellcheck-severity info|warning|error` controls the level of all spelling +findings; the default is `warning`. At `error`, spelling findings cause a nonzero +exit code. At `info`, they produce GitHub notices and do not count toward +`--max-warnings`. Rule codes stay stable at every level so suppressions continue +to work. + +```bash +python -m dayamlchecker path/to/interview.yml +python -m dayamlchecker --spellcheck-severity error path/to/interview.yml +python -m dayamlchecker --no-spellcheck path/to/interview.yml +python -m dayamlchecker --spellcheck-wordlist project-words.txt path/to/interview.yml +python -m dayamlchecker --spellcheck-ignore-word Urbana path/to/interview.yml +python -m dayamlchecker --spellcheck-language es path/to/interview.yml +python -m dayamlchecker --spellcheck-language en --spellcheck-language es path/to/interview.yml +``` + +Word lists are UTF-8, one word per line, with `#` comment lines. The repeatable +`--spellcheck-wordlist` option adds project-specific suppressions. Matching is case +insensitive; standard `--suppress` and `# no-dayc: SP701` suppressions apply. +Use repeatable `--spellcheck-ignore-word WORD` for individual suppressions. +Both options affect only the current invocation; they do not alter dictionaries. + +`--spellcheck-language` replaces the English default. Repeat it for mixed-language +interviews: a word accepted by **any** selected dictionary is accepted, including +when languages mix within one sentence. Block `language:` and interview +`default language:` declarations determine which blocks are eligible; declarations +whose base language is not selected are skipped. Unlabelled text uses all selected +dictionaries. This is explicit dictionary selection, not automatic language detection. +Selecting more languages can also hide a typo that is a valid word in another language. + +Built-in languages are `en` (US English), `es` (US Spanish), `ru` and `sv` (Swedish). +`en-US`/`en_US`, `es-US`, `ru-RU` and `sv-SE` are equivalent aliases. Block tags such +as `es-MX` match selected `es`; regional dictionary selection such as `en-GB` +requires a custom dictionary. English, Russian and Swedish dictionaries come +with Spylls and never require network access. + +Spanish uses the RLA-ES Hunspell dictionary distributed by LibreOffice, +including inflection rules and accented words. It is not shipped with this +package: the first run that selects `es` downloads it (about 850 KB) from a +pinned LibreOffice commit, verifies its SHA-256 hashes, and caches it in +`$XDG_CACHE_HOME/dayamlchecker` (default `~/.cache/dayamlchecker`; +`%LOCALAPPDATA%\dayamlchecker\cache` on Windows). Later runs on the same machine +reuse the cache. Set `DAYAMLCHECKER_CACHE_DIR` to choose another location, for +example one saved with `actions/cache`. Fresh CI runners download it once per +run. To work offline, supply a local copy with `--spellcheck-dictionary es=PATH`. + +For other languages or dialects, provide a Hunspell `.aff`/`.dic` pair: + +```bash +# Loads /path/to/fr_FR.aff and /path/to/fr_FR.dic; selects French. +python -m dayamlchecker --spellcheck-dictionary fr=/path/to/fr_FR interview.yml +# Mixes English with the supplied French dictionary. +python -m dayamlchecker --spellcheck-language en --spellcheck-language fr \ + --spellcheck-dictionary fr=/path/to/fr_FR interview.yml +``` + +Dictionary flags are repeatable and can override built-in dictionaries. With no +language flags, the supplied dictionary language codes become the selected languages. +Unknown language codes and missing dictionary files produce configuration errors. +`--no-spellcheck` disables the pass even when language or suppression options +are supplied. The optional `--spellcheck` flag explicitly enables it. + +Python callers can use `dayamlchecker.find_spelling_findings_from_string(yaml_text)` +or pass options to the existing checker functions: + +```python +from dayamlchecker import ( + RuntimeOptions, + SpellcheckOptions, + find_spelling_findings_from_string, +) +from dayamlchecker.messages import Severity + +options = RuntimeOptions( + spellcheck=SpellcheckOptions( + severity=Severity.WARNING, # Or Severity.INFO / Severity.ERROR. + allowed_words=frozenset({"Urbana", "ProjectName"}), + languages=("en", "es"), + # Optional: dictionaries=(("fr", "/path/to/fr_FR"),), + ) +) +findings = find_spelling_findings_from_string(yaml_text, runtime_options=options) +``` + +For the main Python checker APIs, set `RuntimeOptions(spellcheck=None)` +to disable the pass. The spelling-only convenience API always runs it. + +Common legal spellings have explicit recommendations under `SP702` +(`spelling_common_legal_typo`): `judgement` → `judgment` and `judgements` → +`judgments` in US English, and `HIPPA` → `HIPAA` when English is selected. +These checks also cover capitalization, possessives and hyphenated compounds +such as `judgement-proof` and `HIPPA-compliant`. They bypass dictionary-accepted +variants and acronym filtering. Custom British English dictionaries keep their +own accepted variants for `judgement`. Word lists, `--spellcheck-ignore-word`, +`--suppress SP702` and `# no-dayc: SP702` can suppress these recommendations. + +The pass excludes code, stored choice values, object-choice expressions, Mako +expressions, HTML attributes, links' destinations, icons, and Markdown code. +It accepts possessives, recognized hyphenated compounds and spelling variants. +Blocks outside selected languages and long passages mostly outside the selected +dictionaries are skipped. Metadata is excluded. Capitalized names declared in +`metadata.authors` are recognized if they appear in visible prose. Personal-name +spans in contributor, author, credit and copyright sections are excluded, while +ordinary prose in those sections remains checked. + +Capitalized court, county, parish, city and other geographic choice labels are +recognized from the field's label, variable or help; this works across +jurisdictions without a county-name dictionary. Questions and instructions on +those screens remain checked. Other unfamiliar capitalized words produce a +warning only when a nearby lowercase dictionary word supplies spelling evidence +(a single insertion, deletion or transposition). This rule applies at the start +of sentences as well as within them. Acronyms, mixed-case identifiers and short +tokens also receive conservative treatment. Accented tokens are skipped in the +default English-only mode and checked with multilingual/custom dictionaries. +Consequently it can miss typos, particularly in names, capitalized words and valid +words used incorrectly. Locations point to the containing YAML key or field. + +The October 2026 experiment on `~/all_interviews` is recorded in +[`reports/spelling_corpus_review.json`](reports/spelling_corpus_review.json). +Reproduce the comparison without running unrelated lints or URL requests: + +```bash +python scripts/evaluate_spelling.py ~/all_interviews --output /tmp/spelling.json +# Optional: --language en --language es --wordlist project-words.txt +# Custom dictionaries: --dictionary fr=/path/to/fr_FR +``` + +The evaluation reports raw dictionary warnings versus filtered warnings, with +file, word and context for manual review. A warning count alone is not a +false-positive count. The earlier pass had no corpus-specific name list: of 45 +warnings, contextual review identified 44 typos and one organization-name false +positive. It misses the two misspelled car brands found in the earlier run. +The rules were developed on this corpus, so these figures are not an independent +accuracy measurement. Upstream issue links are saved in +[`reports/spelling_issues.json`](reports/spelling_issues.json). + +Before adding the legal spelling recommendations, the English–Spanish run on +the same 180 files produced 44 warnings: the same 44 +reviewed typos, with `Urbana` accepted by the Spanish dictionary and no new warnings. +The English-only warning set was unchanged. That run gave zero reviewed false positives +on this corpus with both languages enabled; it does not establish the accuracy of +Spanish checking on other interviews. Results and commands are recorded in +[`reports/spelling_language_evaluation.json`](reports/spelling_language_evaluation.json). + +With the default-on pass and legal recommendations, the same corpus produces 50 +English-only findings and 49 English–Spanish findings. Five new recommendations +replace `judgement`/`judgements` in references to court judgments. The earlier +warning sets remain intact; the reviewed false-positive counts remain one for +English and zero for English–Spanish. No visible `HIPPA` occurrences were found +in the corpus. The additions and reproducible commands are in +[`reports/spelling_default_evaluation.json`](reports/spelling_default_evaluation.json). + ## Jinja2 preprocessing Files whose first line is exactly `# use jinja` (LF, CRLF, or end of file) are @@ -115,7 +276,7 @@ Template-aware formatting is outside this feature's scope. ## Suppressing checks -You can suppress specific errors or warnings by their ID or finding class (`accessibility`, `style`, `translatability`, `general`). +You can suppress specific errors or warnings by their ID or finding class (`accessibility`, `style`, `translatability`, `spelling`, `general`). **Inline and block comments in YAML:** To suppress a finding on a specific line, use a `# no-dayc: ` comment: diff --git a/pyproject.toml b/pyproject.toml index e5b07b6..4e13555 100644 --- a/pyproject.toml +++ b/pyproject.toml @@ -15,6 +15,7 @@ dependencies = [ "pypdf>=5.0.0", "requests>=2.31.0", "urllib3>=2.0.0", + "spylls>=0.1.7,<0.2", ] license = "MIT" license-files = ["LICEN[CS]E*"] @@ -27,7 +28,7 @@ authors = [ package-dir = { "" = "src" } [tool.setuptools.package-data] -dayamlchecker = ["py.typed", "data/*.yml"] +dayamlchecker = ["py.typed", "data/*.yml", "data/*.txt"] [build-system] requires = ["setuptools >= 77.0.3"] diff --git a/requirements.txt b/requirements.txt index 252af92..8ead48e 100644 --- a/requirements.txt +++ b/requirements.txt @@ -8,5 +8,6 @@ linkify-it-py>=2.0.3 pypdf>=5.0.0 requests>=2.31.0 urllib3>=2.0.0 +spylls>=0.1.7,<0.2 mypy>=1.11.0 types-requests>=2.32.0 diff --git a/scripts/evaluate_spelling.py b/scripts/evaluate_spelling.py new file mode 100644 index 0000000..b465ca8 --- /dev/null +++ b/scripts/evaluate_spelling.py @@ -0,0 +1,145 @@ +"""Compare naive dictionary warnings and conservative spelling on a YAML corpus. + +Run with: python scripts/evaluate_spelling.py ~/all_interviews --output /tmp/spelling.json +The JSON is an audit queue, not automatically labelled ground truth. +""" + +from __future__ import annotations + +import argparse +from collections import Counter +import json +from pathlib import Path +import re +import sys +import time +from typing import Any + +sys.path.insert(0, str(Path(__file__).resolve().parents[1] / "src")) + +from dayamlchecker._jinja import uses_jinja +from dayamlchecker.spelling import ( + dictionary_accepts, + find_spelling_findings, + options_from_cli, + spelling_entries, +) +from dayamlchecker.yaml_structure import ( + _collect_yaml_files, + parse_interview_documents, +) + + +def main() -> None: + parser = argparse.ArgumentParser(description=__doc__) + parser.add_argument("root", type=Path) + parser.add_argument("--output", type=Path, required=True) + parser.add_argument("--wordlist", type=Path, action="append", default=[]) + parser.add_argument("--language", action="append", default=[]) + parser.add_argument( + "--dictionary", action="append", default=[], metavar="LANG=PATH" + ) + args = parser.parse_args() + try: + options = options_from_cli( + languages=args.language, + dictionary_specs=args.dictionary, + wordlists=args.wordlist, + ) + except ValueError as exc: + parser.error(str(exc)) + start = time.monotonic() + files = sorted(set(path.resolve() for path in _collect_yaml_files([args.root]))) + records: dict[str, list[dict]] = {"baseline": [], "filtered": []} + failures: list[dict] = [] + parsed_files = blocks = entries_count = tokens = 0 + for path in files: + try: + content = ( + path.read_text(encoding="utf-8") + .replace("\r\n", "\n") + .replace("\r", "\n") + ) + if uses_jinja(content): + # Match the production checker's preprocessing, without running + # unrelated lints or making network requests. + from dayamlchecker._jinja import render_yaml + + content, missing, unknown = render_yaml(content, str(path)) + if missing or unknown: + failures.append( + {"file": str(path), "error": "partial Jinja rendering"} + ) + docs, parse_errors = parse_interview_documents(content, str(path)) + if parse_errors: + raise ValueError(parse_errors[0].context.get("error", "YAML error")) + except Exception as exc: + failures.append({"file": str(path), "error": str(exc)}) + continue + parsed_files += 1 + blocks += len(docs) + entries = spelling_entries(docs, options) + entries_count += len(entries) + for entry in entries: + words = re.findall(r"\b[^\W\d_]+(?:['’][^\W\d_]+)*\b", entry.text) + tokens += len(words) + seen = set() + for word in words: + lower = word.lower().replace("’", "'") + if ( + lower in seen + or dictionary_accepts(word, options.dictionary_sources) + or dictionary_accepts(lower, options.dictionary_sources) + ): + continue + seen.add(lower) + records["baseline"].append( + { + "file": str(path.relative_to(args.root.resolve())), + "line": entry.line_number, + "location": entry.location, + "word": word, + "snippet": re.sub(r"\s+", " ", entry.text).strip()[:300], + } + ) + records["filtered"].extend( + { + "file": str(path.relative_to(args.root.resolve())), + "line": finding.line_number, + **finding.context, + } + for finding in find_spelling_findings( + docs=docs, input_file=str(path), options=options + ) + ) + summary: dict[str, Any] = { + "files_discovered": len(files), + "files_parsed": parsed_files, + "blocks": blocks, + "text_entries": entries_count, + "raw_word_tokens": tokens, + "parse_or_render_issues": len(failures), + "elapsed_seconds": round(time.monotonic() - start, 2), + } + for name, findings in records.items(): + counts = Counter(record["word"].lower() for record in findings) + summary[name] = { + "warnings": len(findings), + "unique_words": len(counts), + "files_with_warnings": len({f["file"] for f in findings}), + } + args.output.parent.mkdir(parents=True, exist_ok=True) + args.output.write_text( + json.dumps( + {"summary": summary, "issues": failures, **records}, + ensure_ascii=False, + indent=2, + ) + + "\n", + encoding="utf-8", + ) + print(json.dumps(summary, indent=2)) + + +if __name__ == "__main__": + main() diff --git a/src/dayamlchecker/__init__.py b/src/dayamlchecker/__init__.py index fdc61a6..23f30ab 100644 --- a/src/dayamlchecker/__init__.py +++ b/src/dayamlchecker/__init__.py @@ -2,18 +2,22 @@ Finding, FindingClass, ) +from dayamlchecker.spelling import SpellcheckOptions from dayamlchecker.yaml_structure import ( RuntimeOptions, find_errors, find_errors_from_string, find_style_findings_from_string, + find_spelling_findings_from_string, ) __all__ = [ "Finding", "FindingClass", "RuntimeOptions", + "SpellcheckOptions", "find_errors", "find_errors_from_string", "find_style_findings_from_string", + "find_spelling_findings_from_string", ] diff --git a/src/dayamlchecker/data/spelling_words.txt b/src/dayamlchecker/data/spelling_words.txt new file mode 100644 index 0000000..2ed6360 --- /dev/null +++ b/src/dayamlchecker/data/spelling_words.txt @@ -0,0 +1,71 @@ +# Reviewed Docassemble, legal and public-benefit vocabulary. One word per line. +docassemble +arrearage +arrearages +affiant +affiants +efile +efiling +efiled +litigant +litigants +nonlawyer +nonlawyers +noncitizen +noncitizens +unrepresented +# Legal Latin, court terminology and benefits. +appellee +appellees +bona +bono +fide +litem +civ +conservatorship +dispositive +expungement +habilitation +impoundment +indigency +inspectional +memoranda +obligor +obligors +overpayment +overpayments +prehearing +recertification +submittal +unaffordable +unhoused +# Common UI words, compounds and inclusive pronouns. +barcode +birthdate +birthname +checkboxes +checkmarks +daycares +dropoff +financials +funders +github +grandkids +hotline +motorhome +paystub +paystubs +pdf +preprepared +righthand +themself +touchpad +uncheck +webpage +wellbeing +zipcode +zir +zirs +# Accepted spelling variants; style preferences are separate from spelling. +cancelled +publically diff --git a/src/dayamlchecker/messages.py b/src/dayamlchecker/messages.py index 3a5c357..ac3942f 100644 --- a/src/dayamlchecker/messages.py +++ b/src/dayamlchecker/messages.py @@ -16,9 +16,12 @@ class FindingClass(StrEnum): ACCESSIBILITY = "accessibility" STYLE = "style" TRANSLATABILITY = "translatability" + SPELLING = "spelling" class MessageId(StrEnum): + SPELLING_POSSIBLE_TYPO = "spelling_possible_typo" + SPELLING_COMMON_LEGAL_TYPO = "spelling_common_legal_typo" YAML_DUPLICATE_KEY = "yaml_duplicate_key" YAML_DUPLICATE_BLOCK_ID = "yaml_duplicate_block_id" YAML_PARSE_ERROR = "yaml_parse_error" @@ -301,6 +304,20 @@ class MessageDefinition: MESSAGE_DEFINITIONS: dict[str, MessageDefinition] = { + MessageId.SPELLING_POSSIBLE_TYPO: MessageDefinition( + code="SP701", + severity=Severity.WARNING, + finding_class=FindingClass.SPELLING, + summary="Possible spelling mistake", + template='possible spelling mistake "{word}" in {location}: {snippet}', + ), + MessageId.SPELLING_COMMON_LEGAL_TYPO: MessageDefinition( + code="SP702", + severity=Severity.WARNING, + finding_class=FindingClass.SPELLING, + summary="Common legal spelling mistake", + template='"{word}" is usually spelled "{suggestion}" in legal text ({location}): {snippet}', + ), MessageId.YAML_DUPLICATE_KEY: MessageDefinition( code="EG101", severity=Severity.ERROR, @@ -1707,6 +1724,8 @@ class Finding: # stays the real path so tools can resolve it, but ``line_number`` counts # rendered lines, which need not correspond to lines of that file. rendered_jinja: bool = False + # Keep the rule code stable for suppressions when its configured level changes. + severity_override: Severity | None = None @property def definition(self) -> MessageDefinition: @@ -1718,7 +1737,7 @@ def code(self) -> str: @property def severity(self) -> Severity: - return self.definition.severity + return self.severity_override or self.definition.severity @property def finding_class(self) -> FindingClass: diff --git a/src/dayamlchecker/spelling.py b/src/dayamlchecker/spelling.py new file mode 100644 index 0000000..cb37284 --- /dev/null +++ b/src/dayamlchecker/spelling.py @@ -0,0 +1,693 @@ +"""Conservative spelling checks for visible interview text. + +Checking is local. English, Russian and Swedish dictionaries come with spylls; +the Spanish dictionary is downloaded once on first use and cached. +""" + +from __future__ import annotations + +from dataclasses import dataclass, replace +from functools import lru_cache +import hashlib +from html.parser import HTMLParser +import itertools +import importlib.resources +import inspect +import os +from pathlib import Path +import re +from typing import Iterable +import unicodedata + +import requests +from spylls.hunspell import Dictionary # type: ignore[import-untyped] + +from dayamlchecker.accessibility import ( + FIELD_NON_LABEL_KEYS, + _iter_fields, + _extract_field_label, + _extract_field_variable, +) +from dayamlchecker.messages import Finding, MessageId, Severity, make_finding +from dayamlchecker.style import ( + ParsedInterviewDocument, + TextEntry, + _user_facing_text_entries, + _is_object_choice, +) + +_WORDS = re.compile(r"(? None: + object.__setattr__(self, "severity", Severity(self.severity)) + object.__setattr__( + self, + "allowed_words", + frozenset(filter(None, map(normalize_word, self.allowed_words))), + ) + if isinstance(self.languages, str) or not self.languages: + raise ValueError( + "spellcheck languages must be a nonempty sequence of codes" + ) + languages = tuple(dict.fromkeys(normalize_language(v) for v in self.languages)) + dictionaries = tuple( + (normalize_language(code), str(Path(path).expanduser().resolve())) + for code, path in self.dictionaries + ) + if len({code for code, _ in dictionaries}) != len(dictionaries): + raise ValueError("duplicate spellcheck dictionary language") + for code, path in dictionaries: + for suffix in (".aff", ".dic"): + if not Path(path + suffix).is_file(): + raise ValueError( + f"cannot read spellcheck dictionary {code}: {path + suffix}" + ) + try: + _dictionary(path) + except (OSError, UnicodeError, ValueError, re.error) as exc: + raise ValueError( + f"cannot load spellcheck dictionary {code}: {exc}" + ) from exc + for language in languages: + if language in dict(dictionaries): + continue + if language in _REMOTE_DICTIONARIES: + # Download now so a failure is a configuration error, not a + # crash partway through checking. + _remote_dictionary_prefix(language) + elif language not in _SPYLLS_DICTIONARIES: + raise ValueError( + f"unsupported spellcheck language {language!r}; " + f"built in: {', '.join(_BUILTIN_LANGUAGES)}; " + "supply a custom Hunspell dictionary for other languages or regions" + ) + object.__setattr__(self, "languages", languages) + object.__setattr__(self, "dictionaries", dictionaries) + + @property + def dictionary_sources(self) -> tuple[str, ...]: + overrides = dict(self.dictionaries) + return tuple(overrides.get(language, language) for language in self.languages) + + +@dataclass(frozen=True) +class SpellingTextEntry(TextEntry): + geographic_choices: bool = False + + +def normalize_word(word: str) -> str: + """Normalize a word as spelling_text() normalizes interview prose.""" + return unicodedata.normalize("NFC", word.strip().replace("’", "'")).lower() + + +def _base_language(language: str) -> str: + return language.strip().lower().replace("_", "-").split("-", 1)[0] + + +def normalize_language(language: str) -> str: + """Normalize tags while preserving dialects that require custom dictionaries.""" + code = language.strip().lower().replace("_", "-") + if not re.fullmatch(r"[a-z]{2,3}(?:-[a-z0-9]{2,8})*", code): + raise ValueError(f"invalid spellcheck language code {language!r}") + return {"en-us": "en", "es-us": "es", "sv-se": "sv", "ru-ru": "ru"}.get(code, code) + + +@dataclass(frozen=True) +class _RemoteDictionary: + name: str + base_url: str + sha256: dict[str, str] # by file suffix + + +# Dictionaries shipped with spylls. +_SPYLLS_DICTIONARIES = {"en": "en_US", "ru": "ru", "sv": "sv_SE"} +# Dictionaries downloaded on first use rather than redistributed. The pinned +# upstream commit and hashes keep every run on identical files. +_REMOTE_DICTIONARIES = { + "es": _RemoteDictionary( + # RLA-ES Spanish (US), as distributed by LibreOffice. + name="es_US", + base_url="https://raw.githubusercontent.com/LibreOffice/dictionaries/" + "762abe74008b94b2ff06db6f4024b59a8254c467/es/", + sha256={ + ".aff": "674c5a4b4d39fd3b4452f045a4e6e0649db4a2ce23f5903df8c311e21f1a757c", + ".dic": "d46932a5c0ec3881fdf265333df4de45a73de51741baedcb5bb54d37b03979c8", + }, + ), +} +_BUILTIN_LANGUAGES = sorted(_SPYLLS_DICTIONARIES.keys() | _REMOTE_DICTIONARIES) +_SPYLLS_DATA = Path(inspect.getfile(Dictionary)).parent / "data" + + +def dictionary_cache_dir() -> Path: + """Where downloaded dictionaries are kept: $DAYAMLCHECKER_CACHE_DIR, else + the platform's user cache directory.""" + if configured := os.environ.get("DAYAMLCHECKER_CACHE_DIR"): + return Path(configured).expanduser() + if os.name == "nt" and os.environ.get("LOCALAPPDATA"): + return Path(os.environ["LOCALAPPDATA"]) / "dayamlchecker" / "cache" + return Path(os.environ.get("XDG_CACHE_HOME") or Path.home() / ".cache") / ( + "dayamlchecker" + ) + + +def _remote_dictionary_prefix(language: str) -> str: + """Return the cached dictionary prefix, downloading missing or bad files.""" + remote = _REMOTE_DICTIONARIES[language] + # Key the directory on the content, so a pin update never reuses old files. + directory = ( + dictionary_cache_dir() / "dictionaries" / language / remote.sha256[".dic"][:16] + ) + offline_hint = ( + f"to work offline, supply --spellcheck-dictionary {language}=PATH " + "with a local Hunspell dictionary" + ) + for suffix, expected in remote.sha256.items(): + path = directory / (remote.name + suffix) + try: + if path.is_file() and _sha256(path.read_bytes()) == expected: + continue + except OSError: + pass + url = remote.base_url + remote.name + suffix + try: + response = requests.get(url, timeout=60) + response.raise_for_status() + except requests.RequestException as exc: + raise ValueError( + f"cannot download the {language} spellcheck dictionary from {url}: " + f"{exc}; {offline_hint}" + ) from exc + if _sha256(response.content) != expected: + raise ValueError( + f"the downloaded {language} spellcheck dictionary {url} does not " + f"match its expected SHA-256; {offline_hint}" + ) + try: + directory.mkdir(parents=True, exist_ok=True) + # Write then rename, so concurrent runs never read a partial file. + temporary = path.with_name(f".{path.name}.{os.getpid()}.tmp") + temporary.write_bytes(response.content) + os.replace(temporary, path) + except OSError as exc: + raise ValueError( + f"cannot cache the {language} spellcheck dictionary in " + f"{directory}: {exc}; set DAYAMLCHECKER_CACHE_DIR to a writable " + "directory" + ) from exc + return str(directory / remote.name) + + +def _sha256(data: bytes) -> str: + return hashlib.sha256(data).hexdigest() + + +@lru_cache(maxsize=16) +def _dictionary(source: str) -> Dictionary: + if source in _REMOTE_DICTIONARIES: + return Dictionary.from_files(_remote_dictionary_prefix(source)) + if source in _SPYLLS_DICTIONARIES: + # Resolve spylls' bundled copy explicitly: from_files("en_US") prefers + # an en_US.aff/.dic in the current directory when one exists. + name = _SPYLLS_DICTIONARIES[source] + return Dictionary.from_files( + str(_SPYLLS_DATA / Dictionary.DISTRIBUTED[name] / name) + ) + return Dictionary.from_files(source) + + +@lru_cache(maxsize=32768) +def dictionary_accepts(word: str, sources: tuple[str, ...] = ("en",)) -> bool: + return any(bool(_dictionary(source).lookup(word)) for source in sources) + + +def read_wordlist(text: str) -> frozenset[str]: + """Parse one normalized word per line, ignoring blank and # comment lines.""" + return frozenset( + normalize_word(line) + for line in text.splitlines() + if line.strip() and not line.lstrip().startswith("#") + ) + + +def options_from_cli( + *, + languages: Iterable[str], + dictionary_specs: Iterable[str], + wordlists: Iterable[Path], + ignore_words: Iterable[str] = (), + severity: Severity | str = Severity.WARNING, +) -> SpellcheckOptions: + """Build options from LANG=PATH dictionaries and wordlist files. + + Raise ValueError with a user-facing message for invalid input. + """ + allowed = set(ignore_words) + for wordlist in wordlists: + try: + allowed |= read_wordlist(wordlist.read_text(encoding="utf-8")) + except (OSError, UnicodeError) as exc: + raise ValueError( + f"cannot read spellcheck wordlist {wordlist}: {exc}" + ) from exc + dictionaries: list[tuple[str, str]] = [] + for specification in dictionary_specs: + code, separator, path = specification.partition("=") + if not separator or not path: + raise ValueError( + "spellcheck dictionary requires LANG=PATH (without .aff/.dic)" + ) + dictionaries.append((code, path)) + return SpellcheckOptions( + allowed_words=frozenset(allowed), + languages=tuple(languages) + or tuple(code for code, _ in dictionaries) + or ("en",), + dictionaries=tuple(dictionaries), + severity=Severity(severity), + ) + + +@lru_cache(maxsize=1) +def _domain_words() -> frozenset[str]: + return read_wordlist( + importlib.resources.files("dayamlchecker") + .joinpath("data/spelling_words.txt") + .read_text(encoding="utf-8") + ) + + +class _VisibleHTML(HTMLParser): + def __init__(self) -> None: + super().__init__(convert_charrefs=True) + self.parts: list[str] = [] + self.hidden: list[str] = [] + + def handle_starttag(self, tag: str, attrs: list[tuple[str, str | None]]) -> None: + if tag in {"script", "style", "code", "pre"}: + self.hidden.append(tag) + self.parts.append(" ") + if tag == "img" and not self.hidden: + self.parts.extend(value for name, value in attrs if name == "alt" and value) + + def handle_endtag(self, tag: str) -> None: + if tag in self.hidden: + self.hidden = self.hidden[: self.hidden.index(tag)] + self.parts.append(" ") + + def handle_data(self, data: str) -> None: + if not self.hidden: + self.parts.append(data) + + +def _without_expressions(text: str) -> str: + """Remove balanced Mako expressions, including nested dicts and quotes.""" + parts: list[str] = [] + pos = 0 + while (start := text.find("${", pos)) != -1: + parts.append(text[pos:start]) + end, depth, quote = start + 2, 1, "" + while end < len(text) and depth: + char = text[end] + if quote: + if char == "\\": + end += 2 + continue + if char == quote: + quote = "" + elif char in "\"'": + quote = char + elif char == "{": + depth += 1 + elif char == "}": + depth -= 1 + end += 1 + parts.append(" ") + pos = end + parts.append(text[pos:]) + return "".join(parts) + + +def spelling_text(text: str) -> str: + """Keep prose while discarding template, Markdown, HTML and URL syntax.""" + text = re.sub(r"<%[\s\S]*?%>", " ", text) + text = _without_expressions(text) + text = re.sub(r"(?m)^[ \t]*%.*$", " ", text) + text = re.sub(r"(?ms)^[ \t]*(`{3,}|~{3,}).*?^[ \t]*\1[^\n]*", " ", text) + text = re.sub(r"`+[^`]*`+", " ", text) + text = _LINK.sub(r"\1", text) + text = re.sub(r":[A-Za-z][\w-]*:", " ", text) # Docassemble icons + text = re.sub(r"\b([^\W\d_]+)\([a-z]{1,4}\)", r"\1", text) # child(ren) + text = re.sub(r"\[(?:FILE|EMOJI|[A-Z][A-Z_ ]*)\b[^\]]*\]", " ", text) + text = _TECHNICAL.sub(" ", text) + parser = _VisibleHTML() + parser.feed(text) + parser.close() + return unicodedata.normalize("NFC", "".join(parser.parts).replace("’", "'")) + + +def _geographic_choice_field(field: dict) -> bool: + # Use the interview's schema to identify names, rather than maintaining a + # county/court/city dictionary specific to one jurisdiction. Only choice + # labels are affected: questions, instructions and help remain checked. + context = ( + " ".join( + ( + _extract_field_label(field), + _extract_field_variable(field), + str(field.get("question", "")), + str(field.get("help", "")), + ) + ) + .replace("_", " ") + .replace(".", " ") + ) + if _GEOGRAPHIC_FIELD.search(context): + return True + return bool( + re.search(r"\bhearing\b", context, re.IGNORECASE) + and re.search(r"\blocation\b", context, re.IGNORECASE) + ) + + +def _lowercase_keys( + docs: Iterable[ParsedInterviewDocument], +) -> list[ParsedInterviewDocument]: + return [ + replace(doc, doc={str(key).lower(): value for key, value in doc.doc.items()}) + for doc in docs + ] + + +def _declared_author_names(docs: list[ParsedInterviewDocument]) -> frozenset[str]: + names: set[str] = set() + for parsed in docs: + metadata = parsed.doc.get("metadata") + if not isinstance(metadata, dict): + continue + authors = metadata.get("authors", []) + if not isinstance(authors, list): + continue + for author in authors: + name = author.get("name", "") if isinstance(author, dict) else author + if isinstance(name, str): + names.update(m.group().lower() for m in _WORDS.finditer(name)) + return frozenset(names) + + +def _without_credit_names(text: str) -> str: + """Exclude personal-name spans in explicitly marked attribution sections. + + Retain prose in those sections; a misspelling in a contributor description + is still useful to flag. Heading scope ends at the next peer/parent heading. + """ + credit_level: int | None = None + lines: list[str] = [] + for line in text.splitlines(keepends=True): + heading = _HEADING.match(line) + if heading: + level = len(line.lstrip().split()[0]) + if credit_level is not None and level <= credit_level: + credit_level = None + if _CREDIT_HEADING.search(heading[1]): + credit_level = level + if credit_level is not None or re.search( + r"(?:©|©|\bcopyright\b)", line, re.IGNORECASE + ): + line = _PERSON_NAME.sub(lambda m: " " * len(m.group()), line) + lines.append(line) + return "".join(lines) + + +def spelling_entries( + docs: Iterable[ParsedInterviewDocument], + options: SpellcheckOptions | None = None, +) -> list[SpellingTextEntry]: + """Select visible text in enabled languages, excluding code/stored values.""" + options = options or SpellcheckOptions() + docs = _lowercase_keys(docs) + default_language = next( + (str(p.doc["default language"]) for p in docs if p.doc.get("default language")), + "", + ) + enabled_bases = {_base_language(language) for language in options.languages} + entries: list[SpellingTextEntry] = [] + for parsed in docs: + declared = parsed.doc.get("language") or default_language + if declared: + # A declared translation outside the enabled dictionaries must not + # be checked against an unrelated language. Configured dictionaries + # remain a union within eligible blocks, including bilingual prose. + if _base_language(str(declared)) not in enabled_bases: + continue + fields = list(_iter_fields(parsed.doc)) + for entry in _user_facing_text_entries([parsed]): + if ( + entry.location.endswith(".first_key") + and entry.text in FIELD_NON_LABEL_KEYS + ): + continue + field_choice = re.fullmatch(r"fields\[(\d+)\]\.choices", entry.location) + choice_owner = ( + fields[int(field_choice[1])] + if field_choice + else parsed.doc if entry.location in _CHOICE_LOCATIONS else None + ) + if choice_owner is not None and _is_object_choice(choice_owner): + continue + geographic = choice_owner is not None and _geographic_choice_field( + choice_owner + ) + entries.append( + SpellingTextEntry( + entry.location, + entry.text, + entry.line_number, + entry.screen_id, + geographic, + ) + ) + if ( + "template" in parsed.doc + and isinstance(parsed.doc.get("content"), str) + and parsed.doc.get("content type") in {None, "text/html", "text/plain"} + ): + entries.append( + SpellingTextEntry( + "content", + parsed.doc["content"], + parsed.line_for_key("content"), + parsed.screen_id, + ) + ) + return entries + + +def _accepted( + word: str, allowed: frozenset[str], sources: tuple[str, ...], english: bool +) -> bool: + lower = word.lower() + if ( + lower in allowed + or dictionary_accepts(word, sources) + or dictionary_accepts(lower, sources) + or dictionary_accepts(lower.capitalize(), sources) + ): + return True + if ( + english + and lower.endswith("'s") + and _accepted(word[:-2], allowed, sources, english) + ): + return True + if "-" in word and all( + _accepted(part, allowed, sources, english) for part in word.split("-") + ): + return True + if english and "-" in word: + prefix, rest = word.split("-", 1) + if prefix.lower() in {"pre", "non", "un", "re", "co", "sur"} and _accepted( + rest, allowed, sources, english + ): + return True + return False + + +def _checkable(word: str) -> bool: + """Words eligible for reporting: not short, acronyms or mixed case.""" + return ( + len(word) >= 3 and not word.isupper() and not any(c.isupper() for c in word[1:]) + ) + + +@lru_cache(maxsize=4096) +def _close_dictionary_word(word: str, sources: tuple[str, ...]) -> bool: + """Rescue title-case typos without warning on every unfamiliar name. + + Require a nearby lowercase dictionary word: proper names alone are not + evidence of an error. Restrict this extra work to longer capitalized words. + """ + if len(word) < 6 or len(word) > 40: + return False + splits = [(word[:i], word[i:]) for i in range(len(word) + 1)] + letters = set("abcdefghijklmnopqrstuvwxyz") + for source in sources: + if source != "en": + letters.update((_dictionary(source).aff.TRY or "").lower()) + # Substitutions are especially prone to confusing names with ordinary + # words (e.g. Quinten / quintet), so omit them for this title-case rescue. + # Generate lazily, cheapest edits first, and stop at the first hit. Look + # candidates up directly so they do not evict entries from the word cache. + # The few duplicate candidates (repeated letters) cost less than deduping. + edits = itertools.chain( + (left + right[1:] for left, right in splits if right), + ( + left + right[1] + right[0] + right[2:] + for left, right in splits + if len(right) > 1 + ), + (left + char + right for left, right in splits for char in letters), + ) + dictionaries = [_dictionary(source) for source in sources] + return any( + dictionary.lookup(candidate) + for candidate in edits + for dictionary in dictionaries + ) + + +def find_spelling_findings( + *, + docs: Iterable[ParsedInterviewDocument], + input_file: str | None, + options: SpellcheckOptions | None = None, +) -> list[Finding]: + """Report unknown prose words once per entry; never rewrite interview text. + + Prioritize precision: skip names/acronyms/mixed case, short tokens and text + that appears to be another language. Apart from explicit legal spelling + recommendations, this does not detect real-word typos. + """ + options = options or SpellcheckOptions() + docs = _lowercase_keys(docs) + author_names = _declared_author_names(docs) + sources = options.dictionary_sources + bases = {_base_language(language) for language in options.languages} + english = "en" in bases + allowed = options.allowed_words | (_domain_words() if english else frozenset()) + findings: list[Finding] = [] + for entry in spelling_entries(docs, options): + text = spelling_text(_without_credit_names(entry.text)) + words = list(_WORDS.finditer(text)) + # Accept each reportable word once; acronyms, short and mixed-case words + # are never reported, nor counted in the other-language ratio, so + # acronym-heavy English prose is still checked. + accepted = { + m.group(): _accepted(m.group(), allowed, sources, english) + for m in words + if _checkable(m.group()) + } + checked = [accepted[m.group()] for m in words if m.group() in accepted] + if len(checked) >= 8 and sum(checked) / len(checked) < 0.6: + continue + seen: set[str] = set() + for match in words: + word, lower = match.group(), match.group().lower() + if lower in seen or lower in allowed: + continue + suggestion = _legal_spelling_suggestion(word, options.languages, allowed) + if not suggestion and ( + accepted.get(word, True) or (bases == {"en"} and not word.isascii()) + ): + continue + if ( + not suggestion + and word[0].isupper() + and ( + entry.geographic_choices + or lower in author_names + or not _close_dictionary_word(lower, sources) + ) + ): + continue + seen.add(lower) + finding = make_finding( + ( + MessageId.SPELLING_COMMON_LEGAL_TYPO + if suggestion + else MessageId.SPELLING_POSSIBLE_TYPO + ), + file_name=input_file, + line_number=entry.line_number, + word=word, + suggestion=suggestion, + location=entry.location, + screen_id=entry.screen_id, + snippet=re.sub( + r"\s+", + " ", + text[max(0, match.start() - 80) : match.end() + 120], + ).strip(), + ) + findings.append(replace(finding, severity_override=options.severity)) + return findings + + +def _legal_spelling_suggestion( + word: str, languages: tuple[str, ...], allowed: frozenset[str] +) -> str | None: + """Explicit legal corrections bypass dictionary variants and acronym filters.""" + if "-" in word: + parts = word.split("-") + replacements = [ + _legal_spelling_suggestion(part, languages, allowed) or part + for part in parts + ] + compound = "-".join(replacements) + return compound if compound != word else None + lower = word.lower() + stem = lower.removesuffix("'s") + suffix = word[len(stem) :] + if lower in allowed or stem in allowed: + return None + if stem == "hippa" and any(_base_language(v) == "en" for v in languages): + return "HIPAA" + suffix + # The preference for judgment is US English; a custom British English + # dictionary should retain its own accepted spelling variants. + if "en" not in languages: + return None + corrected = {"judgement": "judgment", "judgements": "judgments"}.get(stem) + if not corrected: + return None + if word.isupper(): + corrected = corrected.upper() + elif word[0].isupper(): + corrected = corrected.capitalize() + return corrected + suffix diff --git a/src/dayamlchecker/yaml_structure.py b/src/dayamlchecker/yaml_structure.py index d1bc471..1de7fa5 100644 --- a/src/dayamlchecker/yaml_structure.py +++ b/src/dayamlchecker/yaml_structure.py @@ -8,7 +8,7 @@ import re import sys -from typing import Any, Optional +from typing import Any, Callable, Iterator, Optional from dayamlchecker.accessibility import ( AccessibilityLintOptions, find_accessibility_findings, @@ -27,6 +27,12 @@ StyleLintOptions, find_style_findings, ) + +from dayamlchecker.spelling import ( + SpellcheckOptions, + find_spelling_findings, + options_from_cli, +) from mako.template import Template as MakoTemplate # type: ignore[import-untyped] from mako.exceptions import ( # type: ignore[import-untyped] SyntaxException, @@ -228,6 +234,8 @@ class RuntimeOptions: style_openai_api_key: str | None = None style_openai_model: str | None = None docx_accessibility_severity: Severity = Severity.WARNING + # None disables spellcheck. + spellcheck: SpellcheckOptions | None = field(default_factory=SpellcheckOptions) def docx_accessibility_options(self) -> DocxAccessibilityOptions: return DocxAccessibilityOptions(max_severity=self.docx_accessibility_severity) @@ -2097,6 +2105,21 @@ def find_errors_from_string( runtime_options: Optional[RuntimeOptions] = None, ) -> list[YAMLError]: """Preprocess opted-in Jinja templates, then run normal YAML validation.""" + return _check_with_jinja( + full_content, + input_file, + lambda content: _find_errors_from_yaml( + content, input_file, lint_mode, runtime_options + ), + ) + + +def _check_with_jinja( + full_content: str, + input_file: Optional[str], + check_yaml: Callable[[str], list[YAMLError]], +) -> list[YAMLError]: + """Render opted-in Jinja templates, then run check_yaml on the YAML.""" # Match the universal newlines that find_errors() gets from open(): a "---\r" # separator would otherwise not split, collapsing the file into one block. full_content = full_content.replace("\r\n", "\n").replace("\r", "\n") @@ -2148,13 +2171,69 @@ def find_errors_from_string( # usable, while suppression and annotation stay off the source lines. return partial_findings + [ replace(finding, rendered_jinja=True) - for finding in _find_errors_from_yaml( - full_content, input_file, lint_mode, runtime_options - ) + for finding in check_yaml(full_content) ] - return partial_findings + _find_errors_from_yaml( - full_content, input_file, lint_mode, runtime_options - ) + return partial_findings + check_yaml(full_content) + + +def _iter_yaml_documents( + full_content: str, input_file: str +) -> Iterator[tuple[int, str, Any, Optional[YAMLError]]]: + """Split and parse each YAML document. + + Yield (start line, normalized source, document, parse error finding). + """ + yaml_parser = _make_yaml_parser() + line_number = 1 + for source_code in document_match.split(full_content): + lines_in_code = sum(l == "\n" for l in source_code) + source_code = remove_trailing_dots.sub("", source_code) + source_code = fix_tabs.sub(" ", source_code) + try: + doc = _with_line_metadata(yaml_parser.load(source_code)) + except Exception as errMess: + error_line_number = line_number + if isinstance(errMess, MarkedYAMLError): + if errMess.context_mark is not None: + errMess.context_mark.line += line_number - 1 + if errMess.problem_mark is not None: + errMess.problem_mark.line += line_number - 1 + error_line_number = _yaml_error_line_number( + errMess, full_content, line_number + ) + rendered_error = str(errMess) + yield line_number, source_code, None, make_finding( + _yaml_error_message_id(rendered_error), + line_number=error_line_number, + file_name=input_file, + error=rendered_error, + ) + else: + yield line_number, source_code, doc, None + line_number += lines_in_code + + +def parse_interview_documents( + full_content: str, input_file: str = "" +) -> tuple[list[ParsedInterviewDocument], list[YAMLError]]: + """Parse the YAML documents of an interview, without running any checks.""" + docs: list[ParsedInterviewDocument] = [] + errors: list[YAMLError] = [] + for line_number, source_code, doc, error in _iter_yaml_documents( + full_content, input_file + ): + if error is not None: + errors.append(error) + elif isinstance(doc, dict): + docs.append( + ParsedInterviewDocument( + doc=doc, + source_code=source_code, + document_start_line=line_number, + index=len(docs), + ) + ) + return docs, errors def _find_errors_from_yaml( @@ -2181,48 +2260,20 @@ def _find_errors_from_yaml( for key in types_of_blocks.keys() if types_of_blocks[key].get("exclusive", True) ] - yaml_parser = _make_yaml_parser() prior_conditional_fields: list[dict[str, Any]] = [] seen_ids: dict[str, int] = {} skip_undefined = False parsed_docs: list[ParsedInterviewDocument] = [] has_yaml_parse_errors = False - line_number = 1 - for source_code in document_match.split(full_content): - lines_in_code = sum(l == "\n" for l in source_code) - source_code = remove_trailing_dots.sub("", source_code) - source_code = fix_tabs.sub(" ", source_code) - try: - doc = _with_line_metadata(yaml_parser.load(source_code)) - except Exception as errMess: - error_line_number = line_number - if isinstance(errMess, MarkedYAMLError): - if errMess.context_mark is not None: - errMess.context_mark.line += line_number - 1 - if errMess.problem_mark is not None: - errMess.problem_mark.line += line_number - 1 - error_line_number = _yaml_error_line_number( - errMess, full_content, line_number - ) - rendered_error = str(errMess) - all_errors.append( - make_finding( - _yaml_error_message_id(rendered_error), - line_number=error_line_number, - file_name=input_file, - error=rendered_error, - ) - ) + for line_number, source_code, doc, parse_error in _iter_yaml_documents( + full_content, input_file + ): + if parse_error is not None: + all_errors.append(parse_error) has_yaml_parse_errors = True - line_number += lines_in_code - continue - - if doc is None: - # Just YAML comments, that's fine - line_number += lines_in_code continue + # None is a comment-only document, which is fine. if not isinstance(doc, dict): - line_number += lines_in_code continue if lint_mode == ACCESSIBILITY_LINT_MODE: @@ -2447,7 +2498,6 @@ def _find_errors_from_yaml( ) ) - line_number += lines_in_code if not has_yaml_parse_errors: all_errors.extend( _find_interview_level_findings(parsed_docs, input_file=input_file) @@ -2468,6 +2518,14 @@ def _find_errors_from_yaml( options=style_options, ) ) + if runtime_options.spellcheck and not has_yaml_parse_errors: + all_errors.extend( + find_spelling_findings( + docs=parsed_docs, + input_file=input_file, + options=runtime_options.spellcheck, + ) + ) return _apply_dayc_suppressions(all_errors, full_content) @@ -2641,7 +2699,7 @@ def find_style_findings_from_string( lint_mode: str = DEFAULT_LINT_MODE, runtime_options: Optional[RuntimeOptions] = None, ) -> list[Finding]: - resolved_options = runtime_options or RuntimeOptions() + resolved_options = replace(runtime_options or RuntimeOptions(), spellcheck=None) if not resolved_options.style_enabled and not resolved_options.style_include_llm: resolved_options = replace(resolved_options, style_enabled=True) return [ @@ -2656,6 +2714,32 @@ def find_style_findings_from_string( ] +def find_spelling_findings_from_string( + full_content: str, + *, + input_file: str | None = None, + runtime_options: Optional[RuntimeOptions] = None, +) -> list[Finding]: + """Return only spelling findings, with normal preprocessing and suppressions.""" + options = (runtime_options or RuntimeOptions()).spellcheck or SpellcheckOptions() + file_name = input_file or "" + + def check_yaml(content: str) -> list[YAMLError]: + docs, parse_errors = parse_interview_documents(content, file_name) + if parse_errors: + return [] + return _apply_dayc_suppressions( + find_spelling_findings(docs=docs, input_file=file_name, options=options), + content, + ) + + return [ + finding + for finding in _check_with_jinja(full_content, input_file, check_yaml) + if finding.finding_class == FindingClass.SPELLING + ] + + def _collect_yaml_files( paths: list[Path], include_default_ignores: bool = True ) -> list[Path]: @@ -2856,6 +2940,44 @@ def main(argv: Optional[list[str]] = None) -> int: action="store_true", help="Enable Assembly Line style lint checks.", ) + parser.add_argument( + "--spellcheck", + action=argparse.BooleanOptionalAction, + default=True, + help="Check visible text for possible spelling mistakes (enabled by default; offline US English).", + ) + parser.add_argument( + "--spellcheck-severity", + choices=[level.value for level in Severity], + default=Severity.WARNING.value, + help="Severity for spelling findings: info, warning or error (default: warning).", + ) + parser.add_argument( + "--spellcheck-wordlist", + type=Path, + action="append", + default=[], + help="Allow words from a UTF-8 file, one per line (# comments); repeatable.", + ) + parser.add_argument( + "--spellcheck-ignore-word", + action="append", + default=[], + help="Suppress a word, case insensitive; repeatable.", + ) + parser.add_argument( + "--spellcheck-language", + action="append", + default=[], + help="Dictionary language code (default: en/US English); repeat for mixed languages.", + ) + parser.add_argument( + "--spellcheck-dictionary", + action="append", + default=[], + metavar="LANG=PATH", + help="Custom Hunspell dictionary prefix (without .aff/.dic); repeatable.", + ) parser.add_argument( "--style-require-custom-theme", action="store_true", @@ -2996,6 +3118,19 @@ def main(argv: Optional[list[str]] = None) -> int: ) args = parser.parse_args(argv) + spellcheck: SpellcheckOptions | None = None + if args.spellcheck: + try: + spellcheck = options_from_cli( + languages=args.spellcheck_language, + dictionary_specs=args.spellcheck_dictionary, + wordlists=args.spellcheck_wordlist, + ignore_words=args.spellcheck_ignore_word, + severity=args.spellcheck_severity, + ) + except ValueError as exc: + parser.error(str(exc)) + lint_mode = ACCESSIBILITY_LINT_MODE if args.wcag else DEFAULT_LINT_MODE runtime_options = RuntimeOptions( accessibility_error_on_widgets=frozenset( @@ -3010,6 +3145,7 @@ def main(argv: Optional[list[str]] = None) -> int: style_openai_api_key=args.openai_api_key, style_openai_model=args.openai_model, docx_accessibility_severity=Severity(args.docx_accessibility_severity), + spellcheck=spellcheck, ) yaml_files = _collect_yaml_files( diff --git a/tests/conftest.py b/tests/conftest.py index 439af90..e0a9257 100644 --- a/tests/conftest.py +++ b/tests/conftest.py @@ -1,8 +1,34 @@ +import os import sys from pathlib import Path +import pytest + ROOT = Path(__file__).resolve().parents[1] SRC = ROOT / "src" if str(SRC) not in sys.path: sys.path.insert(0, str(SRC)) + + +@pytest.fixture(scope="session", autouse=True) +def dictionary_cache(request): + """Keep downloaded dictionaries in pytest's cache, so local runs fetch once.""" + with pytest.MonkeyPatch.context() as patch: + patch.setenv( + "DAYAMLCHECKER_CACHE_DIR", str(request.config.cache.mkdir("dayamlchecker")) + ) + yield + + +@pytest.fixture +def requires_spanish(): + """Fetch the Spanish dictionary; skip when offline, except in CI.""" + from dayamlchecker.spelling import _remote_dictionary_prefix + + try: + _remote_dictionary_prefix("es") + except ValueError as exc: + if os.environ.get("CI"): + raise + pytest.skip(f"Spanish dictionary unavailable: {exc}") diff --git a/tests/test_spelling.py b/tests/test_spelling.py new file mode 100644 index 0000000..32b3572 --- /dev/null +++ b/tests/test_spelling.py @@ -0,0 +1,876 @@ +from pathlib import Path + +import pytest + +from dayamlchecker import ( + RuntimeOptions, + find_errors_from_string, + find_spelling_findings_from_string, +) +from dayamlchecker.messages import ( + MessageId, + Severity, + get_message_definition, + print_github_annotation, +) +from dayamlchecker.spelling import ( + SpellcheckOptions, + find_spelling_findings, + spelling_text, +) +from dayamlchecker.style import ParsedInterviewDocument +from dayamlchecker.yaml_structure import main + + +def spelling(yaml, **options): + return [ + f + for f in find_errors_from_string( + yaml, + input_file="interview.yml", + runtime_options=RuntimeOptions( + spellcheck=SpellcheckOptions( + **{k.removeprefix("spellcheck_"): v for k, v in options.items()} + ) + ), + ) + if f.message_id + in {MessageId.SPELLING_POSSIBLE_TYPO, MessageId.SPELLING_COMMON_LEGAL_TYPO} + ] + + +def test_public_string_api_returns_only_spelling_warnings(): + findings = find_spelling_findings_from_string("question: Recieve\n") + assert len(findings) == 1 and findings[0].code == "SP701" + + +def test_spelling_is_on_by_default_and_reports_locations_without_duplicates(): + yaml = "id: test\nquestion: Adress\nsubquestion: Please recieve teh benefits, not teh letters.\nfield: name\n" + assert any(f.code == "SP701" for f in find_errors_from_string(yaml)) + assert not any( + f.code == "SP701" + for f in find_errors_from_string( + yaml, runtime_options=RuntimeOptions(spellcheck=None) + ) + ) + findings = spelling(yaml) + assert [f.context["word"] for f in findings] == ["Adress", "recieve", "teh"] + assert [f.line_number for f in findings] == [2, 3, 3] + assert all( + f.file_name == "interview.yml" and f.severity == Severity.WARNING + for f in findings + ) + assert all(f.context["screen_id"] == "test" for f in findings) + + +def test_checks_display_labels_and_template_content_but_not_values_or_code(): + yaml = """question: Choose a benefit +fields: + - Street addres: misspelld_variable + help: Please recieve the letter. + choices: + - Valid label: storred_value + - label: violance + value: internnal_value +--- +template: message +content: Your lettter is ready. +--- +code: | + internnal_variable = 'misspelld text' +""" + assert {f.context["word"] for f in spelling(yaml)} == { + "addres", + "recieve", + "violance", + "lettter", + } + + +def test_object_choice_expressions_and_nonlabel_field_keys_are_skipped(): + yaml = """question: Select a person +fields: + - no label: people + datatype: object_radio + choices: + - requested_guardians if len(requested_guardians.complete_elements()) > 0 else [] + - html: Valid instructions +""" + assert spelling(yaml) == [] + + +def test_template_markup_urls_and_identifiers_do_not_leak_into_prose(): + raw = """% if user: +${ {'unrecognizabl': {'inner': 'strangeword'}} } +<% PythonUnknwon = 'unrecognizabl' %> +[FILE unrecognizabl.png, 100%] +[:fab-fa-github: Valid label](${ weird_url }) +[Helpful link](https://example.org/unrecognizabl) +Valid prose + +unrecognizabl
unrecognizabl
+`unrecognizabl` foo_bar_12 +```python +unrecognizabl +``` +https://example.org/unrecognizabl unrecognizabl@example.org filexyz.pdf +% endif +""" + docs = [ParsedInterviewDocument({"question": raw}, raw, 1, 0)] + assert find_spelling_findings(docs=docs, input_file=None) == [] + assert "Valid prose" in spelling_text(raw) + + +def test_visible_link_text_and_image_alt_are_checked(): + assert { + f.context["word"] + for f in spelling( + 'question: |\n [Recieve](https://example.org) lettter\n' + ) + } == {"Recieve", "lettter"} + + +@pytest.mark.parametrize( + "text", + [ + "Appellee's affidavit of indigency and arrearages", + "I can't pay. I won’t pay. The tenant's child's benefits.", + "Pre-filled non-English checkboxes; child(ren) and themself.", + "Cancelled judgments and wellbeing, ze/zir/zirs.", + "Ask Quinten Steenhuis about MassHealth and SNAP.", + "Your café résumé is ready.", + ], +) +def test_valid_legal_terms_inflections_names_and_markup(text): + assert spelling(f"question: {text}\n") == [] + + +def test_explicit_translations_and_long_unlabelled_foreign_text_are_skipped(): + assert spelling("language: es\nquestion: Seleccione sus beneficios\n") == [] + assert ( + spelling( + "question: Seleccione sus beneficios para completar esta solicitud de asistencia jurídica gratuita\n" + ) + == [] + ) + assert [ + f.context["word"] for f in spelling("language: en-US\nquestion: Recieve\n") + ] == ["Recieve"] + + +def test_close_capitalized_typos_are_caught_inside_sentences(): + assert { + f.context["word"] + for f in spelling( + "question: Social Security Adminstration and Plaintiff/Petitoner\n" + ) + } == {"Adminstration", "Petitoner"} + + +def test_allowlist_is_case_insensitive_and_does_not_change_other_runs(): + yaml = "question: foobarbaz\n" + assert spelling(yaml, spellcheck_allowed_words=frozenset({"FOOBARBAZ"})) == [] + assert len(spelling(yaml)) == 1 + assert SpellcheckOptions().allowed_words == frozenset() + + +@pytest.mark.parametrize("level", list(Severity)) +def test_configured_severity_applies_to_all_spelling_rules_only(level): + findings = find_errors_from_string( + "question: Recieve the judgement and HIPPA letter\nfields: []\n", + runtime_options=RuntimeOptions(spellcheck=SpellcheckOptions(severity=level)), + ) + spelling_findings = [f for f in findings if f.code in {"SP701", "SP702"}] + assert len(spelling_findings) == 3 + assert all(f.severity == level for f in spelling_findings) + assert any( + f.severity == Severity.ERROR and f.code not in {"SP701", "SP702"} + for f in findings + ) + assert ( + get_message_definition(MessageId.SPELLING_POSSIBLE_TYPO).severity + == Severity.WARNING + ) + assert spelling("question: Recieve\n")[0].severity == Severity.WARNING + + +@pytest.mark.parametrize("level", list(Severity)) +def test_spelling_github_annotations_use_configured_severity(level, capsys): + finding = spelling("question: HIPPA\n", spellcheck_severity=level)[0] + print_github_annotation(finding) + kind = "notice" if level == Severity.INFO else level.value + output = capsys.readouterr().out + assert output.startswith(f"::{kind} ") + assert "title=SP702" in output and "HIPAA" in output + + +@pytest.mark.parametrize("level,exit_code", [("info", 0), ("warning", 0), ("error", 1)]) +def test_cli_default_spellcheck_configured_levels_and_exit_codes( + tmp_path, capsys, level, exit_code +): + interview = tmp_path / "spelling.yml" + interview.write_text( + "id: test\nquestion: Recieve the judgement and HIPPA letter\nfield: name\n" + ) + assert ( + main( + [ + str(interview), + "--spellcheck-severity", + level, + "--no-url-check", + "--no-docx-accessibility", + "--no-wcag", + ] + ) + == exit_code + ) + output = capsys.readouterr().out + label = {"info": "INFO", "warning": "WARN", "error": "ERROR"}[level] + assert f"{label:<5} [SP701]" in output and f"{label:<5} [SP702]" in output + assert "judgment" in output and "HIPAA" in output + + +def test_cli_default_is_warning_and_disable_takes_precedence_over_configuration( + tmp_path, capsys +): + interview = tmp_path / "spelling.yml" + interview.write_text("id: test\nquestion: Recieve HIPPA\nfield: name\n") + common = [str(interview), "--no-url-check", "--no-docx-accessibility", "--no-wcag"] + assert main(common) == 0 + assert "WARN [SP701]" in capsys.readouterr().out + assert ( + main( + common + + [ + "--no-spellcheck", + "--spellcheck-language", + "en", + "--spellcheck-severity", + "error", + "--spellcheck-ignore-word", + "foo", + ] + ) + == 0 + ) + output = capsys.readouterr().out + assert "[SP701]" not in output and "[SP702]" not in output + + +def test_cli_info_does_not_count_as_warning_but_warning_limit_does(tmp_path, capsys): + interview = tmp_path / "spelling.yml" + interview.write_text("id: test\nquestion: Recieve\nfield: name\n") + common = [ + str(interview), + "--no-url-check", + "--no-docx-accessibility", + "--no-wcag", + "--max-warnings", + "0", + ] + assert main(common + ["--spellcheck-severity", "info"]) == 0 + capsys.readouterr() + assert main(common) == 1 + assert "WARN [SP701]" in capsys.readouterr().out + + +def test_invalid_severity_errors_in_api_and_cli(tmp_path, capsys): + with pytest.raises(ValueError): + spelling("question: Recieve\n", spellcheck_severity="ignore") + with pytest.raises(SystemExit) as exc: + main([str(tmp_path), "--spellcheck-severity", "ignore"]) + assert exc.value.code == 2 + assert "invalid choice" in capsys.readouterr().err + + +@pytest.mark.parametrize( + "word,suggestion", + [ + ("judgement", "judgment"), + ("Judgement", "Judgment"), + ("JUDGEMENT", "JUDGMENT"), + ("judgements", "judgments"), + ("judgement’s", "judgment's"), + ("judgement-proof", "judgment-proof"), + ("HIPPA", "HIPAA"), + ("hippa", "HIPAA"), + ("HIPPA's", "HIPAA's"), + ("HIPPA-compliant", "HIPAA-compliant"), + ], +) +def test_common_legal_typos_bypass_dictionary_and_case_filters(word, suggestion): + findings = find_spelling_findings_from_string(f"question: {word}\n") + assert len(findings) == 1 + finding = findings[0] + assert finding.message_id == MessageId.SPELLING_COMMON_LEGAL_TYPO + assert finding.context["suggestion"] == suggestion + assert suggestion in finding.message + + +def test_legal_typos_are_deduplicated_and_correct_spellings_are_accepted(): + assert [ + f.context["word"] + for f in spelling("question: HIPPA HIPPA judgement judgement\n") + ] == ["HIPPA", "judgement"] + assert ( + spelling("question: HIPAA Judgment judgments judgmental judgment-proof\n") == [] + ) + + +@pytest.mark.parametrize("level", list(Severity)) +def test_legal_typos_respect_suppressions_at_every_severity(level): + assert ( + spelling( + "question: HIPPA judgement\n", + spellcheck_severity=level, + spellcheck_allowed_words=frozenset({"hippa", "JUDGEMENT"}), + ) + == [] + ) + assert ( + spelling( + "question: HIPPA's judgement-proof\n", + spellcheck_severity=level, + spellcheck_allowed_words=frozenset({"hippa", "JUDGEMENT"}), + ) + == [] + ) + assert ( + spelling( + "question: HIPPA judgement # no-dayc: SP702\n", spellcheck_severity=level + ) + == [] + ) + assert ( + spelling( + "question: HIPPA judgement # no-dayc-block: spelling_common_legal_typo\n", + spellcheck_severity=level, + ) + == [] + ) + + +@pytest.mark.usefixtures("requires_spanish") +def test_legal_spelling_rules_are_scoped_to_english_and_us_variants(tmp_path): + assert ( + spelling( + "question: HIPPA judgement\n", + spellcheck_languages=("es",), + spellcheck_allowed_words=frozenset({"judgement"}), + ) + == [] + ) + prefix = tmp_path / "en_GB" + prefix.with_suffix(".aff").write_text("SET UTF-8\n") + prefix.with_suffix(".dic").write_text("1\njudgement\n") + assert ( + spelling( + "question: judgement\n", + spellcheck_languages=("en-GB",), + spellcheck_dictionaries=(("en-GB", str(prefix)),), + ) + == [] + ) + assert [ + f.context["suggestion"] + for f in spelling( + "question: HIPPA\n", + spellcheck_languages=("en-GB",), + spellcheck_dictionaries=(("en-GB", str(prefix)),), + ) + ] == ["HIPAA"] + + +@pytest.mark.parametrize( + "directive", ["# no-dayc: SP701", "# no-dayc-block: spelling_possible_typo"] +) +def test_source_suppressions(directive): + assert spelling(f"question: Recieve {directive}\n") == [] + + +def test_malformed_yaml_does_not_produce_partial_spelling_warnings(): + assert spelling("question: Recieve\n---\nfields: [\n") == [] + + +def test_cli_custom_wordlist_suppresses_only_listed_words(tmp_path, capsys): + interview = tmp_path / "test.yml" + interview.write_text("id: test\nquestion: |\n foobarbaz lettter\nfield: name\n") + wordlist = tmp_path / "words.txt" + wordlist.write_text("# Custom organization\n # comment\nFOOBARBAZ\n") + result = main( + [ + str(interview), + "--spellcheck-wordlist", + str(wordlist), + "--no-url-check", + "--no-docx-accessibility", + "--no-wcag", + ] + ) + output = capsys.readouterr().out + assert result == 0 + assert "[SP701]" in output and '"lettter"' in output + assert '"foobarbaz"' not in output + + +def test_cli_missing_wordlist_has_clear_error(tmp_path, capsys): + with pytest.raises(SystemExit) as exc: + main([str(tmp_path), "--spellcheck-wordlist", str(tmp_path / "missing")]) + assert exc.value.code == 2 + assert "cannot read spellcheck wordlist" in capsys.readouterr().err + + +@pytest.mark.usefixtures("requires_spanish") +def test_spanish_inflections_and_accented_typos(): + findings = spelling( + "question: Seleccione su dirección y beneficios para completar estas soliciitudes\n" + "subquestion: La direccióón está aquí con los niños y abogados\n", + spellcheck_languages=("es",), + ) + assert {f.context["word"] for f in findings} == {"soliciitudes", "direccióón"} + + +@pytest.mark.usefixtures("requires_spanish") +def test_mixed_languages_within_blocks_and_between_translations(): + yaml = """language: en-US +question: Please complete su solicitud and recieve beneficios +--- +language: es-MX +question: Seleccione su dirección para completar esta soliciitud +--- +language: fr +question: Cette letttre +""" + assert [ + f.context["word"] for f in spelling(yaml, spellcheck_languages=("en", "es")) + ] == ["recieve", "soliciitud"] + assert {f.context["word"] for f in spelling(yaml)} == { + "recieve", + "beneficios", + "solicitud", + } + + +@pytest.mark.usefixtures("requires_spanish") +def test_spanish_default_language_and_templates(): + yaml = """default language: es +--- +template: letter +content: La dirección está aquí y contiene una soliciitud +--- +language: en +question: Recieve +""" + assert [f.context["word"] for f in spelling(yaml)] == ["Recieve"] + assert [ + f.context["word"] for f in spelling(yaml, spellcheck_languages=("es",)) + ] == ["soliciitud"] + + +@pytest.mark.usefixtures("requires_spanish") +def test_spanish_capitalized_typo_unicode_and_custom_suppression(): + assert [ + f.context["word"] + for f in spelling("question: Soliciitud\n", spellcheck_languages=("es",)) + ] == ["Soliciitud"] + yaml = "question: La direccio\u0301n está aquí con direccióón\n" + assert [ + f.context["word"] for f in spelling(yaml, spellcheck_languages=("es",)) + ] == ["direccióón"] + assert ( + spelling( + yaml, + spellcheck_languages=("es",), + spellcheck_allowed_words=frozenset({"DIRECCIÓÓN"}), + ) + == [] + ) + + +@pytest.mark.usefixtures("requires_spanish") +def test_language_selection_and_suppressions_do_not_leak_between_calls(): + yaml = "question: beneficios soliciitud\n" + assert [ + f.context["word"] for f in spelling(yaml, spellcheck_languages=("en", "es")) + ] == ["soliciitud"] + assert ( + spelling( + yaml, + spellcheck_languages=("es",), + spellcheck_allowed_words=frozenset({"soliciitud"}), + ) + == [] + ) + assert {f.context["word"] for f in spelling(yaml)} == {"beneficios", "soliciitud"} + + +@pytest.mark.parametrize("languages", [("xx",), (), "es", ("en-GB",)]) +def test_invalid_or_unavailable_languages_have_clear_api_errors(languages): + with pytest.raises(ValueError, match="spellcheck language"): + spelling("question: Recieve\n", spellcheck_languages=languages) + + +@pytest.mark.usefixtures("requires_spanish") +def test_language_aliases_and_duplicates(): + assert SpellcheckOptions(languages=("EN_us", "en", "es_US")).languages == ( + "en", + "es", + ) + + +@pytest.mark.usefixtures("requires_spanish") +def test_cli_mixed_languages_and_inline_suppressions(tmp_path, capsys): + interview = tmp_path / "mixed.yml" + interview.write_text( + "id: test\nquestion: Please complete su solicitud and recieve beneficios\nsubquestion: lettter\nfield: name\n" + ) + assert ( + main( + [ + str(interview), + "--spellcheck-language", + "en", + "--spellcheck-language", + "es", + "--spellcheck-ignore-word", + "RECIEVE", + "--no-url-check", + "--no-docx-accessibility", + "--no-wcag", + ] + ) + == 0 + ) + output = capsys.readouterr().out + assert "[SP701]" in output and '"lettter"' in output + assert '"recieve"' not in output and '"beneficios"' not in output + + +def test_cli_unsupported_language(tmp_path, capsys): + with pytest.raises(SystemExit) as exc: + main([str(tmp_path), "--spellcheck-language", "xx"]) + assert exc.value.code == 2 + assert "unsupported spellcheck language" in capsys.readouterr().err + + +def test_custom_hunspell_dictionary_in_api_and_cli(tmp_path, capsys): + prefix = tmp_path / "custom" + prefix.with_suffix(".aff").write_text("SET UTF-8\nSFX S Y 1\nSFX S 0 s .\n") + prefix.with_suffix(".dic").write_text("1\nfoobarbaz/S\n") + options = { + "spellcheck_languages": ("fr",), + "spellcheck_dictionaries": (("fr", str(prefix)),), + } + assert spelling("question: foobarbaz foobarbazs\n", **options) == [] + assert [ + f.context["word"] for f in spelling("question: foobarbazz\n", **options) + ] == ["foobarbazz"] + interview = tmp_path / "custom.yml" + interview.write_text( + "id: test\nquestion: foobarbaz foobarbazs foobarbazz\nfield: name\n" + ) + assert ( + main( + [ + str(interview), + "--spellcheck-dictionary", + f"fr={prefix}", + "--no-url-check", + "--no-docx-accessibility", + "--no-wcag", + ] + ) + == 0 + ) + output = capsys.readouterr().out + assert '"foobarbazz"' in output and '"foobarbazs"' not in output + + +@pytest.mark.parametrize("specification", ["fr", "fr=", "fr=/missing/dictionary"]) +def test_cli_invalid_custom_dictionary(tmp_path, capsys, specification): + with pytest.raises(SystemExit) as exc: + main([str(tmp_path), "--spellcheck-dictionary", specification]) + assert exc.value.code == 2 + assert "dictionary" in capsys.readouterr().err + + +def test_jinja_preprocessing_preserves_spelling_warning_marker(): + findings = spelling( + "# use jinja\nid: test\nquestion: {{ 'Recieve' }}\nfield: name\n" + ) + assert len(findings) == 1 and findings[0].rendered_jinja + + +def test_metadata_is_not_spellchecked(): + assert spelling("""metadata: + title: Recieve a lettter + description: misspelld descriptions + authors: + - name: Urbana Zafri + organization: misspelld organization +--- +question: Valid question +""") == [] + + +def test_author_names_are_supported_by_declarations_and_scoped_to_proper_names(): + assert spelling("""metadata: + authors: + - name: Urbana Zafri +--- +question: Urbana wrote this lettter. +""")[0].context["word"] == "lettter" + assert [f.context["word"] for f in spelling("""metadata: + authors: + - name: Urbana Zafri +--- +question: urbana +""")] == ["urbana"] + + +@pytest.mark.parametrize( + "field", + [ + "County", + "Parish", + "Borough", + "City", + "Province", + "Trial court division", + "Hearing location preference", + ], +) +def test_geographic_choices_do_not_require_a_jurisdiction_wordlist(field): + # Urbana resembles 'urban', so capitalization alone does not suppress it. + # The schema provides the evidence that these choices are place names. + yaml = f"""question: Choose a location +fields: + - {field}: selected_value + choices: + - Urbana + - Ouachita + - Cuyahoga + - District Court of Nacogdoches + help: Please recieve your notice. +""" + assert [f.context["word"] for f in spelling(yaml)] == ["recieve"] + + +def test_geographic_context_comes_from_the_variable_even_with_no_label(): + assert spelling("""question: Choose a venue +field: trial_court +choices: + - Urbana District Court +""") == [] + + +def test_only_geographic_choice_labels_receive_geographic_suppression(): + yaml = """question: County information +fields: + - County: county + choices: + - Urbana + - I did not recieve a notice +subquestion: Urbana is ready. +""" + assert {f.context["word"] for f in spelling(yaml)} == {"recieve", "Urbana"} + + +def test_credit_section_excludes_names_and_keeps_prose_typos(): + yaml = """question: About this interview +subquestion: | + ### Contributors + 1. Mariah Jennings-Rampsi + 2. Urbana Zafri + They wrote this lettter. + ### Instructions + Please recieve the notice. +""" + assert [f.context["word"] for f in spelling(yaml)] == ["lettter", "recieve"] + + +def test_copyright_attributions_exclude_names_and_keep_prose_typos(): + assert [f.context["word"] for f in spelling("""question: About this interview +subquestion: Copyright 2026 Urbana Zafri. Please recieve the notice. +""")] == ["recieve"] + + +def test_unknown_capitalized_names_are_not_flagged_based_on_sentence_position(): + assert ( + spelling( + "question: Acura\nsubquestion: Pocketalker is ready. Safelink can help.\n" + ) + == [] + ) + assert [f.context["word"] for f in spelling("question: Recieve the lettter\n")] == [ + "Recieve", + "lettter", + ] + + +def test_brand_name_typos_are_a_documented_limit_of_the_english_dictionary(): + assert spelling("question: Pick a car\nchoices:\n - Izuzu\n - Volkswagon\n") == [] + + +@pytest.mark.parametrize( + "text", ["Plese recieve this Q&A", "Plese recieve AT&T", "Plese recieve