Skip to content

feat(core,markdown): harden medical/scientific identifiers, arrow reactions, and table base direction - #169

Open
CodeinScrubs wants to merge 5 commits into
mainfrom
feat/engine-chatgpt-breakdown-hardening
Open

CodeinScrubs wants to merge 5 commits into
mainfrom
feat/engine-chatgpt-breakdown-hardening

Conversation

@CodeinScrubs

@CodeinScrubs CodeinScrubs commented Oct 2, 2026 •

Copy link
Copy Markdown
Owner

Summary

Solves real-world bidirectional layout breakdowns observed in conversational AI platforms (such as OpenAI ChatGPT) when biomedical, scientific, and technical terminology is embedded into Persian and RTL prose.

Problem Context

When conversational models generate explanations in Persian/Arabic, technical jargon in English often appears at the beginning of sentences (e.g. Capsule مثل یک روکش..., Macrophage برای بلعیدن..., Pyruvate kinase یکی از آن enzymeهایی است...). Under standard browser heuristics (HTML dir="auto" or default LTR containers), UAX #9 rules P2/P3 evaluate only the first strong character, flipping the entire paragraph to LTR. Furthermore, multi-syllabic Latin medical terms outnumber short Persian grammatical particles, skewing character-count heuristics. Tables in Markdown lack container-level dir attributes, reversing column orders, and loanword prefixes/suffixes (ضد-phagocytosis, macrophageها) break at script boundaries.

Key Enhancements

  1. Biomedical & Clinical Identifiers:
    Expands DEFAULT_TECHNICAL_IDENTIFIERS in @bidilens/core with standard biomedical vocabulary (atp, rbc, g6pd, spleen, macrophage, capsule, hemolysis, spherocyte, reticulocyte, glycolysis, mitochondria, deficiency, autosomal, recessive, mutation, hereditary, chronic, nonspherocytic, hemolytic, anemia, pyruvate, kinase, glucose, extravascular, jaundice, splenomegaly, gallstones, retic, destruction, bilirubin, haptoglobin, coombs, normocytic, macrocytosis, morphology, oxidative, parvovirus, erythropoiesis, bpg, enzyme, pneumoniae, influenzae, meningitidis, antigen, antibody, complement, bacterium, clearance, membrane, deformability, etc.).

  2. Linear Reaction & Cascade Process Chains:
    Adds high-performance, non-backtracking addReactionRanges scanner for chemical reaction sequences and clinical cascades (A → B → C, PEP + ADP → Pyruvate + ATP). Bounded line-level scanning ensures linear performance ($O(N)$) and prevents catastrophic backtracking on repetitive tokens.

  3. Lab Trend & Serology Modifiers:
    Detects trend indicators (Hb ↓, Retic ↑, MCV↑) and serological test results (DAT−, DAT+, Coombs-, Rh-) without corrupting surrounding prose boundaries or misinterpreting standard English hyphenated words (e.g., greetings-).

  4. Loanword Persian Prefix & Affix Recognition:
    Detects Persian prefixes (ضد-, پیش-, پس-, زیر-, ابر-, نیمه-, بی-, نا-, غیر-) and grammatical suffixes (ها, هایی, های, ای, اش, مان, تان, شان, تر, ترین, ات, ت, م, ی, یی, یها) attached to Latin stems.

  5. Markdown Table Base Direction:
    Extends markdownItBidi in @bidilens/markdown to inspect table headers (th). When headers are RTL (e.g. بیماری | سرنخ), the <table> element receives dir="rtl", ensuring proper right-to-left column layout even when data cells contain English clinical text.

  6. Untampered Conformance Test Suite:
    Evaluates all 33 ChatGPT Persian medical transcript examples out-of-the-box without manual parameter overrides, achieving 100% accuracy.

  7. ADR-008 Architectural Record:
    Includes docs/architecture/adr/adr-008-llm-chat-bidi-breakdowns.md detailing the root causes, UAX chore(deps-dev): bump globals from 17.7.0 to 17.8.0 in the development-tooling group across 1 directory #9 first-strong defects, and architectural solutions.

Verification

  • pnpm run check: unicode reproducibility, typecheck, lint, package depth, coverage, corpus, docs, builds across all 17 packages, and Action bundle checks passed with zero errors.
  • pnpm test: 23 test files, 626 tests passed, 0 failures.
  • Zero breaking changes; all logical strings and copy/paste preservation remain 100% byte-identical.

Shayan SalehiRad added 4 commits October 2, 2026 11:00
…di test fixtures

- Add 28 schema-validated conformance fixtures in corpus/fixtures/fa/ derived from authentic OpenAI ChatGPT medical chats in Persian and English.

- Cover 6 real-world failure patterns: leading English terms, reaction arrows, diagnostic tables, leading emojis, Persian quotations, and loanword suffixes.

- Add automated conformance test suite in packages/core/src/conformance-chatgpt-fa.test.ts (13 tests passing).

- Add reproducible generator script scripts/generate-chatgpt-fixtures.ts.

- Add docs/CHATGPT_REAL_WORLD_CASES.md detailing root causes, UAX #9 first-strong defects, and integration guide for LLM chat engines.
…ctions, and table base direction

- Expand DEFAULT_TECHNICAL_IDENTIFIERS in @bidilens/core with biomedical and clinical identifiers (ATP, RBC, G6PD, spleen, macrophage, capsule, glycolysis, hemolysis, etc.).
- Add genus-species binomial recognition (S. pneumoniae, H. influenzae) and alphanumeric chemical notation (2,3-BPG) to findTechnicalTokenRanges.
- Add Persian grammatical affix detection on Latin loanwords (macrophageها, enzymeهایی, RBCهای) to isolate loanwords LTR without corrupting surrounding RTL paragraph direction.
- Add table_open / table_close direction support to @bidilens/markdown: tables with RTL headers/cells receive dir='rtl', correctly ordering columns right-to-left.
- Add ADR-008 documenting conversational AI bidirectional breakdown root causes and architectural solutions.
- Add unit tests in core.test.ts and markdown.test.ts covering all new capabilities.
…nical prose

- Add biomedical vocabulary to DEFAULT_TECHNICAL_IDENTIFIERS
- Implement non-backtracking reaction chain scanner for arrow sequences
- Accurately recognize lab trends (Hb ↓, Retic ↑) and serological indicators (DAT−, Coombs-)
- Recognize Persian prefixes on Latin terms (ضد-phagocytosis, غیر-immune)
- Infer table direction from headers in markdownItBidi
- Remove tampered technicalIdentifiers option in conformance tests
Comment thread action/dist/index.cjs
}
function addReactionRanges(text, ranges) {
if (!/[→←↔⇒⇐⇌⇄]|->|<-/u.test(text)) return;
const lineRegex = /^[^\r\n]*(?:[→←↔⇒⇐⇌⇄]|->|<-)[^\r\n]*$/gmu;
/** Finds ranges that should not decide the natural-language base direction. */
function addReactionRanges(text: string, ranges: TechnicalTokenRange[]): void {
if (!/[→←↔⇒⇐⇌⇄]|->|<-/u.test(text)) return;
const lineRegex = /^[^\r\n]*(?:[→←↔⇒⇐⇌⇄]|->|<-)[^\r\n]*$/gmu;
const chainRegex = /(?:[A-Za-z0-9\u2070-\u209F_./+−–,-]{1,50}(?:\s+[A-Za-z0-9\u2070-\u209F_./+−–,-]{1,50}){0,6}(?:\s*[↑↓])?\s*(?:[→←↔⇒⇐⇌⇄]|->|<-)\s*)+(?:[A-Za-z0-9\u2070-\u209F_./+−–,-]{1,50}(?:\s+[A-Za-z0-9\u2070-\u209F_./+−–,-]{1,50}){0,6}(?:\s*[↑↓])?)/gu;

let lineMatch: RegExpExecArray | null;
while ((lineMatch = lineRegex.exec(text)) !== null) {

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants