Repository navigation
feat(core,markdown): harden medical/scientific identifiers, arrow reactions, and table base direction - #169
Open
CodeinScrubs wants to merge 5 commits into
Open
CodeinScrubs wants to merge 5 commits into
CodeinScrubs wants to merge 5 commits into
Conversation
added 4 commits
October 2, 2026 11:00
…di test fixtures - Add 28 schema-validated conformance fixtures in corpus/fixtures/fa/ derived from authentic OpenAI ChatGPT medical chats in Persian and English. - Cover 6 real-world failure patterns: leading English terms, reaction arrows, diagnostic tables, leading emojis, Persian quotations, and loanword suffixes. - Add automated conformance test suite in packages/core/src/conformance-chatgpt-fa.test.ts (13 tests passing). - Add reproducible generator script scripts/generate-chatgpt-fixtures.ts. - Add docs/CHATGPT_REAL_WORLD_CASES.md detailing root causes, UAX #9 first-strong defects, and integration guide for LLM chat engines.
…ctions, and table base direction - Expand DEFAULT_TECHNICAL_IDENTIFIERS in @bidilens/core with biomedical and clinical identifiers (ATP, RBC, G6PD, spleen, macrophage, capsule, glycolysis, hemolysis, etc.). - Add genus-species binomial recognition (S. pneumoniae, H. influenzae) and alphanumeric chemical notation (2,3-BPG) to findTechnicalTokenRanges. - Add Persian grammatical affix detection on Latin loanwords (macrophageها, enzymeهایی, RBCهای) to isolate loanwords LTR without corrupting surrounding RTL paragraph direction. - Add table_open / table_close direction support to @bidilens/markdown: tables with RTL headers/cells receive dir='rtl', correctly ordering columns right-to-left. - Add ADR-008 documenting conversational AI bidirectional breakdown root causes and architectural solutions. - Add unit tests in core.test.ts and markdown.test.ts covering all new capabilities.
…-chatgpt-breakdown-hardening
…nical prose - Add biomedical vocabulary to DEFAULT_TECHNICAL_IDENTIFIERS - Implement non-backtracking reaction chain scanner for arrow sequences - Accurately recognize lab trends (Hb ↓, Retic ↑) and serological indicators (DAT−, Coombs-) - Recognize Persian prefixes on Latin terms (ضد-phagocytosis, غیر-immune) - Infer table direction from headers in markdownItBidi - Remove tampered technicalIdentifiers option in conformance tests
| } | ||
| function addReactionRanges(text, ranges) { | ||
| if (!/[→←↔⇒⇐⇌⇄]|->|<-/u.test(text)) return; | ||
| const lineRegex = /^[^\r\n]*(?:[→←↔⇒⇐⇌⇄]|->|<-)[^\r\n]*$/gmu; |
| /** Finds ranges that should not decide the natural-language base direction. */ | ||
| function addReactionRanges(text: string, ranges: TechnicalTokenRange[]): void { | ||
| if (!/[→←↔⇒⇐⇌⇄]|->|<-/u.test(text)) return; | ||
| const lineRegex = /^[^\r\n]*(?:[→←↔⇒⇐⇌⇄]|->|<-)[^\r\n]*$/gmu; |
| const chainRegex = /(?:[A-Za-z0-9\u2070-\u209F_./+−–,-]{1,50}(?:\s+[A-Za-z0-9\u2070-\u209F_./+−–,-]{1,50}){0,6}(?:\s*[↑↓])?\s*(?:[→←↔⇒⇐⇌⇄]|->|<-)\s*)+(?:[A-Za-z0-9\u2070-\u209F_./+−–,-]{1,50}(?:\s+[A-Za-z0-9\u2070-\u209F_./+−–,-]{1,50}){0,6}(?:\s*[↑↓])?)/gu; | ||
|
|
||
| let lineMatch: RegExpExecArray | null; | ||
| while ((lineMatch = lineRegex.exec(text)) !== null) { |
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Solves real-world bidirectional layout breakdowns observed in conversational AI platforms (such as OpenAI ChatGPT) when biomedical, scientific, and technical terminology is embedded into Persian and RTL prose.
Problem Context
When conversational models generate explanations in Persian/Arabic, technical jargon in English often appears at the beginning of sentences (e.g.
Capsule مثل یک روکش...,Macrophage برای بلعیدن...,Pyruvate kinase یکی از آن enzymeهایی است...). Under standard browser heuristics (HTMLdir="auto"or default LTR containers), UAX #9 rules P2/P3 evaluate only the first strong character, flipping the entire paragraph to LTR. Furthermore, multi-syllabic Latin medical terms outnumber short Persian grammatical particles, skewing character-count heuristics. Tables in Markdown lack container-leveldirattributes, reversing column orders, and loanword prefixes/suffixes (ضد-phagocytosis,macrophageها) break at script boundaries.Key Enhancements
Biomedical & Clinical Identifiers:
Expands
DEFAULT_TECHNICAL_IDENTIFIERSin@bidilens/corewith standard biomedical vocabulary (atp,rbc,g6pd,spleen,macrophage,capsule,hemolysis,spherocyte,reticulocyte,glycolysis,mitochondria,deficiency,autosomal,recessive,mutation,hereditary,chronic,nonspherocytic,hemolytic,anemia,pyruvate,kinase,glucose,extravascular,jaundice,splenomegaly,gallstones,retic,destruction,bilirubin,haptoglobin,coombs,normocytic,macrocytosis,morphology,oxidative,parvovirus,erythropoiesis,bpg,enzyme,pneumoniae,influenzae,meningitidis,antigen,antibody,complement,bacterium,clearance,membrane,deformability, etc.).Linear Reaction & Cascade Process Chains:
Adds high-performance, non-backtracking
addReactionRangesscanner for chemical reaction sequences and clinical cascades (A → B → C,PEP + ADP → Pyruvate + ATP). Bounded line-level scanning ensures linear performance ($O(N)$) and prevents catastrophic backtracking on repetitive tokens.Lab Trend & Serology Modifiers:
Detects trend indicators (
Hb ↓,Retic ↑,MCV↑) and serological test results (DAT−,DAT+,Coombs-,Rh-) without corrupting surrounding prose boundaries or misinterpreting standard English hyphenated words (e.g.,greetings-).Loanword Persian Prefix & Affix Recognition:
Detects Persian prefixes (
ضد-,پیش-,پس-,زیر-,ابر-,نیمه-,بی-,نا-,غیر-) and grammatical suffixes (ها,هایی,های,ای,اش,مان,تان,شان,تر,ترین,ات,ت,م,ی,یی,یها) attached to Latin stems.Markdown Table Base Direction:
Extends
markdownItBidiin@bidilens/markdownto inspect table headers (th). When headers are RTL (e.g.بیماری | سرنخ), the<table>element receivesdir="rtl", ensuring proper right-to-left column layout even when data cells contain English clinical text.Untampered Conformance Test Suite:
Evaluates all 33 ChatGPT Persian medical transcript examples out-of-the-box without manual parameter overrides, achieving 100% accuracy.
ADR-008 Architectural Record:
Includes
docs/architecture/adr/adr-008-llm-chat-bidi-breakdowns.mddetailing the root causes, UAX chore(deps-dev): bump globals from 17.7.0 to 17.8.0 in the development-tooling group across 1 directory #9 first-strong defects, and architectural solutions.Verification
pnpm run check: unicode reproducibility, typecheck, lint, package depth, coverage, corpus, docs, builds across all 17 packages, and Action bundle checks passed with zero errors.pnpm test: 23 test files, 626 tests passed, 0 failures.