Skip to content

paper: protocol-matched external benchmark positioning + 2024-25 related work - #52

Merged
ronaldtse merged 1 commit into
mainfrom
docs/external-benchmarks
Aug 26, 2026
Merged

paper: protocol-matched external benchmark positioning + 2024-25 related work#52
ronaldtse merged 1 commit into
mainfrom
docs/external-benchmarks

Conversation

@ronaldtse

Copy link
Copy Markdown
Contributor

Summary

  • New paper results subsection Positioning against public benchmarks: verified SadeedDiac-25 leaderboard (teacher r6 2.5793 = best dedicated model, 2nd overall behind Claude-3.7-Sonnet's 1.3941, ahead of GLM-5.2 reproduction 2.6911 / Gemini-Flash 3.1926 / GPT-4 3.8645 / Sadeed-1.5B 7.2915), Hebrew Nakdimon-domain comparison (s45 16.58% DER vs DictaBERT-large 35.63% on the same test), Thai no-external-benchmark disclosure, Persian SentenceBench scoping
  • Related work upgraded with the 2024-2025 landscape: SadeedDiac-25 (arXiv 2504.21635), multi-reference evaluation (Mohamed & Mubarak, EMNLP 2025), LLM-prompted G2P (arXiv 2409.08554, 8.30% PER on its own benchmark), intermediate-language G2P (arXiv 2505.06599), D-Nikud (arXiv 2402.00075), DiVRit (arXiv 2510.26521)
  • Knowledge distillation section positions on-policy distillation / GKD (MiniLLM, Agarwal et al.) as the open lever against the measured flat-distribution pathology; added as explicit future work in Discussion
  • Discussion adds benchmark-fragmentation limitation; notes teacher-tier beam gain is language-dependent (Hebrew +12 DER points, Arabic noise)
  • RESULTS.md: Persian claim scoped to SentenceBench (concurrent lines use different PER benchmarks); ara-diac-small section gains leaderboard context; 9 new references, all verified against primary sources

Numbers cross-checked against the primary logs (rababa docs/RESULTS.md r6 verdict table; secryst-train arabic comparison table). No cross-protocol comparisons are claimed anywhere.

Test plan

  • asciidoc table/footnote syntax matches existing sections
  • every external number traced to a primary results log or a fetched citation page
  • full-set ara-diac-small number lands separately and updates the subset value

…ted work

- New results subsection: SadeedDiac-25 leaderboard table (r6 2.5793 best
  dedicated, 2nd overall behind Claude-3.7 1.3941; client student 3.66 on
  300-para subset, full-set in flight), Hebrew Nakdimon-domain comparison
  (s45 16.58 vs DictaBERT-large 35.63 same test), Thai no-benchmark note,
  Persian SentenceBench scoping
- Related work: SadeedDiac-25, QCRI multi-reference eval (EMNLP 2025),
  LLM-prompted G2P (8.30% PER, different benchmark), intermediate-language
  G2P, D-Nikud, DiVRit
- Distillation: on-policy/GKD (MiniLLM, Agarwal et al.) positioned as the
  open lever against the flat-distribution decode pathology
- Discussion: benchmark-fragmentation limitation; teacher-tier beam gain is
  language-dependent (Hebrew +12 DER pts, Arabic ~0)
- Nine new references, all verified against primary sources
@ronaldtse
ronaldtse merged commit 4af2a6d into main Aug 26, 2026
9 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant