Machine Translation (MT) Evaluation Scripts
-
Updated
May 19, 2024 - Python
Machine Translation (MT) Evaluation Scripts
The Suboptimal Quality of WMT Test Sets and Their Impact on HumanParity
TamilLingBench: English→Tamil machine translation benchmark for obligatory verb agreement (rationality, gender, number). Tests whether MT evaluation metrics such as COMET, xCOMET, MetricX and CometKiwi catch wrong Tamil agreement. Code, data and model outputs for the WMT 2026 paper "Obligatory Slots".
[Unofficial] Interactive HTML visualization for MQM error typology for analytic Translation Quality Evaluation (TQE)
EN-UK parallel corpus pipeline built from Baldur's Gate 3 localization files: 186k aligned pairs, 7 text types, Gemini API translation and MT evaluation (BLEU, METEOR, TER, ChrF++, BERTScore). BA thesis project.
To associate your repository with the mt-evaluation topic, visit your repo's landing page and select "manage topics."