Review, judge, and the baseline comparison. Spends model calls.
No deterministic check can tell a correct translation from a fluent wrong one. A Vietnamese sentence that says the opposite of the English, with every role span intact, passes all nine invariants. This milestone is the answer to that, and it is honest about being a sample.
Checklist
Exit
Calibration rates published. If the judge does not catch dropped negations at a high rate, this issue says so and the round-trip is reported as unreliable rather than quoted as a result. The three-way comparison table is in a comment here. Tier 1's glossary and prompt fixes are merged and tier 1 is re-run against them.
Review, judge, and the baseline comparison. Spends model calls.
No deterministic check can tell a correct translation from a fluent wrong one. A Vietnamese sentence that says the opposite of the English, with every role span intact, passes all nine invariants. This milestone is the answer to that, and it is honest about being a sample.
Checklist
review.py: round-trip prompt, judge prompt, the three-valued verdict, stratified seeded samplingreview calibrateand its published ratesreports/review.mdwith everydiffersentry in full, grouped by term and by rule rather than by fileMACHINE/tree against tier 1 against the human strings, onP03andG02Exit
Calibration rates published. If the judge does not catch dropped negations at a high rate, this issue says so and the round-trip is reported as unreliable rather than quoted as a result. The three-way comparison table is in a comment here. Tier 1's glossary and prompt fixes are merged and tier 1 is re-run against them.