Skip to content

M6: review, judge, and the baseline comparison #12

Description

@tamnd

Review, judge, and the baseline comparison. Spends model calls.

No deterministic check can tell a correct translation from a fluent wrong one. A Vietnamese sentence that says the opposite of the English, with every role span intact, passes all nine invariants. This milestone is the answer to that, and it is honest about being a sample.

Checklist

  • review.py: round-trip prompt, judge prompt, the three-valued verdict, stratified seeded sampling
  • The 100-entry calibration set: 50 known-good human entries, 50 corrupted by script with a dropped negation, a removed clause or a swapped term of art
  • review calibrate and its published rates
  • Round-trip 100 % of tier 1
  • reports/review.md with every differs entry in full, grouped by term and by rule rather than by file
  • The three-way baseline: the old MACHINE/ tree against tier 1 against the human strings, on P03 and G02
  • A person reads the flagged entries, and fixes land as a string, a glossary row or a prompt change

Exit

Calibration rates published. If the judge does not catch dropped negations at a high rate, this issue says so and the round-trip is reported as unreliable rather than quoted as a result. The three-way comparison table is in a comment here. Tier 1's glossary and prompt fixes are merged and tier 1 is re-run against them.

Metadata

Metadata

Assignees

No one assigned

    Labels

    fleetSpends model calls through the fleetmilestoneA milestone tracking issue, M0 through M10reviewRound-trip, judge, sampling, the human funnel

    Projects

    No projects

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions