From 740c4395b7d1d72ad985182da2faa693d75a1c6b Mon Sep 17 00:00:00 2001 From: Ronald Tse Date: Sat, 5 Sep 2026 18:24:36 +0200 Subject: [PATCH] =?UTF-8?q?docs:=20CORRECTION=20=E2=80=94=20ara-diac-small?= =?UTF-8?q?-2.0=20full-set=20is=205.08,=20not=204.8218?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The published number does not reproduce: two independent later measurements of the shipped bytes (harness re-eval 5.0821; artifact- level runtime scoring of the sha-pinned release zip 5.0329) agree and disagree with 4.8218. Ledger correction entry with consequences (38% not 42% E4 reduction; PKM arm was better, not matched; frontier ordering unchanged). models.yaml corrected; Paper B frontier cell annotated. All other rows re-verified exactly by the same tooling. --- docs/RESULTS.md | 33 +++++++++++++++++++++++++++++++++ docs/paper.adoc | 12 ++++++++++-- models.yaml | 2 +- 3 files changed, 44 insertions(+), 3 deletions(-) diff --git a/docs/RESULTS.md b/docs/RESULTS.md index 38dd6a6..d99dc02 100644 --- a/docs/RESULTS.md +++ b/docs/RESULTS.md @@ -548,3 +548,36 @@ closed as a negative: depth-compressibility is NOT a universal property of pretrained ByT5-small; it held under one recipe on one language. Paper B's depth paragraph is scoped accordingly (this entry is its counterexample). + +## CORRECTION: ara-diac-small-2.0 full-set is 5.08, not 4.8218 (2026-09-05) + +The published 2.0 number (2026-08-30, in this ledger and the shipped +metadata) does not reproduce. Two independent later measurements of +the same checkpoint agree and disagree with it: + +| Measurement | Path | DER-CE | +|---|---|---| +| published (2026-08-30) | harness, in-run | 4.8218 | +| harness re-eval (final_eval.json, bootstrap CI [2.358, 2.823]) | torch, same protocol | 5.0821 | +| **artifact-level (this entry)** | shipped zip sha d9aa95d0 (= index pin = release bytes), windowed ONNX runtime decode, sadeedbench scoring | **5.0321-class** (5.0329) | + +The artifact-level measurement is the governing one: it scores the +exact bytes users download. The 4.8218 figure is withdrawn; the 2.0 +rung's catalog number is 5.08 (harness) / 5.03 (runtime path). The +cause of the original reading is not reconstructed; both later +measurements postdate it and agree to 0.05pp across independent decode +paths. + +Consequences, stated plainly: +- the E4 error-reduction claim becomes 8.259 -> 5.08 (38%, not 42%) +- the "matches the PKM arm (4.829)" statement is wrong: the vanilla + 2.0 rung (5.08) does NOT match the PKM arm; the PKM arm was better + by 0.25pp at its measurement +- frontier ordering 1.0 -> lite -> 2.0 -> 2.1 is unchanged; G2b + (4.8231) sits between 2.0 and 2.1 rather than above 2.0 +- every other frontier row re-verified exactly by the same + artifact/preds-level tooling (8.2576 / 4.5701 / 5.784 / 4.8231) + +Predictions for all five frontier runs publish alongside this entry +(release frontier-predictions-v1) so the numbers above are re-derivable +by anyone. diff --git a/docs/paper.adoc b/docs/paper.adoc index 7baa05f..fb12166 100644 --- a/docs/paper.adoc +++ b/docs/paper.adoc @@ -286,13 +286,21 @@ The Arabic client line has since been extended to a complete, confidence-bracket |from-scratch d384, every lever |30M |73.95 |[70.22, 71.09] |1.0 rung (AdamW, 3 ep) |300M |8.26 |[4.6, 5.7] |layerdrop (enc 12→6, Muon, 6 ep) |190M |5.78 |[3.03, 3.49] -|2.0 rung (Muon, 3 ep) |300M |4.82 |[2.36, 2.82] +|2.0 rung (Muon, 3 ep) |300M |5.08 |[2.36, 2.82] (corrected; see below) |2.1 rung (Muon, 6 ep) |300M |4.57 |[1.91, 2.35] |2.1 + full Tashkeela (G2b) |300M |4.82 |[2.19, 2.55] |teacher r7 |580M |2.29 |— |=== -Intervals are disjoint between adjacent rungs end to end: the frontier's separations are statistically real, not seed noise. Two structural findings sit on it, one now scoped by a cross-lingual counterexample. First, *pretrained width is load-bearing*: SVD width-stitching of the pretrained ByT5-small fails at both tested ratios (3.8× narrow: 82.96 — worse than the from-scratch collapse; 2× narrow: 78.23). *Depth, by contrast, was compressible under our Arabic recipe*: a verbatim layer copy of encoder layers 12→6 trained to 5.78 full-set — a shippable rung at 63% of the parameters whose 1.21pp depth cost at 6 epochs is CI-separated from the full-depth peer. The Hebrew replication of that depth cut, single-variable against its own full-depth lineage (logit-KD recipe), collapsed instead: 77.48 DER versus the full-depth 30.38 (+53.77pp [51.64, 55.92]) despite normal training convergence. Depth-compressibility is therefore *not* a universal property of pretrained ByT5-small — it held under sequence-KD with Muon on Arabic and catastrophically failed under logit-KD on Hebrew; whether the boundary is the distillation regime or the language is open. Width surgery destroys the pretrained representation; depth surgery spends it — but only where the training regime lets it. Second, the epochs lever is real but secondary: doubling 3→6 epochs moves the rung 4.82→4.57 (−0.25pp) — most of the residual is not undertraining. Two pre-registered causal tests closed the domain-coverage attribution. The swap direction — 8k news-domain units replaced by classical-register Tashkeela at constant 30k budget — came back *negative* (5.81, −0.98pp vs control). The add direction — 48k total with the full cleaned Tashkeela corpus (5× classical coverage, all other levers held, 39,018 steps) — came back *flat-negative*: 4.8231 full-set, delta vs teacher 2.3717 [2.194, 2.554], statistically indistinguishable from the 2.1 rung (4.5701 [1.91, 2.35]) with the point estimate 0.25pp worse. Both directions of the classical-corpus lever fail: the residual is *not* a classical-domain coverage deficit. It reframes as a teacher–student interaction the corpus cannot reach — candidate levers are on-policy distillation (in flight) and label-distribution effects, not more in-domain text. +Intervals are disjoint between adjacent rungs end to end: the frontier's separations are statistically real, not seed noise. A correction attaches to one row: the 2.0 rung was first published at +4.8218; two independent re-measurements of the identical shipped bytes +(harness re-evaluation, and an artifact-level scoring of the released +zip through the runtime itself — sha-pinned to the index) agree on +5.03–5.08, and the higher figure stands. Every other row of the table +re-derives exactly by the same tooling from the published prediction +files, which is how the discrepancy was caught. + +Two structural findings sit on it, one now scoped by a cross-lingual counterexample. First, *pretrained width is load-bearing*: SVD width-stitching of the pretrained ByT5-small fails at both tested ratios (3.8× narrow: 82.96 — worse than the from-scratch collapse; 2× narrow: 78.23). *Depth, by contrast, was compressible under our Arabic recipe*: a verbatim layer copy of encoder layers 12→6 trained to 5.78 full-set — a shippable rung at 63% of the parameters whose 1.21pp depth cost at 6 epochs is CI-separated from the full-depth peer. The Hebrew replication of that depth cut, single-variable against its own full-depth lineage (logit-KD recipe), collapsed instead: 77.48 DER versus the full-depth 30.38 (+53.77pp [51.64, 55.92]) despite normal training convergence. Depth-compressibility is therefore *not* a universal property of pretrained ByT5-small — it held under sequence-KD with Muon on Arabic and catastrophically failed under logit-KD on Hebrew; whether the boundary is the distillation regime or the language is open. Width surgery destroys the pretrained representation; depth surgery spends it — but only where the training regime lets it. Second, the epochs lever is real but secondary: doubling 3→6 epochs moves the rung 4.82→4.57 (−0.25pp) — most of the residual is not undertraining. Two pre-registered causal tests closed the domain-coverage attribution. The swap direction — 8k news-domain units replaced by classical-register Tashkeela at constant 30k budget — came back *negative* (5.81, −0.98pp vs control). The add direction — 48k total with the full cleaned Tashkeela corpus (5× classical coverage, all other levers held, 39,018 steps) — came back *flat-negative*: 4.8231 full-set, delta vs teacher 2.3717 [2.194, 2.554], statistically indistinguishable from the 2.1 rung (4.5701 [1.91, 2.35]) with the point estimate 0.25pp worse. Both directions of the classical-corpus lever fail: the residual is *not* a classical-domain coverage deficit. It reframes as a teacher–student interaction the corpus cannot reach — candidate levers are on-policy distillation (in flight) and label-distribution effects, not more in-domain text. == The decode protocol is part of the measurement [[section-decode]] diff --git a/models.yaml b/models.yaml index db2961e..a3ec745 100644 --- a/models.yaml +++ b/models.yaml @@ -240,7 +240,7 @@ models: size: 1418516306 metrics: - {name: der_teacher_fullset, value: 2.289, source: interscript/interscript-ml docs/RESULTS.md#run-006-r7-muon} - - {name: der_student_fullset, value: 4.8218, source: interscript/interscript-ml docs/RESULTS.md#run-006-r7-muon} + - {name: der_student_fullset, value: 5.0821, source: interscript/interscript-ml docs/RESULTS.md#run-006-r7-muon} parity: {samples: 600, cer_delta: 0.1187} license: BSD-3-Clause ara-diac-small-2.0-int8: