Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
33 changes: 33 additions & 0 deletions docs/RESULTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -548,3 +548,36 @@ closed as a negative: depth-compressibility is NOT a universal
property of pretrained ByT5-small; it held under one recipe on one
language. Paper B's depth paragraph is scoped accordingly (this entry
is its counterexample).

## CORRECTION: ara-diac-small-2.0 full-set is 5.08, not 4.8218 (2026-09-05)

The published 2.0 number (2026-08-30, in this ledger and the shipped
metadata) does not reproduce. Two independent later measurements of
the same checkpoint agree and disagree with it:

| Measurement | Path | DER-CE |
|---|---|---|
| published (2026-08-30) | harness, in-run | 4.8218 |
| harness re-eval (final_eval.json, bootstrap CI [2.358, 2.823]) | torch, same protocol | 5.0821 |
| **artifact-level (this entry)** | shipped zip sha d9aa95d0 (= index pin = release bytes), windowed ONNX runtime decode, sadeedbench scoring | **5.0321-class** (5.0329) |

The artifact-level measurement is the governing one: it scores the
exact bytes users download. The 4.8218 figure is withdrawn; the 2.0
rung's catalog number is 5.08 (harness) / 5.03 (runtime path). The
cause of the original reading is not reconstructed; both later
measurements postdate it and agree to 0.05pp across independent decode
paths.

Consequences, stated plainly:
- the E4 error-reduction claim becomes 8.259 -> 5.08 (38%, not 42%)
- the "matches the PKM arm (4.829)" statement is wrong: the vanilla
2.0 rung (5.08) does NOT match the PKM arm; the PKM arm was better
by 0.25pp at its measurement
- frontier ordering 1.0 -> lite -> 2.0 -> 2.1 is unchanged; G2b
(4.8231) sits between 2.0 and 2.1 rather than above 2.0
- every other frontier row re-verified exactly by the same
artifact/preds-level tooling (8.2576 / 4.5701 / 5.784 / 4.8231)

Predictions for all five frontier runs publish alongside this entry
(release frontier-predictions-v1) so the numbers above are re-derivable
by anyone.
12 changes: 10 additions & 2 deletions docs/paper.adoc
Original file line number Diff line number Diff line change
Expand Up @@ -286,13 +286,21 @@ The Arabic client line has since been extended to a complete, confidence-bracket
|from-scratch d384, every lever |30M |73.95 |[70.22, 71.09]
|1.0 rung (AdamW, 3 ep) |300M |8.26 |[4.6, 5.7]
|layerdrop (enc 12→6, Muon, 6 ep) |190M |5.78 |[3.03, 3.49]
|2.0 rung (Muon, 3 ep) |300M |4.82 |[2.36, 2.82]
|2.0 rung (Muon, 3 ep) |300M |5.08 |[2.36, 2.82] (corrected; see below)
|2.1 rung (Muon, 6 ep) |300M |4.57 |[1.91, 2.35]
|2.1 + full Tashkeela (G2b) |300M |4.82 |[2.19, 2.55]
|teacher r7 |580M |2.29 |—
|===

Intervals are disjoint between adjacent rungs end to end: the frontier's separations are statistically real, not seed noise. Two structural findings sit on it, one now scoped by a cross-lingual counterexample. First, *pretrained width is load-bearing*: SVD width-stitching of the pretrained ByT5-small fails at both tested ratios (3.8× narrow: 82.96 — worse than the from-scratch collapse; 2× narrow: 78.23). *Depth, by contrast, was compressible under our Arabic recipe*: a verbatim layer copy of encoder layers 12→6 trained to 5.78 full-set — a shippable rung at 63% of the parameters whose 1.21pp depth cost at 6 epochs is CI-separated from the full-depth peer. The Hebrew replication of that depth cut, single-variable against its own full-depth lineage (logit-KD recipe), collapsed instead: 77.48 DER versus the full-depth 30.38 (+53.77pp [51.64, 55.92]) despite normal training convergence. Depth-compressibility is therefore *not* a universal property of pretrained ByT5-small — it held under sequence-KD with Muon on Arabic and catastrophically failed under logit-KD on Hebrew; whether the boundary is the distillation regime or the language is open. Width surgery destroys the pretrained representation; depth surgery spends it — but only where the training regime lets it. Second, the epochs lever is real but secondary: doubling 3→6 epochs moves the rung 4.82→4.57 (−0.25pp) — most of the residual is not undertraining. Two pre-registered causal tests closed the domain-coverage attribution. The swap direction — 8k news-domain units replaced by classical-register Tashkeela at constant 30k budget — came back *negative* (5.81, −0.98pp vs control). The add direction — 48k total with the full cleaned Tashkeela corpus (5× classical coverage, all other levers held, 39,018 steps) — came back *flat-negative*: 4.8231 full-set, delta vs teacher 2.3717 [2.194, 2.554], statistically indistinguishable from the 2.1 rung (4.5701 [1.91, 2.35]) with the point estimate 0.25pp worse. Both directions of the classical-corpus lever fail: the residual is *not* a classical-domain coverage deficit. It reframes as a teacher–student interaction the corpus cannot reach — candidate levers are on-policy distillation (in flight) and label-distribution effects, not more in-domain text.
Intervals are disjoint between adjacent rungs end to end: the frontier's separations are statistically real, not seed noise. A correction attaches to one row: the 2.0 rung was first published at
4.8218; two independent re-measurements of the identical shipped bytes
(harness re-evaluation, and an artifact-level scoring of the released
zip through the runtime itself — sha-pinned to the index) agree on
5.03–5.08, and the higher figure stands. Every other row of the table
re-derives exactly by the same tooling from the published prediction
files, which is how the discrepancy was caught.

Two structural findings sit on it, one now scoped by a cross-lingual counterexample. First, *pretrained width is load-bearing*: SVD width-stitching of the pretrained ByT5-small fails at both tested ratios (3.8× narrow: 82.96 — worse than the from-scratch collapse; 2× narrow: 78.23). *Depth, by contrast, was compressible under our Arabic recipe*: a verbatim layer copy of encoder layers 12→6 trained to 5.78 full-set — a shippable rung at 63% of the parameters whose 1.21pp depth cost at 6 epochs is CI-separated from the full-depth peer. The Hebrew replication of that depth cut, single-variable against its own full-depth lineage (logit-KD recipe), collapsed instead: 77.48 DER versus the full-depth 30.38 (+53.77pp [51.64, 55.92]) despite normal training convergence. Depth-compressibility is therefore *not* a universal property of pretrained ByT5-small — it held under sequence-KD with Muon on Arabic and catastrophically failed under logit-KD on Hebrew; whether the boundary is the distillation regime or the language is open. Width surgery destroys the pretrained representation; depth surgery spends it — but only where the training regime lets it. Second, the epochs lever is real but secondary: doubling 3→6 epochs moves the rung 4.82→4.57 (−0.25pp) — most of the residual is not undertraining. Two pre-registered causal tests closed the domain-coverage attribution. The swap direction — 8k news-domain units replaced by classical-register Tashkeela at constant 30k budget — came back *negative* (5.81, −0.98pp vs control). The add direction — 48k total with the full cleaned Tashkeela corpus (5× classical coverage, all other levers held, 39,018 steps) — came back *flat-negative*: 4.8231 full-set, delta vs teacher 2.3717 [2.194, 2.554], statistically indistinguishable from the 2.1 rung (4.5701 [1.91, 2.35]) with the point estimate 0.25pp worse. Both directions of the classical-corpus lever fail: the residual is *not* a classical-domain coverage deficit. It reframes as a teacher–student interaction the corpus cannot reach — candidate levers are on-policy distillation (in flight) and label-distribution effects, not more in-domain text.

== The decode protocol is part of the measurement
[[section-decode]]
Expand Down
2 changes: 1 addition & 1 deletion models.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -240,7 +240,7 @@ models:
size: 1418516306
metrics:
- {name: der_teacher_fullset, value: 2.289, source: interscript/interscript-ml docs/RESULTS.md#run-006-r7-muon}
- {name: der_student_fullset, value: 4.8218, source: interscript/interscript-ml docs/RESULTS.md#run-006-r7-muon}
- {name: der_student_fullset, value: 5.0821, source: interscript/interscript-ml docs/RESULTS.md#run-006-r7-muon}
parity: {samples: 600, cer_delta: 0.1187}
license: BSD-3-Clause
ara-diac-small-2.0-int8:
Expand Down
Loading