Skip to content

Commit 405a442

Browse files
committed
site: finish the 2.0 correction — negatives note and blog table
Follows 6df21ff (the ladder rung correction): the negatives note keeps each delta as measured, annotated against the corrected 5.08 control, withdraws the dead 'cancelled the epoch gain' reading, and states the withdrawal itself. The blog frontier table's 2.0 row moves to 5.08 with the correction date; its CI matches the re-eval interval [2.36, 2.82].
1 parent 39eb4bc commit 405a442

2 files changed

Lines changed: 14 additions & 7 deletions

File tree

src/content/blog/2026-09-05-frontier-closed.adoc

Lines changed: 3 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -25,7 +25,7 @@ bootstrap interval on the gap to the teacher:
2525
|from-scratch 30M, every lever |30M |73.95 |[70.22, 71.09]
2626
|1.0 rung |300M |8.26 |[4.6, 5.7]
2727
|lite rung (enc 12→6) |190M |5.78 |[3.03, 3.49]
28-
|2.0 rung |300M |4.82 |[2.36, 2.82]
28+
|2.0 rung |300M |5.08 (corrected 2026-09-05) |[2.36, 2.82]
2929
|2.1 rung |300M |4.57 |[1.91, 2.35]
3030
|teacher r7 |580M |2.29 |—
3131
|===
@@ -52,7 +52,8 @@ per-step confidence.
5252
The residual gap between the 2.1 student and its teacher had one live
5353
attribution left: classical-domain coverage. Both tests of it failed.
5454
Swapping news-domain training units for classical Tashkeela at constant
55-
budget made the student worse (−0.98pp). Adding five times the
55+
budget made the student worse (+0.72pp over the corrected 2.0
56+
control). Adding five times the
5657
classical corpus — the maximal version of the lever — left it
5758
statistically flat (4.82 vs 4.57, intervals overlapping). The residual
5859
is not a coverage deficit the corpus can reach; it lives in the

src/pages/ml.astro

Lines changed: 11 additions & 5 deletions
Original file line numberDiff line numberDiff line change
@@ -338,11 +338,17 @@ const clientModels = [
338338
</ul>
339339
<p class="lb-note">
340340
What didn't move it — every rung pre-registered, measured, and kept in the log:
341-
multi-token-prediction auxiliary 5.09 (+0.26pp); classical-register swap at constant
342-
budget 5.81 (+0.98pp); register add at matched epochs 4.82 (flat — it cancelled the epoch
343-
gain); on-policy GKD 6.00 (+1.18pp, the worst rung); depth-halved 5.78 (63% of the
344-
parameters — shipped anyway as the browser tier). At SFT convergence, supervision quality
345-
dominates. The ladder is closed.
341+
multi-token-prediction auxiliary 5.09 (+0.26pp as measured vs the
342+
then-published control — level with the corrected 5.08); classical-register swap at
343+
constant budget 5.81 (+0.98pp as measured; +0.72pp over the corrected control); register
344+
add at matched epochs 4.82 (statistically flat against news-only 4.57, intervals
345+
overlapping — it sits between the corrected control and the news-only rung); on-policy
346+
GKD 6.00 (+1.18pp as measured; +1.43pp over the rung it was meant to improve — the worst
347+
measured); depth-halved 5.78 (63% of the parameters — shipped anyway as the browser
348+
tier). At SFT convergence, supervision quality dominates. The ladder is closed. The
349+
control's originally published 4.8218 did not reproduce and was withdrawn on
350+
2026-09-05; two independent re-measurements, one at the artifact level on the exact
351+
released bytes, agree on 5.08.
346352
</p>
347353
<p class="lb-note">
348354
Hebrew, same discipline: 16.43% DER on Biblical Hebrew, where the modern-Hebrew state of

0 commit comments

Comments
 (0)