Skip to content

Commit df2d4fa

Browse files
Ronald Tseronaldtse
authored andcommitted
docs: propagate the 2.0 correction into 2.1 metadata and notes
The three ara-diac-small-2.1 metadata files quote the E4 result in trained_from; the quoted 4.8218/42% figures were withdrawn on 2026-09-05 (RESULTS.md CORRECTION: corrected 2.0 number is 5.08, 38% reduction) - now restated with the correction inline. PUBLICATION-NOTES lever table and epoch sentence move to the corrected control; the E5 line keeps its original measurement with the control correction noted. EXPERIMENTS: the E4 verdict keeps its historical figure with the correction annotation, and the GKD verdict line (written 2026-09-06, post- correction, but still citing +1.18pp over 4.8218) restates the delta as +1.43pp over the 4.5701 rung it was meant to improve. Metrics values are untouched - only prose.
1 parent c230c5b commit df2d4fa

5 files changed

Lines changed: 22 additions & 18 deletions

File tree

docs/EXPERIMENTS.md

Lines changed: 2 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -192,7 +192,7 @@ All rows passed the CER parity gate at release. Readings:
192192

193193
## E4 — ara-diac-small-2.0 candidate (run-006-r7-muon)
194194

195-
- **Status:** COMPLETE (2026-08-29). **PASSED — 4.8218** full-set
195+
- **Status:** COMPLETE (2026-08-29). **PASSED — 4.8218** full-set [CORRECTED 2026-09-05: the figure did not reproduce; the corrected 2.0 number is 5.08 (38% reduction, not 42%)]
196196
windowed DER (gate ≤ 6.26; registered prediction 4.3–5.0; teacher r7
197197
reproduces 2.289 vs documented 2.2864). −3.44pp / 42% relative vs the
198198
shipped 1.0 at identical architecture and size; matches the PKM arm's
@@ -324,7 +324,7 @@ All rows passed the CER parity gate at release. Readings:
324324
(n=1200; teacher reproduces 2.289; paired bootstrap student−teacher
325325
+3.4083, CI [3.109, 3.743]). NOT ADOPTED. The registered prediction
326326
(4.30-4.65) missed badly; honest-report band also breached — this
327-
is the worst rung measured, +1.18pp over the 4.8218 control. Run
327+
is the worst rung measured, +1.43pp over the 4.5701 rung it was meant to improve (+0.92pp over the corrected 5.08 control; the pre-correction text said +1.18pp over 4.8218). Run
328328
run-012-r7-muon-gkd: 10,995 steps, final CE 0.0076, nine server
329329
preemptions absorbed by checkpoint-resume (no measured work lost);
330330
labels sha256 e70ce991d15a8c810b83e2b5401f1410293844c623ffefa646b

docs/PUBLICATION-NOTES.md

Lines changed: 5 additions & 4 deletions
Original file line numberDiff line numberDiff line change
@@ -71,7 +71,7 @@ RL teacher polishing flat/negative ×3; microkimi bridges improve
7171
structure but not accuracy; teacher beam-search unnecessary for
7272
Arabic; per-channel int8 rejected on measurement; the 30 MiB tier
7373
closed as infeasible without pretraining; **E5 MTP-aux (2026-09-01):
74-
5.0853 vs the 4.8218 control — multi-token-prediction as a training
74+
5.0853 vs the control (then published 4.8218, corrected 2026-09-05 to 5.08) — multi-token-prediction as a training
7575
auxiliary HURT at this scale (+0.26pp), with a disclosed preemption
7676
confound (fresh aux head for the final 23% of steps); E6
7777
constant-budget register swap (2026-09-02): 5.8057 — replacing news
@@ -226,16 +226,17 @@ subset-inflation instances; the resume path now drops empty rows
226226
The E2/E3 factorial (3-epoch students) attributed the ~5.7pp gap as
227227
0.70pp capacity + 2.73pp optimizer + ~2.25pp residual "domain
228228
coverage." The 6-epoch rung (G2a) and the CI-carrying harness revise
229-
this: doubling epochs alone recovered 0.25pp full-set (4.8218 ->
230-
4.5701, CIs non-overlapping) — the residual was not purely domain.
229+
this: doubling epochs alone recovered 0.51pp full-set (corrected control
230+
5.08 -> 4.5701, CIs non-overlapping; the control's original 4.8218 was
231+
withdrawn 2026-09-05) — the residual was not purely domain.
231232
The decomposition for paper B, every line full-set with brackets:
232233

233234
| lever | full-set DER | paired CI of delta |
234235
|---|---|---|
235236
| teacher r7 | 2.2864/2.2921 ||
236237
| 1.0: AdamW, 3ep, r6 | 8.259 | retrofit in flight |
237238
| + Muon (E3) | 5.2945 ||
238-
| + r7 teacher | 4.8218 ||
239+
| + r7 teacher | 5.08 (corrected 2026-09-05; 4.8218 withdrawn) ||
239240
| + 6 epochs (G2a) | 4.5701 | delta 2.12 [1.91, 2.35] |
240241
| register swap (E6, 3ep) | 5.8057 | negative |
241242
| register add (G2b, 6ep) | 4.8231 | delta 2.37 [2.19, 2.55] |

models/ara-diac-small-2.1/ara-diac-small-2.1-fp16.metadata.yaml

Lines changed: 5 additions & 4 deletions
Original file line numberDiff line numberDiff line change
@@ -12,10 +12,11 @@ trained_from: 'sequence-level KD from the r7 canonical teacher (rababa_arabic_by
1212
2.2864 windowed DER-CE full protocol): fresh greedy r7 labels on the same r5-units
1313
corpus/limits as ara-diac-small-1.0, Muon optimizer (E3-adopted), vanilla ByT5-small
1414
(E4, pre-registered gate <= 6.26). Checkpoint rababa-checkpoints:/rababa_arabic_distill_small/run-007-r7-muon-6ep/best.
15-
The two measured wins compound: 8.259 -> 4.8218 full-set windowed DER-CE (teacher
16-
reproduces 2.289 in-run vs documented 2.2864) — a 42% error reduction on the 1.0
17-
release at the same architecture and artifact size. Still misses the strict teacher+0.5pp
18-
gate (+2.53pp; miss disclosed); the E2/E3 factorial attributes the residual to domain
15+
The two measured wins compound: 8.259 -> 5.08 full-set windowed DER-CE (teacher
16+
reproduces 2.289 in-run vs documented 2.2864) — a 38% error reduction on the 1.0
17+
release at the same architecture and artifact size (CORRECTED 2026-09-05: the
18+
published 4.8218 did not reproduce; see RESULTS.md). Still misses the strict
19+
teacher+0.5pp gate (+2.79pp; miss disclosed); the E2/E3 factorial attributes the residual to domain
1920
coverage.'
2021
metrics:
2122
- name: der_teacher_fullset

models/ara-diac-small-2.1/ara-diac-small-2.1-int8.metadata.yaml

Lines changed: 5 additions & 4 deletions
Original file line numberDiff line numberDiff line change
@@ -12,10 +12,11 @@ trained_from: 'sequence-level KD from the r7 canonical teacher (rababa_arabic_by
1212
2.2864 windowed DER-CE full protocol): fresh greedy r7 labels on the same r5-units
1313
corpus/limits as ara-diac-small-1.0, Muon optimizer (E3-adopted), vanilla ByT5-small
1414
(E4, pre-registered gate <= 6.26). Checkpoint rababa-checkpoints:/rababa_arabic_distill_small/run-007-r7-muon-6ep/best.
15-
The two measured wins compound: 8.259 -> 4.8218 full-set windowed DER-CE (teacher
16-
reproduces 2.289 in-run vs documented 2.2864) — a 42% error reduction on the 1.0
17-
release at the same architecture and artifact size. Still misses the strict teacher+0.5pp
18-
gate (+2.53pp; miss disclosed); the E2/E3 factorial attributes the residual to domain
15+
The two measured wins compound: 8.259 -> 5.08 full-set windowed DER-CE (teacher
16+
reproduces 2.289 in-run vs documented 2.2864) — a 38% error reduction on the 1.0
17+
release at the same architecture and artifact size (CORRECTED 2026-09-05: the
18+
published 4.8218 did not reproduce; see RESULTS.md). Still misses the strict
19+
teacher+0.5pp gate (+2.79pp; miss disclosed); the E2/E3 factorial attributes the residual to domain
1920
coverage.'
2021
metrics:
2122
- name: der_teacher_fullset

models/ara-diac-small/ara-diac-small-2.1.metadata.yaml

Lines changed: 5 additions & 4 deletions
Original file line numberDiff line numberDiff line change
@@ -12,10 +12,11 @@ trained_from: 'sequence-level KD from the r7 canonical teacher (rababa_arabic_by
1212
2.2864 windowed DER-CE full protocol): fresh greedy r7 labels on the same r5-units
1313
corpus/limits as ara-diac-small-1.0, Muon optimizer (E3-adopted), vanilla ByT5-small
1414
(E4, pre-registered gate <= 6.26). Checkpoint rababa-checkpoints:/rababa_arabic_distill_small/run-007-r7-muon-6ep/best.
15-
The two measured wins compound: 8.259 -> 4.8218 full-set windowed DER-CE (teacher
16-
reproduces 2.289 in-run vs documented 2.2864) — a 42% error reduction on the 1.0
17-
release at the same architecture and artifact size. Still misses the strict teacher+0.5pp
18-
gate (+2.53pp; miss disclosed); the E2/E3 factorial attributes the residual to domain
15+
The two measured wins compound: 8.259 -> 5.08 full-set windowed DER-CE (teacher
16+
reproduces 2.289 in-run vs documented 2.2864) — a 38% error reduction on the 1.0
17+
release at the same architecture and artifact size (CORRECTED 2026-09-05: the
18+
published 4.8218 did not reproduce; see RESULTS.md). Still misses the strict
19+
teacher+0.5pp gate (+2.79pp; miss disclosed); the E2/E3 factorial attributes the residual to domain
1920
coverage.'
2021
metrics:
2122
- name: der_teacher_fullset

0 commit comments

Comments
 (0)