From d431b2be715cf4f2442d3e91870376d05d607875 Mon Sep 17 00:00:00 2001 From: Ronald Tse Date: Mon, 7 Sep 2026 02:57:58 +0200 Subject: [PATCH] docs: propagate the 2.0 correction into 2.1 metadata and notes The three ara-diac-small-2.1 metadata files quote the E4 result in trained_from; the quoted 4.8218/42% figures were withdrawn on 2026-09-05 (RESULTS.md CORRECTION: corrected 2.0 number is 5.08, 38% reduction) - now restated with the correction inline. PUBLICATION-NOTES lever table and epoch sentence move to the corrected control; the E5 line keeps its original measurement with the control correction noted. EXPERIMENTS: the E4 verdict keeps its historical figure with the correction annotation, and the GKD verdict line (written 2026-09-06, post- correction, but still citing +1.18pp over 4.8218) restates the delta as +1.43pp over the 4.5701 rung it was meant to improve. Metrics values are untouched - only prose. --- docs/EXPERIMENTS.md | 4 ++-- docs/PUBLICATION-NOTES.md | 9 +++++---- .../ara-diac-small-2.1-fp16.metadata.yaml | 9 +++++---- .../ara-diac-small-2.1-int8.metadata.yaml | 9 +++++---- models/ara-diac-small/ara-diac-small-2.1.metadata.yaml | 9 +++++---- 5 files changed, 22 insertions(+), 18 deletions(-) diff --git a/docs/EXPERIMENTS.md b/docs/EXPERIMENTS.md index 0baa7e0..f25689f 100644 --- a/docs/EXPERIMENTS.md +++ b/docs/EXPERIMENTS.md @@ -192,7 +192,7 @@ All rows passed the CER parity gate at release. Readings: ## E4 — ara-diac-small-2.0 candidate (run-006-r7-muon) -- **Status:** COMPLETE (2026-08-29). **PASSED — 4.8218** full-set +- **Status:** COMPLETE (2026-08-29). **PASSED — 4.8218** full-set [CORRECTED 2026-09-05: the figure did not reproduce; the corrected 2.0 number is 5.08 (38% reduction, not 42%)] windowed DER (gate ≤ 6.26; registered prediction 4.3–5.0; teacher r7 reproduces 2.289 vs documented 2.2864). −3.44pp / 42% relative vs the shipped 1.0 at identical architecture and size; matches the PKM arm's @@ -324,7 +324,7 @@ All rows passed the CER parity gate at release. Readings: (n=1200; teacher reproduces 2.289; paired bootstrap student−teacher +3.4083, CI [3.109, 3.743]). NOT ADOPTED. The registered prediction (4.30-4.65) missed badly; honest-report band also breached — this - is the worst rung measured, +1.18pp over the 4.8218 control. Run + is the worst rung measured, +1.43pp over the 4.5701 rung it was meant to improve (+0.92pp over the corrected 5.08 control; the pre-correction text said +1.18pp over 4.8218). Run run-012-r7-muon-gkd: 10,995 steps, final CE 0.0076, nine server preemptions absorbed by checkpoint-resume (no measured work lost); labels sha256 e70ce991d15a8c810b83e2b5401f1410293844c623ffefa646b diff --git a/docs/PUBLICATION-NOTES.md b/docs/PUBLICATION-NOTES.md index d1111ac..03525d3 100644 --- a/docs/PUBLICATION-NOTES.md +++ b/docs/PUBLICATION-NOTES.md @@ -71,7 +71,7 @@ RL teacher polishing flat/negative ×3; microkimi bridges improve structure but not accuracy; teacher beam-search unnecessary for Arabic; per-channel int8 rejected on measurement; the 30 MiB tier closed as infeasible without pretraining; **E5 MTP-aux (2026-09-01): -5.0853 vs the 4.8218 control — multi-token-prediction as a training +5.0853 vs the control (then published 4.8218, corrected 2026-09-05 to 5.08) — multi-token-prediction as a training auxiliary HURT at this scale (+0.26pp), with a disclosed preemption confound (fresh aux head for the final 23% of steps); E6 constant-budget register swap (2026-09-02): 5.8057 — replacing news @@ -226,8 +226,9 @@ subset-inflation instances; the resume path now drops empty rows The E2/E3 factorial (3-epoch students) attributed the ~5.7pp gap as 0.70pp capacity + 2.73pp optimizer + ~2.25pp residual "domain coverage." The 6-epoch rung (G2a) and the CI-carrying harness revise -this: doubling epochs alone recovered 0.25pp full-set (4.8218 -> -4.5701, CIs non-overlapping) — the residual was not purely domain. +this: doubling epochs alone recovered 0.51pp full-set (corrected control +5.08 -> 4.5701, CIs non-overlapping; the control's original 4.8218 was +withdrawn 2026-09-05) — the residual was not purely domain. The decomposition for paper B, every line full-set with brackets: | lever | full-set DER | paired CI of delta | @@ -235,7 +236,7 @@ The decomposition for paper B, every line full-set with brackets: | teacher r7 | 2.2864/2.2921 | — | | 1.0: AdamW, 3ep, r6 | 8.259 | retrofit in flight | | + Muon (E3) | 5.2945 | — | -| + r7 teacher | 4.8218 | — | +| + r7 teacher | 5.08 (corrected 2026-09-05; 4.8218 withdrawn) | — | | + 6 epochs (G2a) | 4.5701 | delta 2.12 [1.91, 2.35] | | register swap (E6, 3ep) | 5.8057 | negative | | register add (G2b, 6ep) | 4.8231 | delta 2.37 [2.19, 2.55] | diff --git a/models/ara-diac-small-2.1/ara-diac-small-2.1-fp16.metadata.yaml b/models/ara-diac-small-2.1/ara-diac-small-2.1-fp16.metadata.yaml index 26fcaf7..1d5d07d 100644 --- a/models/ara-diac-small-2.1/ara-diac-small-2.1-fp16.metadata.yaml +++ b/models/ara-diac-small-2.1/ara-diac-small-2.1-fp16.metadata.yaml @@ -12,10 +12,11 @@ trained_from: 'sequence-level KD from the r7 canonical teacher (rababa_arabic_by 2.2864 windowed DER-CE full protocol): fresh greedy r7 labels on the same r5-units corpus/limits as ara-diac-small-1.0, Muon optimizer (E3-adopted), vanilla ByT5-small (E4, pre-registered gate <= 6.26). Checkpoint rababa-checkpoints:/rababa_arabic_distill_small/run-007-r7-muon-6ep/best. - The two measured wins compound: 8.259 -> 4.8218 full-set windowed DER-CE (teacher - reproduces 2.289 in-run vs documented 2.2864) — a 42% error reduction on the 1.0 - release at the same architecture and artifact size. Still misses the strict teacher+0.5pp - gate (+2.53pp; miss disclosed); the E2/E3 factorial attributes the residual to domain + The two measured wins compound: 8.259 -> 5.08 full-set windowed DER-CE (teacher + reproduces 2.289 in-run vs documented 2.2864) — a 38% error reduction on the 1.0 + release at the same architecture and artifact size (CORRECTED 2026-09-05: the + published 4.8218 did not reproduce; see RESULTS.md). Still misses the strict + teacher+0.5pp gate (+2.79pp; miss disclosed); the E2/E3 factorial attributes the residual to domain coverage.' metrics: - name: der_teacher_fullset diff --git a/models/ara-diac-small-2.1/ara-diac-small-2.1-int8.metadata.yaml b/models/ara-diac-small-2.1/ara-diac-small-2.1-int8.metadata.yaml index 23847e1..9e4ba85 100644 --- a/models/ara-diac-small-2.1/ara-diac-small-2.1-int8.metadata.yaml +++ b/models/ara-diac-small-2.1/ara-diac-small-2.1-int8.metadata.yaml @@ -12,10 +12,11 @@ trained_from: 'sequence-level KD from the r7 canonical teacher (rababa_arabic_by 2.2864 windowed DER-CE full protocol): fresh greedy r7 labels on the same r5-units corpus/limits as ara-diac-small-1.0, Muon optimizer (E3-adopted), vanilla ByT5-small (E4, pre-registered gate <= 6.26). Checkpoint rababa-checkpoints:/rababa_arabic_distill_small/run-007-r7-muon-6ep/best. - The two measured wins compound: 8.259 -> 4.8218 full-set windowed DER-CE (teacher - reproduces 2.289 in-run vs documented 2.2864) — a 42% error reduction on the 1.0 - release at the same architecture and artifact size. Still misses the strict teacher+0.5pp - gate (+2.53pp; miss disclosed); the E2/E3 factorial attributes the residual to domain + The two measured wins compound: 8.259 -> 5.08 full-set windowed DER-CE (teacher + reproduces 2.289 in-run vs documented 2.2864) — a 38% error reduction on the 1.0 + release at the same architecture and artifact size (CORRECTED 2026-09-05: the + published 4.8218 did not reproduce; see RESULTS.md). Still misses the strict + teacher+0.5pp gate (+2.79pp; miss disclosed); the E2/E3 factorial attributes the residual to domain coverage.' metrics: - name: der_teacher_fullset diff --git a/models/ara-diac-small/ara-diac-small-2.1.metadata.yaml b/models/ara-diac-small/ara-diac-small-2.1.metadata.yaml index 1e8bb8d..1ef387a 100644 --- a/models/ara-diac-small/ara-diac-small-2.1.metadata.yaml +++ b/models/ara-diac-small/ara-diac-small-2.1.metadata.yaml @@ -12,10 +12,11 @@ trained_from: 'sequence-level KD from the r7 canonical teacher (rababa_arabic_by 2.2864 windowed DER-CE full protocol): fresh greedy r7 labels on the same r5-units corpus/limits as ara-diac-small-1.0, Muon optimizer (E3-adopted), vanilla ByT5-small (E4, pre-registered gate <= 6.26). Checkpoint rababa-checkpoints:/rababa_arabic_distill_small/run-007-r7-muon-6ep/best. - The two measured wins compound: 8.259 -> 4.8218 full-set windowed DER-CE (teacher - reproduces 2.289 in-run vs documented 2.2864) — a 42% error reduction on the 1.0 - release at the same architecture and artifact size. Still misses the strict teacher+0.5pp - gate (+2.53pp; miss disclosed); the E2/E3 factorial attributes the residual to domain + The two measured wins compound: 8.259 -> 5.08 full-set windowed DER-CE (teacher + reproduces 2.289 in-run vs documented 2.2864) — a 38% error reduction on the 1.0 + release at the same architecture and artifact size (CORRECTED 2026-09-05: the + published 4.8218 did not reproduce; see RESULTS.md). Still misses the strict + teacher+0.5pp gate (+2.79pp; miss disclosed); the E2/E3 factorial attributes the residual to domain coverage.' metrics: - name: der_teacher_fullset