Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 2 additions & 2 deletions docs/EXPERIMENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -192,7 +192,7 @@ All rows passed the CER parity gate at release. Readings:

## E4 — ara-diac-small-2.0 candidate (run-006-r7-muon)

- **Status:** COMPLETE (2026-08-29). **PASSED — 4.8218** full-set
- **Status:** COMPLETE (2026-08-29). **PASSED — 4.8218** full-set [CORRECTED 2026-09-05: the figure did not reproduce; the corrected 2.0 number is 5.08 (38% reduction, not 42%)]
windowed DER (gate ≤ 6.26; registered prediction 4.3–5.0; teacher r7
reproduces 2.289 vs documented 2.2864). −3.44pp / 42% relative vs the
shipped 1.0 at identical architecture and size; matches the PKM arm's
Expand Down Expand Up @@ -324,7 +324,7 @@ All rows passed the CER parity gate at release. Readings:
(n=1200; teacher reproduces 2.289; paired bootstrap student−teacher
+3.4083, CI [3.109, 3.743]). NOT ADOPTED. The registered prediction
(4.30-4.65) missed badly; honest-report band also breached — this
is the worst rung measured, +1.18pp over the 4.8218 control. Run
is the worst rung measured, +1.43pp over the 4.5701 rung it was meant to improve (+0.92pp over the corrected 5.08 control; the pre-correction text said +1.18pp over 4.8218). Run
run-012-r7-muon-gkd: 10,995 steps, final CE 0.0076, nine server
preemptions absorbed by checkpoint-resume (no measured work lost);
labels sha256 e70ce991d15a8c810b83e2b5401f1410293844c623ffefa646b
Expand Down
9 changes: 5 additions & 4 deletions docs/PUBLICATION-NOTES.md
Original file line number Diff line number Diff line change
Expand Up @@ -71,7 +71,7 @@ RL teacher polishing flat/negative ×3; microkimi bridges improve
structure but not accuracy; teacher beam-search unnecessary for
Arabic; per-channel int8 rejected on measurement; the 30 MiB tier
closed as infeasible without pretraining; **E5 MTP-aux (2026-09-01):
5.0853 vs the 4.8218 control — multi-token-prediction as a training
5.0853 vs the control (then published 4.8218, corrected 2026-09-05 to 5.08) — multi-token-prediction as a training
auxiliary HURT at this scale (+0.26pp), with a disclosed preemption
confound (fresh aux head for the final 23% of steps); E6
constant-budget register swap (2026-09-02): 5.8057 — replacing news
Expand Down Expand Up @@ -226,16 +226,17 @@ subset-inflation instances; the resume path now drops empty rows
The E2/E3 factorial (3-epoch students) attributed the ~5.7pp gap as
0.70pp capacity + 2.73pp optimizer + ~2.25pp residual "domain
coverage." The 6-epoch rung (G2a) and the CI-carrying harness revise
this: doubling epochs alone recovered 0.25pp full-set (4.8218 ->
4.5701, CIs non-overlapping) — the residual was not purely domain.
this: doubling epochs alone recovered 0.51pp full-set (corrected control
5.08 -> 4.5701, CIs non-overlapping; the control's original 4.8218 was
withdrawn 2026-09-05) — the residual was not purely domain.
The decomposition for paper B, every line full-set with brackets:

| lever | full-set DER | paired CI of delta |
|---|---|---|
| teacher r7 | 2.2864/2.2921 | — |
| 1.0: AdamW, 3ep, r6 | 8.259 | retrofit in flight |
| + Muon (E3) | 5.2945 | — |
| + r7 teacher | 4.8218 | — |
| + r7 teacher | 5.08 (corrected 2026-09-05; 4.8218 withdrawn) | — |
| + 6 epochs (G2a) | 4.5701 | delta 2.12 [1.91, 2.35] |
| register swap (E6, 3ep) | 5.8057 | negative |
| register add (G2b, 6ep) | 4.8231 | delta 2.37 [2.19, 2.55] |
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -12,10 +12,11 @@ trained_from: 'sequence-level KD from the r7 canonical teacher (rababa_arabic_by
2.2864 windowed DER-CE full protocol): fresh greedy r7 labels on the same r5-units
corpus/limits as ara-diac-small-1.0, Muon optimizer (E3-adopted), vanilla ByT5-small
(E4, pre-registered gate <= 6.26). Checkpoint rababa-checkpoints:/rababa_arabic_distill_small/run-007-r7-muon-6ep/best.
The two measured wins compound: 8.259 -> 4.8218 full-set windowed DER-CE (teacher
reproduces 2.289 in-run vs documented 2.2864) — a 42% error reduction on the 1.0
release at the same architecture and artifact size. Still misses the strict teacher+0.5pp
gate (+2.53pp; miss disclosed); the E2/E3 factorial attributes the residual to domain
The two measured wins compound: 8.259 -> 5.08 full-set windowed DER-CE (teacher
reproduces 2.289 in-run vs documented 2.2864) — a 38% error reduction on the 1.0
release at the same architecture and artifact size (CORRECTED 2026-09-05: the
published 4.8218 did not reproduce; see RESULTS.md). Still misses the strict
teacher+0.5pp gate (+2.79pp; miss disclosed); the E2/E3 factorial attributes the residual to domain
coverage.'
metrics:
- name: der_teacher_fullset
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -12,10 +12,11 @@ trained_from: 'sequence-level KD from the r7 canonical teacher (rababa_arabic_by
2.2864 windowed DER-CE full protocol): fresh greedy r7 labels on the same r5-units
corpus/limits as ara-diac-small-1.0, Muon optimizer (E3-adopted), vanilla ByT5-small
(E4, pre-registered gate <= 6.26). Checkpoint rababa-checkpoints:/rababa_arabic_distill_small/run-007-r7-muon-6ep/best.
The two measured wins compound: 8.259 -> 4.8218 full-set windowed DER-CE (teacher
reproduces 2.289 in-run vs documented 2.2864) — a 42% error reduction on the 1.0
release at the same architecture and artifact size. Still misses the strict teacher+0.5pp
gate (+2.53pp; miss disclosed); the E2/E3 factorial attributes the residual to domain
The two measured wins compound: 8.259 -> 5.08 full-set windowed DER-CE (teacher
reproduces 2.289 in-run vs documented 2.2864) — a 38% error reduction on the 1.0
release at the same architecture and artifact size (CORRECTED 2026-09-05: the
published 4.8218 did not reproduce; see RESULTS.md). Still misses the strict
teacher+0.5pp gate (+2.79pp; miss disclosed); the E2/E3 factorial attributes the residual to domain
coverage.'
metrics:
- name: der_teacher_fullset
Expand Down
9 changes: 5 additions & 4 deletions models/ara-diac-small/ara-diac-small-2.1.metadata.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -12,10 +12,11 @@ trained_from: 'sequence-level KD from the r7 canonical teacher (rababa_arabic_by
2.2864 windowed DER-CE full protocol): fresh greedy r7 labels on the same r5-units
corpus/limits as ara-diac-small-1.0, Muon optimizer (E3-adopted), vanilla ByT5-small
(E4, pre-registered gate <= 6.26). Checkpoint rababa-checkpoints:/rababa_arabic_distill_small/run-007-r7-muon-6ep/best.
The two measured wins compound: 8.259 -> 4.8218 full-set windowed DER-CE (teacher
reproduces 2.289 in-run vs documented 2.2864) — a 42% error reduction on the 1.0
release at the same architecture and artifact size. Still misses the strict teacher+0.5pp
gate (+2.53pp; miss disclosed); the E2/E3 factorial attributes the residual to domain
The two measured wins compound: 8.259 -> 5.08 full-set windowed DER-CE (teacher
reproduces 2.289 in-run vs documented 2.2864) — a 38% error reduction on the 1.0
release at the same architecture and artifact size (CORRECTED 2026-09-05: the
published 4.8218 did not reproduce; see RESULTS.md). Still misses the strict
teacher+0.5pp gate (+2.79pp; miss disclosed); the E2/E3 factorial attributes the residual to domain
coverage.'
metrics:
- name: der_teacher_fullset
Expand Down
Loading