From caf7342e26b6edc1369dde59c07d1df28af02a94 Mon Sep 17 00:00:00 2001 From: Ronald Tse Date: Sat, 29 Aug 2026 10:09:31 +0800 Subject: [PATCH] =?UTF-8?q?docs:=20ara-diac-tiny=20clean-label=20verdict?= =?UTF-8?q?=20=E2=80=94=20from-scratch=20collapses=20(74.68)?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Closes the retraction arc; pretrained-backbone law confirmed on Arabic evidence. Browser-tier path forward: width-stitch from ByT5-small. --- docs/PUBLICATION-NOTES.md | 14 ++++++++++++++ docs/RESULTS.md | 22 ++++++++++++++++++++++ 2 files changed, 36 insertions(+) diff --git a/docs/PUBLICATION-NOTES.md b/docs/PUBLICATION-NOTES.md index 4edb439..922aa66 100644 --- a/docs/PUBLICATION-NOTES.md +++ b/docs/PUBLICATION-NOTES.md @@ -116,3 +116,17 @@ data, not same capacity law. Two lessons for the experiments section: inferred from filenames; (2) a reproduced number is only evidence of reproducibility when the input pipeline is versioned. run-005 (regenerating labels live from the r6 teacher) is the honest test. + +## ara-diac-tiny run-005: the law holds (2026-08-29) + +The clean-label rerun scores 74.68 vs the poisoned run's 83.08 (gate +3.07): corruption explained 8pp, capacity explains the rest. The +"pretrained backbone or nothing" claim now has Arabic evidence +matching the Thai ablation, closing the retraction arc +(verdict -> retraction -> poisoned rerun -> clean rerun). For the +paper: report the pair (83.08, 74.68) as the label-quality and +capacity bounds of the same from-scratch rung, with the teacher +reproducing 1.3205 across all three measurements as the harness +control. Next rung on the frontier: stitch-down from ByT5-small +pretraining (ridge-fit width bridge), which changes init, not data or +capacity alone. diff --git a/docs/RESULTS.md b/docs/RESULTS.md index 501e1a3..61a74c4 100644 --- a/docs/RESULTS.md +++ b/docs/RESULTS.md @@ -195,6 +195,28 @@ Verdict: run-004 says nothing about from-scratch capacity; the retraction stands. **run-005** (fresh teacher labels, no snapshot, 2026-08-29) is the actual clean-label falsification test — in flight. +## ara-diac-tiny run-005 — clean-label verdict: from-scratch collapses (2026-08-29) + +The falsification test the Aug-24 retraction called for: same 33M +from-scratch student (d384, 8+8), labels regenerated live from the r6 +teacher (11,793 units, no snapshot), 4,422 steps / 3 epochs, final CE +~0.9. Windowed gate, 300 SadeedDiac-25 paragraphs, teacher reproduces +1.3205: + +| Model | DER-CE (300) | +|---|---| +| Teacher (r6) | 1.3205% | +| Tiny, mojibake labels (run-004) | 83.08% | +| **Tiny, clean labels (run-005)** | **74.68% — REJECTED** | + +Clean labels recover ~8pp of the collapse — the student learns real +signal — but remains two orders off the <= 3.07 gate. **The +pretrained-backbone law now rests on Arabic evidence as well as +Thai**: from-scratch byte-level students at this width do not work. +The viable path to a sub-100MB browser tier is width reduction FROM a +pretrained ByT5-small (closed-form stitch across widths, microkimi +protocol), not from-scratch training. + ## ara-diac-small-1.0 — Arabic client tier (2026-08-24) Sequence-level KD from the r6 teacher (rababa_arabic_byt5/run-006-morph,