@@ -6,28 +6,34 @@ target: Arab
66tokenizer : bytes
77opset : 14
88decoder : kv
9- precision : fp32
9+ precision : int8
1010license : BSD-3-Clause
1111trained_from : ' sequence-level KD from the r7 canonical teacher (rababa_arabic_byt5/run-007-news/best,
1212 2.2864 windowed DER-CE full protocol): fresh greedy r7 labels on the same r5-units
1313 corpus/limits as ara-diac-small-1.0, Muon optimizer (E3-adopted), vanilla ByT5-small
14- (E4, pre-registered gate <= 6.26). Checkpoint
15- rababa-checkpoints:/rababa_arabic_distill_small/run-006-r7-muon/best. The two
16- measured wins compound: 8.259 -> 4.8218 full-set windowed DER-CE (teacher reproduces
17- 2.289 in-run vs documented 2.2864) — a 42% error reduction on the 1.0 release at the
18- same architecture and artifact size. Still misses the strict teacher+0.5pp gate
19- (+2.53pp; miss disclosed); the E2/E3 factorial attributes the residual to domain
14+ (E4, pre-registered gate <= 6.26). Checkpoint rababa-checkpoints:/rababa_arabic_distill_small/run-006-r7-muon/best.
15+ The two measured wins compound: 8.259 -> 4.8218 full-set windowed DER-CE (teacher
16+ reproduces 2.289 in-run vs documented 2.2864) — a 42% error reduction on the 1.0
17+ release at the same architecture and artifact size. Still misses the strict teacher+0.5pp
18+ gate (+2.53pp; miss disclosed); the E2/E3 factorial attributes the residual to domain
2019 coverage.'
2120metrics :
2221- name : der_teacher_fullset
2322 value : 2.289
24- protocol : windowed DER-CE (1400-byte windows, word-boundary split, greedy,
25- haraqat-projected, Misraj evaluator); full 1,200-paragraph SadeedDiac-25;
26- in-run reproduction of the documented 2.2864 (r7 canonical teacher)
23+ protocol : windowed DER-CE (1400-byte windows, word-boundary split, greedy, haraqat-projected,
24+ Misraj evaluator); full 1,200-paragraph SadeedDiac-25; in-run reproduction of
25+ the documented 2.2864 (r7 canonical teacher)
2726 source : interscript/interscript-ml docs/RESULTS.md#run-006-r7-muon
2827- name : der_student_fullset
2928 value : 4.8218
30- protocol : same full-set harness; E4 (r7 teacher labels + Muon, vanilla
31- ByT5-small) vs the 8.259 AdamW/r6-labels 1.0 release — a 42% error
32- reduction at identical architecture and artifact size
29+ protocol : same full-set harness; E4 (r7 teacher labels + Muon, vanilla ByT5-small)
30+ vs the 8.259 AdamW/r6-labels 1.0 release — a 42% error reduction at identical
31+ architecture and artifact size
3332 source : interscript/interscript-ml docs/RESULTS.md#run-006-r7-muon
33+ parity :
34+ samples : 600
35+ cer_delta : 0.1393
36+ sha256 :
37+ decoder-kv.onnx : 324e8df388538f7ab22eae88cf78b9bcfe45a7edc704fa4864397014d00e38b8
38+ decoder.onnx : d31d1c5329e16a583372a0679651d0af37e93b9fbb22e3cd7b49ab8bc0c6a40d
39+ encoder.onnx : 9079b8058216af6d7475d63c8286c80cc2cbe701c9eba179f569c6c84d8d750a
0 commit comments