docs: Paper B v1.1 — clean-label closure, CI frontier, width/depth axis - #155
Merged
Conversation
Slots the post-08-25 evidence: the Arabic clean-label falsification closes finding 3's corrupted-labels caveat (74.68 clean; 73.95 with every lever, CI [70.22, 71.09]); the full-set Arabic frontier gains a paired-bootstrap table (73.95 / 8.26 / 5.78 / 4.82 / 4.57 / 2.29 with disjoint intervals); the width-load-bearing vs depth-compressible finding (SVD stitch fails both ratios, verbatim depth cut ships); the E6 register-swap negative and the G2b live-test slot; catalog gains 2.0/2.1/layerdrop rows (twelve models, three tiers); leaderboard gains the 2.1 row; subset lesson updated to five instances.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What
Item 06 of TODO.publish-client: slots the post-08-25 evidence into
docs/paper.adoc (v1.0 -> v1.1).
Changes
falsification run (74.68) + the maximal-lever tiny-max run
(73.95, delta CI [70.22, 71.09], p=0) — the pretrained-or-collapse
law now rests on clean-label evidence in both task families
(adjacent-rung CIs disjoint end to end)
fails both ratios (82.96 / 78.23), verbatim enc 12->6 depth cut
ships at 5.784 [3.03, 3.49], 1.21pp CI-separated depth cost
live (cell slots in when the gate lands)
three tiers (client-lite = the ~95MiB int4 browser-budget artifact)
five-instance record with cross-ref to Paper A 3.2
Honesty guards
its public release rides the index-v3 chain this week