Training that uses the weight budget as a capacity-achieving code can make calibration machinery unnecessary; this lab measures when that statement holds.
- Masking a 30B-class MoE to its top-45.3%-demand experts beats the full model on the lab's mathematics gate (+14.7 pooled, 6/6 paired seeds); random masks at the same fraction score 0/120 — selection, not sparsity (FINDINGS, routing crest).
- Capability follows expert identity, not any aggregate of the keep-set: with recall and coverage pinned, exchanging 4 experts per layer per class between matched keep-set draws moved 23–52 gate solves per side and largely inverted the pair outcomes; recall- and coverage-organization both returned registered nulls first.
- A transformer birth (multi-block, real math diet, 1000 steps) runs bit-identically across Mac CPU, RTX 3080 CUDA, and an external lab's C++ reimplementation — integer training with sha-equality as the acceptance bar, reproducible by one command (REPRODUCE).
These are measured, fenced claims — the tags and scope vocabulary in FINDINGS are part of each statement, not optional reading.
Start with the glossary, curated findings — organized by evidence maturity and scope rather than chronology, one tag per claim — the reproduction walkthrough, and the living verdict ledger. The measured-history appendix keeps supported earlier mathematics and physics results without making this page chronological.
Citation policy: name the exact repository commit SHA and the exact verdict
entry in docs/RESULTS.md that supports the claim. The
ledger is living, so an unpinned citation is not reproducible. Repository
metadata is in CITATION.cff.
llmopt is a mathematics and physics research lab organized around
executable instruments, declared comparisons, and oracle-checked readouts.
The main paths are:
llmopt/search/— symbolic derivation search with explicit rewrite rules, structural and learned evaluators, proposal policies, transposition memory, and verification at the boundary.llmopt/mathgen/— seeded generators for calculus, linear algebra, ordinary differential equations, mechanics, and proofs, with symbolic checks built into generation or evaluation.llmopt/quantum/— model-Hamiltonian ground-state instruments and a verified ZX-graph search path for circuit reduction.llmopt/train/andllmopt/intmath.py— closed-system model births, exact integer primitives, controlled diets, and training interventions that can be compared trajectory by trajectory.llmopt/quantize/— weight diagnostics, sensitivity probes, closed-form allocation, packed crystal artifacts, and the capacity meter that selects which allocation regime deserves testing.llmopt/eval/— equivalence, calibration, latency, and statistical instruments used as supporting readouts rather than substitutes for capability gates.
Experiments are recorded as named arms and cells. A gate accepts or rejects a
declared result; a pin fixes an artifact or trajectory contract. Those words,
the maturity labels, and the controlled scope vocabulary are defined in
GLOSSARY.md.
Install the editable package and development dependencies with:
python -m pip install -e ".[dev]"The status board, theory map, riff ledger, handoffs, and machine-readable
index remain living authority surfaces: docs/BOARD.md,
docs/THEORY.md,
docs/RIFF-LEDGER.md,
docs/handoffs/, and
docs/results-index.jsonl.
-
[SINGLE-SEED] [REGIME-SCOPED: at-capacity house crystals]The house crystals tested at capacity could be packed from weight sigma alone, with no calibration data or allocator search, while remaining inside their declared capability gate. This is a Mac MPS,n=1-per-crystal result, not a universal quantization rule: PACKED CRYSTAL C0+C1 VERDICT indocs/RESULTS.md. -
[SINGLE-SEED] [DEVICE-SCOPED] [FORMAT-BOUND] [REGIME-SCOPED: Qwen2.5-0.5B]THE SIGMA-PACK CLAIM DOES NOT TRANSPORT TO Qwen2.5-0.5B. On the registered RTX 3080 fake-quant arm, max-anchored and calibrated grids exploited weight-tail structure that the house crystals did not have. The boundary is one model, one device, andn=1; it is not a law of web-trained dense models: PACKED CRYSTAL C6 VERDICT and PACKED CRYSTAL C6c VERDICT indocs/RESULTS.md. -
[REPLICATED] [DEVICE-SCOPED] [FORMAT-BOUND] [REGIME-SCOPED: measured deployment artifacts]The packed integer GEMM path reproduced its digest across the registered Mac MPS and RTX 3080 CUDA devices. This independent-device route covers one house crystal's integer GEMM path, not a full integer end-to-end decoder: PACKED CRYSTAL C4 VERDICT and R-PASS VERDICT indocs/RESULTS.md. -
[NULL] [DEVICE-SCOPED] [FREE-RUN-GATED] [REGIME-SCOPED: specified diet and recipe]Removing scaffold-token loss did not repair the registered gravmoe capability gate; the arm degraded while its format diagnostics remained intact. This is a Mac-localn=1null: VERDICT SOL-ADOPTION-1 indocs/RESULTS.md. -
[REPLICATED] [FORMAT-BOUND] [FREE-RUN-GATED] [REGIME-SCOPED: measured deployment artifacts]A resident 30B-class MoE scored HIGHER on the registered mathematics gate with 45.3% of its experts masked out than at full width, +14.7 pooled across six paired seeds. The effect is selection, not sparsity: random and anti-demand masks at the identical fraction scored0/120. It is also domain-specific — the same recipe on mechanics returned+3and then-59pooled, with open and closed recall matched to the mathematics arm to within0.0001and0.003, so coverage and recall do not predict the sign of the effect. The scope is one vehicle, one keep rule, one gate, and Mac MLX: VERDICT MOE-GT-1-R5, VERDICT MOE-GT-1-R6, and VERDICT MOE-GT-2-D4-PHYS-B indocs/RESULTS.md.
The broader closed-system record, including positive results, nulls,
retractions, and amendments, is curated by evidence maturity in
docs/FINDINGS.md. Historical mathematics and physics
measurements that are still supported but are not part of that main arc are in
docs/MEASURED-HISTORY.md.
RJOB_LOCAL=1 python -m llmopt.reproduce gravmoe-rb1On the current Mac-local contract, PASS means the final training-trajectory
digest exactly matches the committed gravmoe-rb1 pin. The adopted command is
booked by VERDICT SOL-ADOPTION-1 in docs/RESULTS.md;
repository commit 4ef9cd511369023d69db7332aebf36517de62951 made the path self-contained from committed windows.
Trajectory agreement is not oracle correctness. It certifies the pinned weight path and teacher-forced training readouts. Free-run symbolic solve scoring additionally needs the uncommitted diet row text and its oracle; the artifact-backed gate arms therefore run in an explicit trajectory-only mode.
The walkthrough in docs/REPRODUCE.md explains setup,
expected output, the full pin registry, and where oracle-backed correctness
claims live. Use python -m llmopt.reproduce --list to inspect the registry;
do not interpret digest equality as a capability score.
The calibration-free packing law is established only for the measured
at-capacity house crystals. The negative PACKED CRYSTAL C6 VERDICT is
equally narrow: Qwen2.5-0.5B, one registered RTX 3080 device, fake quantization,
and n=1. More models, independent implementations, and capability-gated
deployment artifacts are needed before either boundary can move.
Many training comparisons remain single-seed and device-scoped. The README
keeps those fences visible; docs/FINDINGS.md carries the
current maturity labels, and docs/BOARD.md separates live
work from closed results.
Why masking a deployed MoE to its demand coalition beats full width on mathematics is unexplained, and the two quantities a keep rule optimizes — coverage and recall of demanded experts — were measured not to predict even the sign of that effect. One piece is now measured, and a same-night control inverted it. A matched-size random fill resurrected a dead core about as well as the verbal-branch fill (51/36/48 versus 55 of 120), which read as coverage-not-class — but that random pool was itself about 45 percent verbal-branch experts. Fills that exclude the verbal branch at the same recall score 0 and 7 of 120, against 16 to 55 for fills that include it, and a recall ladder over random fills is non-monotone (a 0.795-recall draw scores 9; a 0.792 draw scores 66). So the verbal population is necessary for the resurrection and recall does not organize it; what is sufficient is still unmeasured, with two named anomalies open. Why the crest beats full width remains unexplained, and the crest stays a booked observation on one vehicle and one domain, not a deployment recommendation.
Trajectory reproduction still stops short of self-contained free-run oracle scoring because the row text is not committed. Until that input can be shared lawfully, the public artifact proves the trajectory and teacher-forced readouts only.
The mathematics and physics charter is intentionally narrow. Candidate work
belongs in docs/RIFF-LEDGER.md; uncertainty becomes a
claim only after a declared arm, a named verdict, and the applicable regime,
device, format, and replication fences are booked.
Licensed under Apache-2.0.