diff --git a/README.md b/README.md index b76e663..d39d295 100644 --- a/README.md +++ b/README.md @@ -1,125 +1,71 @@ -# interscript/ml-models +# interscript-ml -Unified training framework for Interscript ML-powered maps. +The **contract** for Interscript's phonological layer — the normative +definition of what a "hidden reading" model is, and the zoo that +publishes models conforming to it. -**Status:** skeleton (P0 in TODO.rababa/10). Framework abstractions -implemented and tested. Task data modules + configs ready; full -training requires GPU + dataset fetch. +This repo owns three things and nothing else: -## What this is +1. **The `models.yaml` index** — the stable URL every runtime resolves + model ids against (with per-artifact sha256s and split-part support). +2. **The IMF v1 model-zip format** — the artifact contract: + `metadata.yaml` + ONNX graphs + member sha256 manifest. The normative + text is [SPEC.md](SPEC.md); the reference loader now lives in the + [Python crystal](https://github.com/secryst/secryst-py). +3. **The model zoo + publish pipeline** — teachers from + [interscript-ml-train](https://github.com/interscript/interscript-ml-train) + are distilled, gated (parity written into the artifact), and released + as index entries here. -One training repo for every ML map in Interscript: +## The system -- **rababa_arabic** — Arabic diacritization (adds harakat) -- **rababa_hebrew** — Hebrew diacritization (adds nikud) -- **secryst_thai_ipa** — Thai → IPA transliteration - -Each task is a **config + data module**. The training framework is -shared. Adding a new transliteration pair (Khmer → IPA, Japanese → -Romaji) is one new directory under `src/tasks/` — zero edits to -framework code. - -## Architecture - -``` -src/ -├── framework/ # SHARED abstractions (MECE) -│ ├── config.py # TaskConfig loaded from YAML -│ ├── registry.py # Plugin registry (OCP) -│ ├── data.py # DataModule ABC -│ ├── model.py # ModelModule ABC (teacher + student) -│ ├── trainer.py # BaseTrainer + FineTune + Distill (DRY) -│ ├── evaluator.py # BaseEvaluator + edit_distance + DER/PER utils -│ ├── exporter.py # OnnxExporter ABC -│ └── pipeline.py # TrainingPipeline orchestrator -├── tasks/ -│ ├── rababa_arabic/ # config.yaml + data.py + student.py + metrics.py -│ ├── rababa_hebrew/ -│ └── secryst_thai_ipa/ -└── cli.py # python -m src.cli train --task rababa_arabic -``` - -## Design principles (project conventions) - -- **OCP** — adding a task = one new directory. Adding a metric, model - architecture, or data source = one new subclass + `@register_*` - decorator. Framework code is never edited. -- **MECE** — each module owns one concern. Data has no knowledge of - model architecture. Model has no knowledge of trainer. Trainer has - no knowledge of evaluator. -- **DRY** — the epoch loop, edit-distance math, and ONNX export - wrapper are written once. -- **Model-driven, semantically-driven** — class names mirror domain - concepts (`RababaArabicData`, `DEREvaluator`, `SecrystThaiIpaStudent`). -- **Performance** — frozen dataclasses for config; lazy imports for - torch so framework tests run without GPU deps. - -## Quick start - -```bash -scripts/setup_env.sh # creates .venv, installs deps -scripts/fetch_data.sh # fetch raw datasets (set env vars first) -scripts/train.sh rababa_arabic # full training pipeline -scripts/export.sh rababa_arabic # export student to ONNX -scripts/publish.sh rababa_arabic # upload to HuggingFace Hub ``` - -Or via the CLI directly: - -```bash -python -m src.cli list -python -m src.cli train --task rababa_arabic --data-root data --out-root models -python -m src.cli evaluate --task secryst_thai_ipa -python -m src.cli export --task rababa_hebrew +interscript deterministic transliteration maps + engines +(ruby · js · py) │ maps that need vocalization dispatch to a + │ crystal through stdlib adapters (optional) + ▼ +secryst crystals Ruby gem · pip install secryst · npm i secryst +(secryst org) implement IMF v1 + models.yaml — nothing else + │ + ▼ +interscript-ml ◄────── models/zips resolve through this index +(THIS repo) ──────► golden sets: crystals diffed against each other + +interscript-ml-train teachers (arabic · persian · urdu + hebrew docs); +(interscript org) students distilled here enter the zoo above ``` -## Adding a new task - -1. Create `src/tasks//config.yaml` (copy from an existing task). -2. Create `src/tasks//data.py` extending `DataModule`, decorated - with `@register_data_module("_data")`. -3. Create `src/tasks//student.py` extending `ModelModule`, - decorated with `@register_model_module("_student")`. -4. Create `src/tasks//metrics.py` extending `BaseEvaluator`, - decorated with `@register_evaluator("")`. -5. Run `python -m src.cli train --task `. - -That's it. No framework edits. - -## Tests - -```bash -pytest -v -``` - -Framework tests run without torch (CPU-only, fast). Training and ONNX -export tests are gated behind `@pytest.mark.gpu` and require the -`[train]` and `[export]` extras. - -## Distribution +Dependency directions, stated once: -Models ship as **IMF v1** zips (Interscript Model Format — spec in -[`docs/imf-v1.md`](./docs/imf-v1.md)): byte-level tokenizer only, ONNX -opset 14, sha256-verified graphs, metrics traceable to `RESULTS.md` -anchors. Build/validate with `PYTHONPATH=src python -m imf pack|validate`. +- **interscript-ml depends on nothing.** It is the contract: an index, + a format, golden sets, and release tooling. +- **Crystals depend only on the contract.** A crystal has zero + interscript-core dependency — a TTS front-end can phonemize Khmer + with `pip install secryst` and nothing else. +- **Engines depend on crystals only optionally.** An engine without a + crystal simply cannot execute maps that declare a vocalization step. +- **Training depends on nothing downstream.** Teachers never import the + contract; they're consumed by it (via the export gate). -Models reach end users through three channels (full plan in -[`TODO.distribution/`](./TODO.distribution/)): +## Repositories -| Channel | Audience | Why | -|---|---|---| -| **GitHub Releases** (primary) | All consumers | Versioned, immutable, checksums, tied to source tags | -| **HuggingFace Hub** (canonical) | Researchers | Model cards, datasets, auto-conversion, inference API | -| **jsdelivr CDN** (edge) | Browser | Edge-cached, CORS-friendly, no rate limits | +| repo | role | +|---|---| +| [interscript/interscript-ml](https://github.com/interscript/interscript-ml) | this — contract + zoo | +| [secryst/secryst](https://github.com/secryst/secryst) | Ruby crystal (the original, est. 2020) | +| [secryst/secryst-py](https://github.com/secryst/secryst-py) | Python crystal — reference, owns golden generation | +| [secryst/secryst-ts](https://github.com/secryst/secryst-ts) | TypeScript crystal (npm `secryst`) | +| [secryst/secryst.github.io](https://www.secryst.org) | the crystals' documentation site | +| [interscript/interscript-ml-train](https://github.com/interscript/interscript-ml-train) | training monorepo (arabic/persian/urdu) | +| [interscript/rababa](https://github.com/interscript/rababa) · [rababa-farsi](https://github.com/interscript/rababa-farsi) · [rababa-urdu](https://github.com/interscript/rababa-urdu) | archived origins of the train monorepo (full history merged there) | -Per-task versioning: `rababa_arabic-v1.0.0`, `secryst_thai_ipa-v1.2.0`, -etc. Each release ships fp32 + int8 + int4 variants with SHA256 -sidecars, SLSA provenance, and Sigstore signatures. +`runtime/` in this repo is the **frozen origin** of the Python crystal — +kept for provenance; live code and releases are in secryst-py. -Distribution phases (P2–P8) are tracked in `TODO.distribution/`. The -first production release lands when phase P6 (first trained model) -completes. +## Environment (as implemented by every crystal) -## License +`SECRYST_INDEX` (index URL or path; default: `models.yaml` on this +repo's main) · `SECRYST_CACHE` (default `~/.cache/secryst`). Cache hits +are re-verified against the index on every load. -BSD-3-Clause, for code and model weights alike (see `LICENSE`). +License: BSD-3-Clause. diff --git a/SPEC.md b/SPEC.md new file mode 100644 index 0000000..0b48976 --- /dev/null +++ b/SPEC.md @@ -0,0 +1,106 @@ +# interscript-ml contract — Specification (v1) + +Normative. Key words MUST / MUST NOT / SHALL / SHOULD / MAY are to be +interpreted as described in RFC 2119. This file is the canonical text; +the [crystals' documentation site](https://www.secryst.org/spec.html) +renders the same content for users. + +## Conformance + +An implementation conforms to interscript-ml v1 when it: + +- **C1** — MUST resolve model ids against a `models.yaml` index of + `version: 1`, honoring `SECRYST_INDEX`. +- **C2** — MUST verify the whole-artifact sha256 before installation + and re-verify every cache hit; a mismatch MUST fail loudly. +- **C3** — MUST verify every `.onnx` member against the manifest + sha256 map on load; members not covered by the manifest MUST NOT + load. +- **C4** — MUST implement the byte tokenizer exactly (§4); id + sequences MUST NOT be treated as raw bytes. +- **C5** — MUST produce byte-identical outputs to the reference + crystal on the shared golden sets. +- **C6** — SHOULD install artifacts atomically (temp file + rename). + +## §1 The models.yaml index + + version: 1 + models: + : + filename: string # artifact file name + url: string # single-file channel (http(s):// or file://) + sha256: string # whole-artifact digest + size: int + precision: fp32 | fp16 # default fp32 + task: string + parts: # OPTIONAL: split artifacts + - url: string + sha256: string # per-part digest, verified as it lands + size: int + +The `parts` mechanism exists for artifacts exceeding GitHub's 2 GiB +per-asset cap: parts stream into one file in index order, each verified +on arrival; the assembled file is then checked against the entry-level +`sha256` exactly as a single-file model — the cache contract is +identical. + +## §2 Resolution algorithm + +1. Resolve `` in `models.models`. Unknown ids MUST raise an error + enumerating known ids. +2. If a cached copy exists at `/models//` whose + whole-file sha256 matches, use it — cache hits are re-verified, + never trusted blindly. +3. Otherwise download (single URL, or parts in order) to a temporary + file in the target directory, verifying digests as data lands. +4. Verify the assembled artifact against the index sha256. +5. Atomically rename into place, then load. + +## §3 IMF v1 model zips + +A zip containing at minimum `metadata.yaml`, `encoder.onnx`, and +`decoder.onnx`. Zips MUST NOT rely on zip-level integrity; integrity +is the manifest's job. + + format: imf-v1 + tokenizer: bytes + id: + task: string + decoder: plain | kv # kv iff decoder-kv.onnx is present + precision: fp32 + opset: 14 + sha256: # every .onnx member MUST be covered + encoder.onnx: + decoder.onnx: + +A conforming loader MUST reject: any other `format`, any `tokenizer` +other than `bytes`, missing required members, and uncovered or +mismatched member digests. + +## §4 Byte tokenizer + +| concept | rule | +|---|---| +| encoding a string | UTF-8 bytes `b` → ids `b + 3`, then one trailing `eos` | +| decoding ids | stop at `eos`; skip `pad`/`unk`; `(id − 3) mod 256` per byte; reassemble as UTF-8 | +| special ids | `pad = 0`, `eos = 1`, `unk = 2` | + +Warning — the classic silent-garbage bug: ids are offset by 3 and carry +a trailing EOS. Feeding `text.bytes` directly, or forgetting the EOS, +produces plausible-but-wrong outputs that pass shape checks. Interop +tests MUST cover both encode and decode round-trips, including +multi-byte scripts. + +## §5 Environment + +| variable | meaning | default | +|---|---|---| +| `SECRYST_INDEX` | index URL or local path | `models.yaml` on this repo's main | +| `SECRYST_CACHE` | artifact cache directory | `~/.cache/secryst` | + +## Conformance kits + +- Golden parity kit (deterministic decode-loop fixture + reference + goldens): [secryst-py/parity](https://github.com/secryst/secryst-py/tree/main/parity). +- Export gate (parity written into released artifacts): this repo's + release pipeline (`src/imf/`, WO03). diff --git a/runtime/README.md b/runtime/README.md index a8b47fc..0bca2e0 100644 --- a/runtime/README.md +++ b/runtime/README.md @@ -11,7 +11,7 @@ Home repo: https://github.com/secryst/secryst-py (this copy in ml-models/runtime is the frozen origin; the package now lives there). ```python -from interscript_ml import Model +from secryst import Model model = Model.load("khm-latn-1.0") # id: index resolve -> download # -> sha256-verify -> cache -> load