Skip to content

Add config-first ingestion and verified Modal training path - #160

Merged
ProfSynapse merged 119 commits into
feat/submodule-cloud-api-v1from
feat/submodule-cloud-api-v1-windows
Sep 30, 2026
Merged

ProfSynapse merged 119 commits into
feat/submodule-cloud-api-v1from
feat/submodule-cloud-api-v1-windows

Conversation

@ProfSynapse

@ProfSynapse ProfSynapse commented Sep 19, 2026 •

Copy link
Copy Markdown
Owner

What this branch adds

  • A source-agnostic, config-driven Markdown ingestion API and CLI, with optional YAML frontmatter, deterministic normalized bundles, and dataset preparation.
  • A pinned Qwen3.5-4B prompt/completion SFT recipe and model profile for the prepared 220-row dataset.
  • Packaged Modal training behind the public TrainingAPI: isolated trainer, exact source and runtime commitments, private model preparation, durable artifact publication, and verified artifact streaming.
  • Scoped tests, runtime lock updates, operational guidance, and the investigation record in docs/review/modal-volume-root-diagnostic-20260927.md.

This PR grew beyond its original ingestion-only scope. It is stacked on feat/submodule-cloud-api-v1; the head branch remains feat/submodule-cloud-api-v1-windows.

Verified training milestone

The normal CLI completed a fresh L40S smoke from source df1a0d556d6d30c01ccfb8180301cfee751e04a7, attempt modal-4c9ccfa5f57b6e4b4b0e01c6: two training steps, trainer exit 0, completed lineage, five verified downloaded artifacts, and a saved Qwen3.5-4B LoRA adapter (rank 32, alpha 64). Independent read-only inspection confirmed 181 training and 39 validation examples and a nonempty adapter with 256 LoRA tensor headers. The evidence document records exact identities, sizes, and hashes.

This bounded training and artifact milestone has now been extended to one submitted training-to-vLLM evaluation job. A fresh L40S rehearsal on source 6e3f4f142a5778f73b3c866bb5775475a2becca8 completed two optimizer steps in 131.1 seconds, saved the LoRA adapter, reused the prepared base snapshot for serving, passed all three same-job evaluation cases, and downloaded all five verified training artifacts. Independent archive inspection confirmed rank-32 LoRA with 256 A/B tensors and adapter/evaluation digest agreement.

The generic config now supports deliberate training/serving chat-template options and an omitted request output-token ceiling (max_tokens: null). This Qwen recipe disables thinking consistently. All three saved replies were independently read as complete, relevant output without visible thinking text; natural-stop assertions passed. This is workflow qualification, not chapter-style, factual-quality, or long-context writing qualification after two steps. Full training and GGUF remain on hold.

Exact-head CI exposed two independent-chat integration failures when new shared startup fields were not explicitly selected there. Commit 352a6a2bc8858325805c0d3cb12e9cea89ad6e03 preserves existing behavior by choosing the existing defaults, retains the full-field regression assertion, and refreshes only affected inference source hashes. Independent review accepted the change; 17 focused integration tests and a broader 182-case Linux suite passed. All four runtime/closure commitment checks are CURRENT. Fresh CI on this commit must be green before merge; the live rehearsal evidence belongs to the earlier exact source above, not an invented new GPU run.

Review map

  1. Public ingestion and dataset preparation contracts, then the generic CLI.
  2. Pinned SFT recipe, text tokenizer output, and prepared input handling.
  3. Modal release, source/model commitments, isolated execution, artifact publication, and host verification.
  4. Tests, locked runtime inventories, skills, and investigation notes.

The diff is large; review these boundaries and the current CI results before merging.

@ProfSynapse ProfSynapse changed the title Add source-agnostic Markdown ingestion API and CLI Add config-first ingestion and verified Modal training path Sep 29, 2026
@ProfSynapse
ProfSynapse merged commit f78874a into feat/submodule-cloud-api-v1 Sep 30, 2026
4 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant