Add config-first ingestion and verified Modal training path - #160
Merged
ProfSynapse merged 119 commits intoSep 30, 2026
Merged
ProfSynapse merged 119 commits into
ProfSynapse merged 119 commits into
Conversation
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What this branch adds
This PR grew beyond its original ingestion-only scope. It is stacked on
feat/submodule-cloud-api-v1; the head branch remainsfeat/submodule-cloud-api-v1-windows.Verified training milestone
The normal CLI completed a fresh L40S smoke from source
df1a0d556d6d30c01ccfb8180301cfee751e04a7, attemptmodal-4c9ccfa5f57b6e4b4b0e01c6: two training steps, trainer exit 0, completed lineage, five verified downloaded artifacts, and a saved Qwen3.5-4B LoRA adapter (rank 32, alpha 64). Independent read-only inspection confirmed 181 training and 39 validation examples and a nonempty adapter with 256 LoRA tensor headers. The evidence document records exact identities, sizes, and hashes.This bounded training and artifact milestone has now been extended to one submitted training-to-vLLM evaluation job. A fresh L40S rehearsal on source
6e3f4f142a5778f73b3c866bb5775475a2becca8completed two optimizer steps in 131.1 seconds, saved the LoRA adapter, reused the prepared base snapshot for serving, passed all three same-job evaluation cases, and downloaded all five verified training artifacts. Independent archive inspection confirmed rank-32 LoRA with 256 A/B tensors and adapter/evaluation digest agreement.The generic config now supports deliberate training/serving chat-template options and an omitted request output-token ceiling (
max_tokens: null). This Qwen recipe disables thinking consistently. All three saved replies were independently read as complete, relevant output without visible thinking text; natural-stop assertions passed. This is workflow qualification, not chapter-style, factual-quality, or long-context writing qualification after two steps. Full training and GGUF remain on hold.Exact-head CI exposed two independent-chat integration failures when new shared startup fields were not explicitly selected there. Commit
352a6a2bc8858325805c0d3cb12e9cea89ad6e03preserves existing behavior by choosing the existing defaults, retains the full-field regression assertion, and refreshes only affected inference source hashes. Independent review accepted the change; 17 focused integration tests and a broader 182-case Linux suite passed. All four runtime/closure commitment checks are CURRENT. Fresh CI on this commit must be green before merge; the live rehearsal evidence belongs to the earlier exact source above, not an invented new GPU run.Review map
The diff is large; review these boundaries and the current CI results before merging.