PA-8: judge whether this machine can run a local model, and say what is short - #65
Merged
Merged
Conversation
… what is short The two halves that decide whether this feature is honest. The surface that shows a running download is deliberately not here. Measuring first shrank the item: the RUN half already exists. A local runtime is already an allowed provider -- the one trusted OpenAI-compatible package, local URL only -- and the model listing already asks it what it serves. And the probe is not native: the engine is a process on the user's own machine, so memory, cores and arch come straight from the OS. Only the GPU needs a shell out. `local-model-fit.ts` is built on three rules. An unseen GPU is reported as NOT MEASURED, never as absent. Telling a user with a discrete card that their machine has none is a false statement about their own computer, and it is the one they would push back on hardest and be right about. So `undefined` means we did not look, which is a third answer. It never invents a model's size. The figure comes from whoever is offering the model. A fit check against a size this module guessed would be a guess wearing a verdict's clothes. A refusal carries the numbers that produced it, and distinguishes the two cases that matter: short by an amount this machine has if something is closed, versus short by more than it has in total. "Your machine cannot run this" is not actionable and cannot be checked by the person it is about. Memory is checked against AVAILABLE rather than total, and `os.freemem()` understates on Linux because it counts reclaimable page cache as used. That direction is deliberate: it can refuse a model that would in fact have squeezed in, and it will not promise one that gets killed mid-generation. The opposite error loses the user's work. The runtime's own listing and fetch shapes are in `local-model-pull.ts`, because the OpenAI-compatible `GET /models` the provider listing uses reports an id and nothing else -- no size -- so it cannot answer the question this feature turns on. We do not fetch weights ourselves: no mirror choice, no quantisation choice, no digest we chose to trust. The runtime already resumes interrupted downloads and shares progress between concurrent callers; re-implementing that would mean owning model provenance, which is a much larger promise than running a model locally. Shapes come from the runtime's published API reference rather than from reading our own code back, and two traps it states explicitly have a test each: `completed` may be absent while a layer is downloading, and `total` is per-LAYER not per-model -- a caller reading it as whole-model progress shows a bar that resets on every layer. An in-band `error` frame arrives after a 200 is already sent and is terminal, not progress; treated as progress the pull looks like it is still running and never finishes. There is also a test asserting that a whole-body JSON parse FAILS on the real stream, so nobody simplifies the line-by-line reader back into one parse. core 1034 tests 0 fail, 36 new. The fit judgement is pure and needs no machine. The GPU parsers run against the reference's own recorded output, so they are verified on a host with no GPU. No cloud fallback exists here and none should be added. A user who asked for a local model asked for the data to stay on the machine; answering from a hosted model would defeat the only reason to want this, and would do it invisibly.
…a 404 it found The fit judgement and the probes had no caller. That is the failure this queue keeps finding -- an arming matcher nobody invoked, a model preference read and discarded, builtin skills registered into a list with no consumer -- so the surface is the point of this change, not an extra. `redrob models-local` prints the machine, then per configured local endpoint every installed model with either "runs here" or the refusal and its numbers. DRIVING IT END TO END FOUND A REAL DEFECT, and nothing short of that run would have. A local runtime is configured with its OpenAI base, conventionally `/v1`, but `/api/tags` is a SIBLING of `/v1`, not a child. Joining onto the configured base produced `/v1/api/tags`, got a 404, and the command printed "answering, with nothing installed" against a server serving two models. A 404 and an empty list are both "no models" to a caller that does not separate them, so the bug read as a working feature. Fixed, with a regression test that says why, and reverse-verified: removing the fix fails it. Only a TRAILING version segment is stripped. A runtime served under a prefix by a reverse proxy keeps its path, which is why `/v1/engine` is left alone. Measured on this host (16 cores, 61.8 GiB): a 2 GiB model runs here; a 120 GiB model is refused with "needs about 144.0 GiB of memory (120.0 GiB of weights plus runtime headroom) and 38.9 GiB is free of 61.8 GiB total -- short by 105.1 GiB, more than this machine has in total". Top-level `models-local` rather than `models local`, because `models` takes a positional provider filter and a subcommand there is ambiguous with a provider named `local`. The GPU probe's failure is swallowed and reported as NOT MEASURED, never as "no GPU". On a machine with no discrete card the tool simply does not exist, which is the normal state of the world rather than a condition to report -- and the memory judgement, the part that actually decides whether a model runs, does not depend on it. HttpClient and AppProcess are provided by this command rather than added to the CLI runtime. It is the only command that speaks HTTP to a local runtime or shells out for a GPU, and widening the shared runtime would hand both to every other command that has no use for them. core 1031 tests 0 fail across 123 files. redrob cli 380 tests 0 fail. Generated SDK clean, tsc clean.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The two halves of PA-8 that decide whether the feature is honest. The surface that shows a
running download is deliberately not here — this does not close the item.
Measuring first shrank it
The run half already exists: a local runtime is already an allowed provider (the one
trusted OpenAI-compatible package, local URL only) and the model listing already asks it
what it serves. And the probe is not native — the engine is a process on the user's own
machine, so memory, cores and arch come straight from the OS. Only the GPU needs a shell
out.
Three rules the fit module is built on
An unseen GPU is NOT MEASURED, never absent. Telling a user with a discrete card that
their machine has none is a false statement about their own computer.
undefinedis a thirdanswer, and the refusal text is asserted never to mention a GPU it did not look at.
It never invents a model's size. The figure comes from whoever offers the model. A fit
check against a size this module guessed would be a guess wearing a verdict's clothes.
A refusal carries its numbers, and separates the two cases that matter:
short by 4.0 GiB, which this machine has if you close somethingshort by 4.0 GiB, more than this machine has in totalMemory is checked against available, not total, and
os.freemem()understates on Linux.That direction is chosen: it can refuse a model that would have squeezed in, and it will not
promise one that gets killed mid-generation. The opposite error loses the user's work.
Why a second set of routes
The OpenAI-compatible
GET /modelsthe provider listing uses reports an id and nothingelse — no size — so it cannot answer the question this feature turns on. Sizes and the
fetch live on the runtime's own API.
We do not fetch weights ourselves. No mirror choice, no quantisation choice, no digest
we chose to trust. The runtime already resumes interrupted downloads and shares progress
between callers; re-implementing that means owning model provenance, a far bigger promise
than "run a model locally".
Shapes come from the runtime's published API reference, not from reading our own code back.
Two traps it states explicitly have a test each:
completedmay be absent while a layer downloads.totalis per-layer, not per-model — read as whole-model progress it shows a bar thatresets on every layer.
errorframe arrives after a 200 is already sent, so it is terminal; treated asprogress the pull looks alive and never finishes.
JSON.parsefails on the real stream, so nobodysimplifies the line reader back into one parse.
Verification
tscclean for core and redrobreference's own recorded output, so they are verified on a host with no GPU
No cloud fallback exists here and none should be added: a user who asked for a local model
asked for the data to stay on the machine.