Skip to content

PA-8: judge whether this machine can run a local model, and say what is short - #65

Merged
its-janghoon merged 2 commits into
developfrom
feature/pa8-local-model-fit
Oct 3, 2026
Merged

its-janghoon merged 2 commits into
developfrom
feature/pa8-local-model-fit

Conversation

@its-janghoon

Copy link
Copy Markdown
Contributor

The two halves of PA-8 that decide whether the feature is honest. The surface that shows a
running download is deliberately not here — this does not close the item.

Measuring first shrank it

The run half already exists: a local runtime is already an allowed provider (the one
trusted OpenAI-compatible package, local URL only) and the model listing already asks it
what it serves. And the probe is not native — the engine is a process on the user's own
machine, so memory, cores and arch come straight from the OS. Only the GPU needs a shell
out.

Three rules the fit module is built on

An unseen GPU is NOT MEASURED, never absent. Telling a user with a discrete card that
their machine has none is a false statement about their own computer. undefined is a third
answer, and the refusal text is asserted never to mention a GPU it did not look at.

It never invents a model's size. The figure comes from whoever offers the model. A fit
check against a size this module guessed would be a guess wearing a verdict's clothes.

A refusal carries its numbers, and separates the two cases that matter:

Case What the user is told
short, machine has the RAM short by 4.0 GiB, which this machine has if you close something
short, machine does not short by 4.0 GiB, more than this machine has in total
unsupported CPU named before memory is mentioned at all

Memory is checked against available, not total, and os.freemem() understates on Linux.
That direction is chosen: it can refuse a model that would have squeezed in, and it will not
promise one that gets killed mid-generation. The opposite error loses the user's work.

Why a second set of routes

The OpenAI-compatible GET /models the provider listing uses reports an id and nothing
else
— no size — so it cannot answer the question this feature turns on. Sizes and the
fetch live on the runtime's own API.

We do not fetch weights ourselves. No mirror choice, no quantisation choice, no digest
we chose to trust. The runtime already resumes interrupted downloads and shares progress
between callers; re-implementing that means owning model provenance, a far bigger promise
than "run a model locally".

Shapes come from the runtime's published API reference, not from reading our own code back.
Two traps it states explicitly have a test each:

  • completed may be absent while a layer downloads.
  • total is per-layer, not per-model — read as whole-model progress it shows a bar that
    resets on every layer.
  • An in-band error frame arrives after a 200 is already sent, so it is terminal; treated as
    progress the pull looks alive and never finishes.
  • There is a test asserting a whole-body JSON.parse fails on the real stream, so nobody
    simplifies the line reader back into one parse.

Verification

  • core 1034 tests, 0 fail (36 new)
  • tsc clean for core and redrob
  • The fit judgement is pure and needs no machine; the GPU parsers run against the
    reference's own recorded output, so they are verified on a host with no GPU

No cloud fallback exists here and none should be added: a user who asked for a local model
asked for the data to stay on the machine.

… what is short

The two halves that decide whether this feature is honest. The surface that
shows a running download is deliberately not here.

Measuring first shrank the item: the RUN half already exists. A local runtime is
already an allowed provider -- the one trusted OpenAI-compatible package, local
URL only -- and the model listing already asks it what it serves. And the probe
is not native: the engine is a process on the user's own machine, so memory,
cores and arch come straight from the OS. Only the GPU needs a shell out.

`local-model-fit.ts` is built on three rules.

An unseen GPU is reported as NOT MEASURED, never as absent. Telling a user with a
discrete card that their machine has none is a false statement about their own
computer, and it is the one they would push back on hardest and be right about.
So `undefined` means we did not look, which is a third answer.

It never invents a model's size. The figure comes from whoever is offering the
model. A fit check against a size this module guessed would be a guess wearing a
verdict's clothes.

A refusal carries the numbers that produced it, and distinguishes the two cases
that matter: short by an amount this machine has if something is closed, versus
short by more than it has in total. "Your machine cannot run this" is not
actionable and cannot be checked by the person it is about.

Memory is checked against AVAILABLE rather than total, and `os.freemem()`
understates on Linux because it counts reclaimable page cache as used. That
direction is deliberate: it can refuse a model that would in fact have squeezed
in, and it will not promise one that gets killed mid-generation. The opposite
error loses the user's work.

The runtime's own listing and fetch shapes are in `local-model-pull.ts`, because
the OpenAI-compatible `GET /models` the provider listing uses reports an id and
nothing else -- no size -- so it cannot answer the question this feature turns
on. We do not fetch weights ourselves: no mirror choice, no quantisation choice,
no digest we chose to trust. The runtime already resumes interrupted downloads
and shares progress between concurrent callers; re-implementing that would mean
owning model provenance, which is a much larger promise than running a model
locally.

Shapes come from the runtime's published API reference rather than from reading
our own code back, and two traps it states explicitly have a test each:
`completed` may be absent while a layer is downloading, and `total` is per-LAYER
not per-model -- a caller reading it as whole-model progress shows a bar that
resets on every layer. An in-band `error` frame arrives after a 200 is already
sent and is terminal, not progress; treated as progress the pull looks like it is
still running and never finishes. There is also a test asserting that a
whole-body JSON parse FAILS on the real stream, so nobody simplifies the
line-by-line reader back into one parse.

core 1034 tests 0 fail, 36 new. The fit judgement is pure and needs no machine.
The GPU parsers run against the reference's own recorded output, so they are
verified on a host with no GPU.

No cloud fallback exists here and none should be added. A user who asked for a
local model asked for the data to stay on the machine; answering from a hosted
model would defeat the only reason to want this, and would do it invisibly.
…a 404 it found

The fit judgement and the probes had no caller. That is the failure this queue
keeps finding -- an arming matcher nobody invoked, a model preference read and
discarded, builtin skills registered into a list with no consumer -- so the
surface is the point of this change, not an extra.

`redrob models-local` prints the machine, then per configured local endpoint
every installed model with either "runs here" or the refusal and its numbers.

DRIVING IT END TO END FOUND A REAL DEFECT, and nothing short of that run would
have. A local runtime is configured with its OpenAI base, conventionally `/v1`,
but `/api/tags` is a SIBLING of `/v1`, not a child. Joining onto the configured
base produced `/v1/api/tags`, got a 404, and the command printed "answering,
with nothing installed" against a server serving two models. A 404 and an empty
list are both "no models" to a caller that does not separate them, so the bug
read as a working feature. Fixed, with a regression test that says why, and
reverse-verified: removing the fix fails it.

Only a TRAILING version segment is stripped. A runtime served under a prefix by
a reverse proxy keeps its path, which is why `/v1/engine` is left alone.

Measured on this host (16 cores, 61.8 GiB): a 2 GiB model runs here; a 120 GiB
model is refused with "needs about 144.0 GiB of memory (120.0 GiB of weights plus
runtime headroom) and 38.9 GiB is free of 61.8 GiB total -- short by 105.1 GiB,
more than this machine has in total".

Top-level `models-local` rather than `models local`, because `models` takes a
positional provider filter and a subcommand there is ambiguous with a provider
named `local`.

The GPU probe's failure is swallowed and reported as NOT MEASURED, never as "no
GPU". On a machine with no discrete card the tool simply does not exist, which is
the normal state of the world rather than a condition to report -- and the memory
judgement, the part that actually decides whether a model runs, does not depend
on it.

HttpClient and AppProcess are provided by this command rather than added to the
CLI runtime. It is the only command that speaks HTTP to a local runtime or shells
out for a GPU, and widening the shared runtime would hand both to every other
command that has no use for them.

core 1031 tests 0 fail across 123 files. redrob cli 380 tests 0 fail. Generated
SDK clean, tsc clean.
@its-janghoon
its-janghoon merged commit f8c7752 into develop Oct 3, 2026
15 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant