Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
46 commits
Select commit Hold shift + click to select a range
852e980
fix: open a bridged history on the user after every trim
REPPL Oct 4, 2026
1980bf9
chore: resolve iss-2610032306154631
REPPL Oct 4, 2026
a66e8c8
fix(selftest): never unload a model pinned during the run
REPPL Oct 4, 2026
4b25f77
chore: resolve the self-test pin issue
REPPL Oct 4, 2026
a6d586a
fix: refuse a completions prompt that is not a JSON string
REPPL Oct 4, 2026
1e3e179
chore: resolve the non-string completions prompt record
REPPL Oct 4, 2026
2ae3853
fix: drop a context probe's bounds when the provenance moves
REPPL Oct 4, 2026
ba2181d
chore: resolve iss-2610032241096901
REPPL Oct 4, 2026
0b1d1cc
Merge branch 'fix/bridge-history-opens-on-user' into integrate/0.10.0
REPPL Oct 4, 2026
7359d68
Merge branch 'fix/selftest-honours-pins' into integrate/0.10.0
REPPL Oct 4, 2026
60f56cd
Merge branch 'fix/refuse-non-string-prompt' into integrate/0.10.0
REPPL Oct 4, 2026
6259343
Merge branch 'fix/probe-resume-clamps-window' into integrate/0.10.0
REPPL Oct 4, 2026
4a92d50
chore: capture two bridge history findings from review
REPPL Oct 4, 2026
bb43455
fix: send a bridged history that alternates after an unanswered message
REPPL Oct 4, 2026
d3681e1
chore: capture the time-of-check gap in the idle jobs' unload
REPPL Oct 4, 2026
856cb5f
fix(selftest): never load a model that is pinned and not loaded
REPPL Oct 4, 2026
3fe9263
fix: fit the bridge's single-turn fallback to the window in encoded b…
REPPL Oct 4, 2026
6a0982c
docs: say the self-test measures a pinned model only while it is loaded
REPPL Oct 4, 2026
d2cccd3
test: hold the bridged history tests to how much history is kept
REPPL Oct 4, 2026
a003d6a
chore: resolve iss-2610042030092101 and iss-2610042030098334
REPPL Oct 4, 2026
7164c10
docs: record why a string prompt is not held to the stop-string rule
REPPL Oct 4, 2026
0c770a1
chore: capture two pre-existing gateway findings from the security re…
REPPL Oct 4, 2026
d76105c
fix(selftest): tell a pinned refusal from a failed unload
REPPL Oct 4, 2026
46d754a
Merge branch 'fix/bridge-history-alternates' into integrate/0.10.0
REPPL Oct 4, 2026
d96945f
Merge branch 'fix/selftest-skips-unloaded-pins' into integrate/0.10.0
REPPL Oct 4, 2026
bb42ceb
chore: capture a budget change judged against the old budget
REPPL Oct 4, 2026
c576c65
test: show a resumed probe starts over rather than resuming old bounds
REPPL Oct 4, 2026
5d975f2
test: hold the window-trim test to keeping more than the newest turn
REPPL Oct 4, 2026
9d91069
fix: judge a probe's figure against the settings in force as it is saved
REPPL Oct 4, 2026
a2bf405
fix: take a probe's ceiling from the provenance it is stamped with
REPPL Oct 4, 2026
e3b7024
docs: say an unfinished measurement starts over when its settings move
REPPL Oct 4, 2026
e96a915
fix: re-judge measurements after a save puts the new budget in force
REPPL Oct 4, 2026
a41a0a9
chore: resolve iss-2610042040451551
REPPL Oct 4, 2026
f4e93c6
Merge branch 'fix/probe-provenance-followup' into integrate/0.10.0
REPPL Oct 4, 2026
4595e1a
chore: defer the two open majors out loud for the 0.10.0 cut
REPPL Oct 4, 2026
d590e89
chore: record the findings and the idea filed on 2026-10-04
REPPL Oct 4, 2026
dd929d8
chore: close the update-check and unload specs, shipping their intents
REPPL Oct 4, 2026
b0bf996
chore: record the unload intent's fidelity review and its two concerns
REPPL Oct 4, 2026
1be355f
chore: record the update-check intent's fidelity review and its findings
REPPL Oct 4, 2026
1054453
fix: judge a probe's figure even when the registry's write fails
REPPL Oct 4, 2026
8352c94
test: show a budget-only save lifts a recorded load failure at once
REPPL Oct 4, 2026
782b5f9
docs: correct the budget re-judging's reasoning and narrow its claim
REPPL Oct 4, 2026
24240c1
refactor: drop the context probe's unused Candidate.Served
REPPL Oct 4, 2026
4b29fcf
test: say the save in the probe-save test raised the served window
REPPL Oct 4, 2026
05e5a5f
Merge branch 'fix/probe-provenance-final' into integrate/0.10.0
REPPL Oct 4, 2026
5b01245
docs: cut 0.10.0
REPPL Oct 4, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
Original file line number Diff line number Diff line change
@@ -1,10 +1,10 @@
---
id: adr-2610040749545010
slug: the-gateway-may-refuse-a-request-whose-prompt-or-messages
status: accepted
status: accepted; superseded in part by adr-2610042021365934 (a /v1/completions prompt that is not a JSON string is refused too, reading its kind)
date: 2026-10-04
supersedes: [adr-2609201008470380, adr-2609061610102325]
superseded_by: null
superseded_by: adr-2610042021365934
related_intents: []
related_rfcs: []
related_adrs: [adr-2609201008470380, adr-2609061610102325]
Expand Down
Original file line number Diff line number Diff line change
@@ -0,0 +1,116 @@
---
id: adr-2610042021365934
slug: the-gateway-may-refuse-a-completions-prompt-that-is-not-a
status: accepted
date: 2026-10-04
supersedes: [adr-2610040749545010]
superseded_by: null
related_intents: []
related_rfcs: []
related_adrs: [adr-2610040749545010, adr-2609201008470380, adr-2609061610102325]
---

# ADR-2610042021365934: The gateway may refuse a completions prompt that is not a JSON string, reading its kind and nothing else

**Supersedes in part:**
[adr-2610040749545010](2610040749545010-the-gateway-may-refuse-a-request-whose-prompt-or-messages.md),
on two of its sentences, and for a `/v1/completions` prompt only: that the
check "reads whether the field is empty and nothing else", and that "a prompt
or messages value of another type ... [is] left to the model server". The
check now also reads the prompt's JSON kind. Every other decision of that
record stands: what counts as an empty string prompt, the refusal of empty
`messages` on the network and bridge paths, a missing field left to the model
server, keeping nothing, and the one file the check lives in.

## Context

adr-2610040749545010 granted the gateway one bit about a request's prompt:
whether it is empty. Its adversarial review found that a `/v1/completions`
prompt that is an array of strings reaches the generation thread of mlx-lm
0.32.0 outside its `try` (iss-2610040805353721). The tokenizer treats a list
of strings as a batch, so `["hi"]` tokenizes to `[[72, 73]]` and `[""]` to
`[[]]`, and the nested list raises in the prompt cache's search and in the
batch generator. The emptiness check refused only an empty array and let
`[""]` and `["hi"]` through, because telling them apart from a string reads
more than emptiness.

The maintainer set the condition on 2026-10-03: reproduce it on the Mac first,
and refuse it if it freezes or crashes the server. On 2026-10-04, on Dessau
v0.9.3 with mlx-lm 0.32.0 and a 30B mixture-of-experts 4-bit model, each prompt
on a fresh load:

- `["hi"]` and `[""]` each got no answer, and the ordinary prompt sent after
each got none either: the model server was frozen, as the empty prompt had
left it.
- `[72,73]`, an array of token ids, was refused by mlx-lm with a 404 ("text
input must be of type `str`"), and the ordinary prompt after it was
answered.

So on the pinned server no prompt that is not a string yields a generation: a
token array is refused upstream, and an array of strings freezes the model for
every client. The health watch on main turns such a freeze into a crash and a
restart (iss-2610040752568866), which still fails every request in flight on
that model.

## Decision

We will refuse, with a 400 and before a model is resolved or loaded, a
`/v1/completions` request whose `prompt` is present and is not a JSON string:
an array of any kind (of strings, of token ids, or empty), a number, an
object, a boolean or `null`. The refusal names the field and its required
kind, `"prompt" must be a string`, and never the value.

The check reads the prompt's JSON kind, from the first byte of its raw value,
and nothing else: it never looks inside an array or an object, never counts
elements, and decodes a value only when it is already a string, to apply the
emptiness test adr-2610040749545010 grants. It keeps nothing. A missing
`prompt` is still left to the model server, whose check for it runs on its
handler thread. The chat path and the bridge's `Ask` path are unchanged: they
read `messages`, whose emptiness is all they read.

The check stays in `internal/gateway/emptyprompt.go`, in the same function as
the emptiness check, and its entry on the prompt-content readers' list in
`internal/archtest/prompt_content_test.go` names the wider reading.

We will treat any further reading of prompt content by the gateway as a new
decision that supersedes this record, as the records before it said.

## Alternatives Considered

1. **Refuse any prompt that is not a JSON string (chosen).** It closes a
reproduced freeze that any client on the network can cause, and refuses
nothing that works on the pinned server. The grant grows from one bit to
the value's JSON kind, read from its first byte.
2. **Refuse only an array whose elements include a string.** It is the
narrowest refusal of the reproduced freeze, but it reads inside the array,
element by element, which is more of the prompt than its kind. A token
array and a number would still reach the model server, which refuses them
with a 404 that says nothing a client can act on. Rejected as a wider read
for a smaller result.
3. **Leave it to the health watch.** The watch restarts a frozen model server
within a tick, but every request in flight on that model fails, and any
client can repeat the request at will. Rejected as the only defence, as
adr-2610040749545010 rejected it for the empty prompt.
4. **Accept arrays and fan them out as one request per element, as OpenAI
does.** That reads every element, writes new requests from client text and
adds batching the gateway does not otherwise have. Rejected as a feature
nobody has asked for, and a far wider grant than the defect calls for.

## Consequences

- adr-2610040749545010 changes only its status fields, to say it is
superseded in part by this record; its text is not edited.
- A client that sends an array, number, object, boolean or `null` prompt gets
a 400 naming the field. Before, an array of strings froze the model for
everyone and the rest were refused by the model server with a 404. That is
a breaking change on the wire, recorded as `impact: breaking`, though no
such request produced an answer before.
- A client written for OpenAI's batched or tokenized `prompt` forms cannot use
them through Dessau. It could not before either.
- If a later mlx-lm pin accepts a list prompt and generates from it, this
refusal would withhold a working feature. Moving the pin is the moment to
check, and lifting the refusal is a new decision.
- Obligations: the adversarial security review of the change that lands this,
the gateway being a trust boundary; a test that each non-string kind is
refused before the pool or the model server is reached, and that a string
prompt is not; and the readers' list entry updated to name the kind read.
3 changes: 2 additions & 1 deletion .abcd/development/decisions/adrs/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -43,4 +43,5 @@ hand edit.
| [adr-2609201008477513](2609201008477513-a-deliberately-invoked-per-model-diagnostic-may-write-prompt.md) | A deliberately invoked per-model diagnostic may write prompts and answers to the operator's own log, and the client is not told | accepted | 2026-09-20 |
| [adr-2609202200197975](2609202200197975-who-counts-as-one-client-for-scheduling-a-paired-client-is-o.md) | Who counts as one client for scheduling: a paired client is one person by its pairing, unpaired clients are one class served last, and what is written down is the contention, never the client | accepted | 2026-09-20 |
| [adr-2610030906462776](2610030906462776-dessau-serves-from-one-macos-account-the-machine-wide-shared.md) | Dessau serves from one macOS account; the machine-wide shared model cache is retired | accepted | 2026-10-03 |
| [adr-2610040749545010](2610040749545010-the-gateway-may-refuse-a-request-whose-prompt-or-messages.md) | The gateway may refuse a request whose prompt or messages are empty, reading emptiness and nothing else | accepted | 2026-10-04 |
| [adr-2610040749545010](2610040749545010-the-gateway-may-refuse-a-request-whose-prompt-or-messages.md) | The gateway may refuse a request whose prompt or messages are empty, reading emptiness and nothing else | superseded in part by [adr-2610042021365934](2610042021365934-the-gateway-may-refuse-a-completions-prompt-that-is-not-a.md) on the prompt's kind | 2026-10-04 |
| [adr-2610042021365934](2610042021365934-the-gateway-may-refuse-a-completions-prompt-that-is-not-a.md) | The gateway may refuse a completions prompt that is not a JSON string, reading its kind and nothing else | accepted | 2026-10-04 |
Original file line number Diff line number Diff line change
@@ -0,0 +1,43 @@
---
id: itd-2610041336141168
slug: a-dessau-agent-that-selects-the-most-appropriate-model-for-a
spec_id: null
kind: null
suggested_kind: null
reclassification_history: []
builds_on: []
severity: minor
impact: additive
origin: researcher-authored
production_mode: hand-written
---

# Dessau picks the model that answers each request

## Press Release

> A 'Dessau' agent that selects the most appropriate model for a given response. It parses the question, decides, answers, parses the answer and only then returns it, depending on the client (e.g. Discord: in chunks; Dessau Chat: whole with thinking; etc.). This agent can later be expanded, e.g. by using the Brave search API.

## Why This Matters

> _Why this matters to the user — replace before planning._

## Mechanism

> _Prompted (the claim-recording gradient): why the authors expect this to work, as a falsifiable "we expect X because Y" — not the outcome restated. Replace this line with the claim, or with the exact token `None stated.` alone on its line to record the claim as considered and declined._

## Scope Conditions

> _Required (the claim-recording gradient): the population, platform, scale, or assumptions this claim holds under, one per top-level bullet — `abcd intent plan` stamps each with a persistent identity. Replace this line with those bullets, or with the exact token `None stated.` alone on its line._

## Acceptance Criteria

> _Required (the itd-1 discipline): add at least one Given-When-Then bullet describing the verifiable bar for "shipped" before this draft can be planned._

## Open Questions

_None recorded yet._

## Audit Notes

_Empty. Populated by intent-auditor when intent moves to shipped/._

This file was deleted.

Loading
Loading