Skip to content
Draft
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 2 additions & 2 deletions .github/workflows/e2e.yml
Original file line number Diff line number Diff line change
Expand Up @@ -66,8 +66,8 @@ jobs:

# The exact command a developer runs locally (--print-build-logs is only CI
# verbosity), so a green run here and a green run on a laptop mean the same
# thing. It brings up both clusters, deploys the mock model, waits for the
# ModelService, and asserts a live 200 — exiting non-zero on any failure.
# thing. It brings up both clusters, deploys the mock model, and runs the
# tests in e2e/, exiting non-zero if any fails.
- name: Run e2e
run: nix run .#e2e --print-build-logs -- --verify

Expand Down
141 changes: 91 additions & 50 deletions CONTRIBUTING.md
Original file line number Diff line number Diff line change
Expand Up @@ -105,7 +105,8 @@ curl -fsSL https://install.determinate.systems/nix | sh -s -- install

`nix flake check` runs all of the project's checks inside the Nix sandbox:
Python, shell, and Nix linters and formatters, the [ty](https://docs.astral.sh/ty)
type checker on every composition function, plus unit tests for every function.
type checker on every composition function and the end-to-end tests, plus unit
tests for every function.
Run `nix flake show` to see what else is available.

```bash
Expand All @@ -120,7 +121,7 @@ composition function renders the right resources. The integration layer is
`nix run .#e2e`, which brings up two local `kind` clusters and runs the
whole path — scheduling, the serving-stack install on a registered cluster,
gateway routing, a live request — with no cloud credentials. Add `-- --verify`
and it waits for readiness, asserts a 200, and exits non-zero on failure. That
and it runs the pytest suite in `e2e/`, exiting non-zero if a test fails. That
verify command is what the label-gated `E2E` workflow runs on CI (add the
`test-e2e` label to a PR), so a green local `--verify` and a green CI run mean
the same thing. See `e2e/README.md`.
Expand Down Expand Up @@ -245,7 +246,7 @@ functions/<name>/
main.py # CLI entrypoint (boilerplate)
fn.py # FunctionRunner gRPC service and Composer logic
tests/
test_fn.py # unittest-based tests for fn.py
test_fn.py # pytest tests for fn.py
```

The `Composer.compose()` method in `fn.py` reads the XR from the request,
Expand Down Expand Up @@ -274,12 +275,15 @@ XRDs or dependencies you've removed don't linger.

### Tests

Every function has tests under `functions/<name>/tests/test_fn.py`. The
canonical form is a table of `Case`s, each running the function on a
Every function has tests under `functions/<name>/tests/`, run with
[pytest](https://docs.pytest.org/). `test_fn.py` tests the function as a whole.
A module with logic of its own, such as `compose-model-deployment`'s scheduler,
can have its own `test_<module>.py` too.

The canonical form is a table of `Case`s, each running the function on a
`RunFunctionRequest` and comparing the whole `RunFunctionResponse` against an
expected one — not asserting on individual fields. `compose-usages` is a clean
example; `compose-model-cache` shows the same form scaled up to a multi-pass
reconcile. The skeleton:
expected one, rather than asserting on individual fields. `compose-model-cache`
is a good example. The skeleton:

```python
@dataclasses.dataclass
Expand All @@ -289,50 +293,87 @@ class Case:
want: fnv1.RunFunctionResponse


def setUpModule() -> None:
logging.configure(level=logging.Level.DISABLED)


class TestFunctionRunner(unittest.IsolatedAsyncioTestCase):
maxDiff = None

@classmethod
def setUpClass(cls) -> None:
cls.runner = fn.FunctionRunner()

async def test_compose(self) -> None:
cases = [
Case(
name="describes what this case exercises",
req=fnv1.RunFunctionRequest(...),
want=fnv1.RunFunctionResponse(...),
),
]
for case in cases:
with self.subTest(case.name):
got = await self.runner.RunFunction(case.req, None)
self.assertEqual(
json_format.MessageToDict(case.want),
json_format.MessageToDict(got),
"-want, +got",
)
COMPOSE_CASES = [
Case(
name="describes what this case exercises",
req=fnv1.RunFunctionRequest(...),
want=fnv1.RunFunctionResponse(...),
),
]


def _to_dict(msg: message.Message) -> dict:
"""msg as a dict with sorted keys, so pytest's diff of two lines them up."""
return json.loads(json_format.MessageToJson(msg, sort_keys=True))


@pytest.mark.parametrize("case", COMPOSE_CASES, ids=lambda case: case.name)
def test_compose(case: Case) -> None:
"""RunFunction composes the resources an XR needs."""
got = asyncio.run(fn.FunctionRunner().RunFunction(case.req, None))
assert _to_dict(got) == _to_dict(case.want)
```

Name a table for the test that runs it, and put it just above that test. Each
case becomes its own test, named for the case, so `pytest -k` can select it.
With `got` on the left, pytest's diff shows the expected lines as `-` and the
actual lines as `+`, the same way round as Go's `cmp.Diff(want, got)`. Tests are
plain functions, with no classes, fixtures, or `conftest.py`. They call the
async `RunFunction` with `asyncio.run` rather than needing a plugin, and check
errors with `pytest.raises(..., match=...)`.

Cases are data, so a reader should be able to see everything a case asserts by
reading it:

- **Write each case out in full.** Repetition between cases is fine. Don't
derive one case from another, or from a shared base, by copying and mutating
it, and don't change a request or response once it's built. Pass
requirements, conditions, and results to the constructor.
- **A resource that appears in three or more cases gets a helper,** the XR
included. Count resources by the role they play, such as "the GPU node pool"
or "an endpoint's Backend". An observed resource plays a different role from
the desired resource it reflects, so it gets its own helper. A helper builds
that one resource and returns the `fnv1.Resource` that carries it, or a dict
where another resource embeds it. Everything that varies between the cases
that use it is a keyword argument with no default, including readiness as an
`fnv1.Ready` value, so every call shows every value that varies. Write a
resource that appears in one or two cases inline. Never write a helper that
builds a whole request, response, map of resources, or case.
- **Name a case for what its input sets up** and what it expects. Put a comment
on the case as a whole directly above its `Case(`. A short comment beside a
single value can explain that value.
- **Compare the whole output, once.** A test that calls the same entry point
with different data belongs in that entry point's table as another case.
- **Write values as literals,** in requests and expectations alike, including
names the function hashes. An expectation computed by code, whether the code
under test or the SDK's `child_name`, passes whatever that code does.
- **Build the XR from its generated model,** with
`resource.dict_to_struct(xr.model_dump(exclude_none=True, mode="json",
by_alias=True))`. Write composed and observed resources as dicts in their wire
form. The generated models include schema defaults, so a model doesn't fix its
own wire form: the SDK sends only the fields a function sets, while the API
server fills in the defaults.
- **Say why where a test departs from a rule,** in a comment beside the
departure.

Because `want` is the whole response, it must include the parts the function
always emits: `meta.ttl` (60s), an empty `context`, and any conditions, results,
and requirements. Give observed conditions a fixed `lastTransitionTime` so the
input is deterministic. Protobuf maps (`desired.resources`,
`requirements.resources`) compare order-independently, but repeated fields
(`conditions`, `results`, status arrays) must match the order the function
emits.

`nix flake check` runs every function's tests, and so does `nix run .#test`,
outside the sandbox. Name a function to run only its tests, and pass pytest
arguments after it:

```bash
nix run .#test -- compose-usages -k namespace
```

Build the XR with
`resource.dict_to_struct(xr.model_dump(exclude_none=True, mode="json"))` from a
generated Pydantic model; build other observed, desired, and required resources
as plain dicts. Because `want` is the whole response, it must include the parts
the function always emits: `meta.ttl` (60s), an empty `context`, and any
conditions, results, and requirements. Give observed conditions a fixed
`lastTransitionTime` so the input is deterministic. Protobuf maps
(`desired.resources`, `requirements.resources`) compare order-independently, but
repeated fields (`conditions`, `results`, status arrays) must match the order
the function emits.

Some existing tests (`compose-serving-stack`, the second method in
`compose-eks-cluster`) predate this form and assert on individual fields. Don't
model new tests on them. Add new cases to the function's `test_fn.py` and run
`nix flake check` to verify they pass.
Each function runs in a pytest session of its own, because every function names
its package `function`.

### Running locally

Expand Down
60 changes: 40 additions & 20 deletions e2e/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -80,10 +80,10 @@ runs, so prefer a separate control plane for cloud work.
(kube-prometheus-stack et al.) add up fast; a full Docker disk surfaces as
`no space left on device`. Reclaim between runs with `docker builder prune -af`
and `docker image prune -af`.
- Everything else — `kind`, `kubectl`, `curl`, `git`, and the `crossplane` CLI —
is provided by the flake via `nix run .#e2e`.
- Everything else — `kind`, `kubectl`, the `docker` CLI, Python and pytest, and
the `crossplane` CLI — is provided by the flake via `nix run .#e2e`.

The workload cluster is pinned to **k8s v1.34** (in `run.sh`) for the
The workload cluster is pinned to **k8s v1.34** (in `environment.py`) for the
`resource.k8s.io` (DRA) APIs, on-by-default in 1.34: both the serving stack's
NVIDIA DRA driver and the dra-example-driver register DeviceClasses, and the
example driver publishes the `ResourceSlice`s the engine's `ResourceClaim` binds
Expand All @@ -93,10 +93,20 @@ against. The control-plane cluster needs no DRA.

```bash
nix run .#e2e # bring up both clusters + deploy the mock model
nix run .#e2e -- --verify # same, then wait for readiness and assert a live 200
nix run .#e2e -- --verify # same, then run the tests
nix run .#e2e -- --test # run the tests against clusters already up
nix run .#e2e -- --clean # tear both clusters down
```

Arguments after `--verify` go to pytest, so `nix run .#e2e -- --verify -k usage`
runs only the tests whose names match. Bring-up reuses clusters that are already
up. To rerun the tests against an environment that's up, without bringing it up
again, use `--test`:

```bash
nix run .#e2e -- --test -k usage
```

`crossplane project run` installs the config and applies the resources, then
returns; the serving-stack install and model rollout reconcile in the background.
So wait for the `ModelService` to become ready before curling. The gateway's
Expand All @@ -122,16 +132,17 @@ kubectl run curl -n ml-team --rm -it --image=curlimages/curl@sha256:7c12af72ceb3
-d '{"model":"ml-team/mock","max_tokens":16,"messages":[{"role":"user","content":"hi"}]}'
```

`--verify` runs both of those, among other checks, and exits non-zero on
failure. It's the exact command the `E2E` CI workflow runs, so a green `--verify`
locally and a green CI run mean the same thing; use the manual curls above to
poke the endpoints interactively.
`--verify` runs the tests in `test_serving.py`, which send both of those
requests among others, and exits non-zero if any fails. It's the exact command
the `E2E` CI workflow runs, so a green `--verify` locally and a green CI run
mean the same thing; use the manual curls above to poke the endpoints
interactively.

## How it's structured

`nix run .#e2e` materialises the Nix-built function images and hands off to
`run.sh`, which does the cross-cluster orchestration that `crossplane project
run` flags can't express:
`environment.py`, which does the cross-cluster orchestration that `crossplane
project run` flags can't express:

1. Create the **workload** kind cluster (pinned v1.34).
2. Install MetalLB on it (the serving stack doesn't) with a pool inside the
Expand All @@ -147,13 +158,21 @@ run` flags can't express:
control-plane pods over the shared kind network), then apply the
Modelplane manifests.

Everything the control plane needs is a declarative manifest; the shell in
`run.sh` is only the irreducible cross-cluster setup (a second cluster, its
MetalLB and DRA driver, and the cross-cluster kubeconfig).
Everything the control plane needs is a declarative manifest; `environment.py`
is only the irreducible cross-cluster setup (a second cluster, its MetalLB and
DRA driver, and the cross-cluster kubeconfig). With `--verify`, pytest then runs
`test_serving.py`. Its fixtures in `conftest.py` wait for the model to serve,
then start a curl pod on each cluster to send requests from. The tests read the
clusters through the Kubernetes API, with the official Python client, while
bring-up drives the kind, crossplane, docker and kubectl CLIs.

```
e2e/
run.sh # two-cluster orchestration
environment.py # two-cluster bring-up and teardown
conftest.py # fixtures: the clusters, curl pods, readiness
test_serving.py # the tests
kube.py, gateway.py # Kubernetes API and curl helpers
wait.py # polling until a condition holds
dra-example-driver.yaml # vendored fake DRA GPU driver (applied to workload)
manifests/ # applied to the control plane after setup
00-namespaces.yaml
Expand All @@ -169,24 +188,24 @@ e2e/
- **MetalLB on the workload cluster.** Both gateways run there, and both need
`LoadBalancer` addresses kind can't provide: the serving stack gates the
cluster gateway's readiness on having one (`READY_CEL` in its `gateway.py`).
Nothing Modelplane composes installs MetalLB, so `run.sh` does, with a pool
Nothing Modelplane composes installs MetalLB, so bring-up does, with a pool
inside the detected kind Docker subnet (see caveat) so the control plane can
route to the addresses it hands out.
- **Fake DRA driver.** A `claim: DRA` engine emits a `ResourceClaim`; with no DRA
driver it stays Pending and the pod never schedules. `run.sh` applies the
driver it stays Pending and the pod never schedules. Bring-up applies the
vendored **dra-example-driver**, which publishes fake `gpu.example.com` devices
so the claim binds on a GPU-less node.
- **Cross-cluster kubeconfig.** `source: Existing` needs a kubeconfig the
control-plane provider pods can use to reach the workload API server.
`kind get kubeconfig --internal` gives an address routable across the shared
kind network; a host kubeconfig (`127.0.0.1:<port>`) wouldn't be.
- **Node label.** On a BYO cluster Modelplane doesn't provision/label pools, so
`run.sh` labels the workload node `modelplane.ai/pool=gpu-synthetic` (matching
bring-up labels the workload node `modelplane.ai/pool=gpu-synthetic` (matching
`nodePools[].name`); without it worker pods stay Pending.

## Caveats / open questions

- **Cross-cluster networking uses the detected kind subnet.** `run.sh` reads the
- **Cross-cluster networking uses the detected kind subnet.** Bring-up reads the
`kind` Docker network's subnet (usually 172.18.0.0/16, but kind bumps to
172.19/... when earlier networks already hold 172.18) and derives the
workload cluster's MetalLB pool from it. A hardcoded 172.18 would leave the LB
Expand All @@ -195,10 +214,11 @@ e2e/
for the config to install, then applies the resources and exits — it doesn't
block on XR readiness. The serving-stack install (the long pole) and the model
rollout happen after, so watch the `ModelService`'s `RoutingReady` rather than
the command's exit. `--timeout` in `run.sh` bounds the build and config install.
the command's exit. `--timeout` in `environment.py` bounds the build and config
install.
- **Two DRA drivers on a GPU-less node.** The serving stack's **NVIDIA** DRA
driver targets NFD-GPU-labelled nodes, so it sits at 0/0 (inert) yet its Helm
release still reports Ready. The **dra-example-driver** `run.sh` installs is the
release still reports Ready. The **dra-example-driver** bring-up installs is the
active one — it publishes the fake `gpu.example.com` devices the engine binds.
- **Serving-stack weight.** cert-manager, Envoy Gateway, Envoy AI Gateway, GAIE
CRDs, kube-prometheus-stack, LeaderWorkerSet, NFD, DRA driver — all on the
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -12,4 +12,4 @@
# See the License for the specific language governing permissions and
# limitations under the License.


"""End-to-end tests that bring Modelplane up on kind and send it traffic. See README.md."""
Loading
Loading