Skip to content

Port the unit and end-to-end tests to pytest - #482

Draft
negz wants to merge 7 commits into
modelplaneai:mainfrom
negz:fixture-upper
Draft

negz wants to merge 7 commits into
modelplaneai:mainfrom
negz:fixture-upper

Conversation

@negz

@negz negz commented Oct 1, 2026 •

Copy link
Copy Markdown
Collaborator

Description of your changes

Fixes #473.

#473 asks how we should write e2e tests, and the bake-off pointed at pytest. This PR moves the e2e tests and the function unit tests to pytest, so the repo has one test runner, and rewrites every unit test to one consistent pattern, which CONTRIBUTING's Tests section now describes.

To judge the pattern, read compose-model-cache's tests. No unit test's input or expectation changed: every distinct request and response is the same as before.

nix run .#test runs every function's unit tests. Name a function to run only its tests, and pass pytest arguments after it. Each case is a test of its own:

$ nix run .#test -- compose-model-cache -v
...::test_compose[GKE cluster first pass composes RWX PVC and hydration Job] PASSED
...::test_compose[PVC bound and Job complete reports Ready] PASSED
...
14 passed in 0.12s

When a case fails, pytest prints the whole response, with the expected lines as - and the actual lines as +, the same way round as Go's cmp.Diff(want, got). Here's one case with an expected value broken on purpose:

$ nix run .#test -- compose-model-cache -k "Job and complete"
...
E         -                             'phase': 'Hydrating',
E         ?                                       ^ -------
E         +                             'phase': 'Ready',
E         ?                                       ^^^^
...
FAILED functions/compose-model-cache/tests/test_fn.py::test_compose[PVC bound and Job complete reports Ready]

nix run .#e2e -- --verify brings up both kind clusters as before, then runs the e2e tests, passing any further arguments to pytest:

$ nix run .#e2e -- --verify
...
e2e/test_serving.py::test_chat_completion_succeeds PASSED
e2e/test_serving.py::test_chat_completion_without_a_valid_key_is_refused[no key] PASSED
e2e/test_serving.py::test_chat_completion_without_a_valid_key_is_refused[a key no Secret holds] PASSED
e2e/test_serving.py::test_response_reports_the_served_model PASSED
e2e/test_serving.py::test_unclaimed_model_is_not_routed PASSED
e2e/test_serving.py::test_models_lists_the_model_service PASSED
e2e/test_serving.py::test_anthropic_message_succeeds PASSED
e2e/test_serving.py::test_gateway_logs_a_usage_record PASSED
e2e/test_serving.py::test_cluster_gateway_refuses_a_caller_without_a_client_certificate PASSED
e2e/test_serving.py::test_cluster_gateway_has_no_plaintext_listener PASSED
10 passed in 532.32s (0:08:52)

--test reruns the e2e tests against clusters that are already up, and --clean tears them down:

$ nix run .#e2e -- --test -k usage
1 passed, 9 deselected in 25.39s
$ nix run .#e2e -- --clean

I have:

  • Read and followed Modelplane's contribution process.
  • Run nix flake check (or ./nix.sh flake check) and made sure it passes.
  • Added or updated tests covering any composition function changes. This PR changes only tests.
  • Signed off every commit with git commit -s.

The e2e tests are moving to pytest, and one test runner for the whole repo is
simpler than two. pytest 9 runs the existing unittest suites unchanged and
reports each subTest on its own, so this commit switches the runner without
touching a test. It configures pytest in strict mode, and to print whole
assertion diffs, since the tests compare whole responses.

Towards modelplaneai#473.

Signed-off-by: Nic Cope <nicc@rk0n.org>
The previous commit runs the unittest suites under pytest unchanged. This
commit ports them to pytest's own style, so each case in a table is a test of
its own that -k can select. Async tests call RunFunction with asyncio.run
rather than needing a plugin. pytest captures a test's output and shows it only
when the test fails, so the tests no longer disable the functions' logging.

pytest diffs two dicts in insertion order, where unittest sorted their keys.
The Structs a function returns rarely list their keys in the same order as the
test's expected ones, so the tests compare dicts with sorted keys, which keeps
a diff to the fields that differ. No case's request or expected response
changes.

Towards modelplaneai#473.

Signed-off-by: Nic Cope <nicc@rk0n.org>
The tests built their cases in several ways. Some derived a case from another
by copying and mutating it, or patched a request after building it. Others
built whole requests and responses with helpers, asserted on individual fields,
or computed expected values with the code under test or the SDK. A reader often
couldn't tell what a case checked without tracing code. Some case names also
claimed conditions their input didn't set up.

This commit rewrites the tests to the rules CONTRIBUTING now states, and says
why in a comment wherever a test departs from them. Helpers take everything
that varies as a required keyword argument, because a defaulted argument had
hidden that a case named for an unpinned replica was pinned. A field-level test
became a case in its entry point's table, or was deleted where a case already
sent the same request, and misleading case names are corrected.

Every distinct RunFunction request and response is unchanged. Writing cases
out in full takes the tests from about 19,000 lines to 39,000.

Towards modelplaneai#473.

Signed-off-by: Nic Cope <nicc@rk0n.org>
The checks build their virtualenvs from a uv2nix package set that checks.nix
defines for itself. The e2e app needs a virtualenv from the same set, so this
commit moves the set to nix/python.nix, and flake.nix passes it to the checks.

Signed-off-by: Nic Cope <nicc@rk0n.org>
The local e2e was a shell script, e2e/run.sh. Most of it checked the running
environment, with polling loops written out by hand, JSON matched as text, and
an exit at the first failed check, so one fault hid every check after it.

This commit replaces run.sh with a Python package. environment.py brings the two
clusters up and tears them down as run.sh did, and a pytest suite makes the
checks run.sh's --verify made, one test each. Each test waits only for what it
needs, so a gateway that refuses every caller fails only the tests that need
it to serve. nix run .#e2e keeps its flags.

Four checks are stricter. The response must name the served model in its model
field, where the name anywhere in the body passed. An unclaimed model must get a
404, where any status but 200 passed. /v1/models must list the model by its
exact name, where a substring of another name passed. And the usage record must
be a new one with the full endpoint name, where any earlier matching record
passed.

Towards modelplaneai#473.

Signed-off-by: Nic Cope <nicc@rk0n.org>
@negz negz changed the title Port the function unit tests to pytest and write their cases out in full Port the unit and end-to-end tests to pytest Oct 1, 2026
@negz negz added the test-e2e label Oct 1, 2026
Running one function's unit tests, or the e2e tests against clusters that were
already up, meant entering the dev shell and driving uv, a second tool to learn
alongside nix. This commit adds nix run .#test, which runs every function's unit
tests, or one function's with pytest arguments after its name, against the
virtualenvs the flake checks use. nix run .#e2e -- --test runs the e2e tests
against clusters that are already up.

Towards modelplaneai#473.

Signed-off-by: Nic Cope <nicc@rk0n.org>
The e2e tests shelled out to kubectl for every read, wait and exec. Its errors
reached the tests as stderr strings, a failed kubectl exec and a failed curl
both showed up as an exit code, and built-in objects came back as untyped
dicts. The official Kubernetes Python client returns typed objects and typed
errors. Requests to the gateways still run curl in a pod, because only a pod
can reach their published addresses from macOS.

Towards modelplaneai#473.

Signed-off-by: Nic Cope <nicc@rk0n.org>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Decide how to write e2e tests

1 participant