Conversation
The e2e tests are moving to pytest, and one test runner for the whole repo is simpler than two. pytest 9 runs the existing unittest suites unchanged and reports each subTest on its own, so this commit switches the runner without touching a test. It configures pytest in strict mode, and to print whole assertion diffs, since the tests compare whole responses. Towards modelplaneai#473. Signed-off-by: Nic Cope <nicc@rk0n.org>
The previous commit runs the unittest suites under pytest unchanged. This commit ports them to pytest's own style, so each case in a table is a test of its own that -k can select. Async tests call RunFunction with asyncio.run rather than needing a plugin. pytest captures a test's output and shows it only when the test fails, so the tests no longer disable the functions' logging. pytest diffs two dicts in insertion order, where unittest sorted their keys. The Structs a function returns rarely list their keys in the same order as the test's expected ones, so the tests compare dicts with sorted keys, which keeps a diff to the fields that differ. No case's request or expected response changes. Towards modelplaneai#473. Signed-off-by: Nic Cope <nicc@rk0n.org>
The tests built their cases in several ways. Some derived a case from another by copying and mutating it, or patched a request after building it. Others built whole requests and responses with helpers, asserted on individual fields, or computed expected values with the code under test or the SDK. A reader often couldn't tell what a case checked without tracing code. Some case names also claimed conditions their input didn't set up. This commit rewrites the tests to the rules CONTRIBUTING now states, and says why in a comment wherever a test departs from them. Helpers take everything that varies as a required keyword argument, because a defaulted argument had hidden that a case named for an unpinned replica was pinned. A field-level test became a case in its entry point's table, or was deleted where a case already sent the same request, and misleading case names are corrected. Every distinct RunFunction request and response is unchanged. Writing cases out in full takes the tests from about 19,000 lines to 39,000. Towards modelplaneai#473. Signed-off-by: Nic Cope <nicc@rk0n.org>
The checks build their virtualenvs from a uv2nix package set that checks.nix defines for itself. The e2e app needs a virtualenv from the same set, so this commit moves the set to nix/python.nix, and flake.nix passes it to the checks. Signed-off-by: Nic Cope <nicc@rk0n.org>
The local e2e was a shell script, e2e/run.sh. Most of it checked the running environment, with polling loops written out by hand, JSON matched as text, and an exit at the first failed check, so one fault hid every check after it. This commit replaces run.sh with a Python package. environment.py brings the two clusters up and tears them down as run.sh did, and a pytest suite makes the checks run.sh's --verify made, one test each. Each test waits only for what it needs, so a gateway that refuses every caller fails only the tests that need it to serve. nix run .#e2e keeps its flags. Four checks are stricter. The response must name the served model in its model field, where the name anywhere in the body passed. An unclaimed model must get a 404, where any status but 200 passed. /v1/models must list the model by its exact name, where a substring of another name passed. And the usage record must be a new one with the full endpoint name, where any earlier matching record passed. Towards modelplaneai#473. Signed-off-by: Nic Cope <nicc@rk0n.org>
Running one function's unit tests, or the e2e tests against clusters that were already up, meant entering the dev shell and driving uv, a second tool to learn alongside nix. This commit adds nix run .#test, which runs every function's unit tests, or one function's with pytest arguments after its name, against the virtualenvs the flake checks use. nix run .#e2e -- --test runs the e2e tests against clusters that are already up. Towards modelplaneai#473. Signed-off-by: Nic Cope <nicc@rk0n.org>
The e2e tests shelled out to kubectl for every read, wait and exec. Its errors reached the tests as stderr strings, a failed kubectl exec and a failed curl both showed up as an exit code, and built-in objects came back as untyped dicts. The official Kubernetes Python client returns typed objects and typed errors. Requests to the gateways still run curl in a pod, because only a pod can reach their published addresses from macOS. Towards modelplaneai#473. Signed-off-by: Nic Cope <nicc@rk0n.org>
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Description of your changes
Fixes #473.
#473 asks how we should write e2e tests, and the bake-off pointed at pytest. This PR moves the e2e tests and the function unit tests to pytest, so the repo has one test runner, and rewrites every unit test to one consistent pattern, which CONTRIBUTING's Tests section now describes.
To judge the pattern, read compose-model-cache's tests. No unit test's input or expectation changed: every distinct request and response is the same as before.
nix run .#testruns every function's unit tests. Name a function to run only its tests, and pass pytest arguments after it. Each case is a test of its own:When a case fails, pytest prints the whole response, with the expected lines as
-and the actual lines as+, the same way round as Go'scmp.Diff(want, got). Here's one case with an expected value broken on purpose:nix run .#e2e -- --verifybrings up both kind clusters as before, then runs the e2e tests, passing any further arguments to pytest:--testreruns the e2e tests against clusters that are already up, and--cleantears them down:I have:
nix flake check(or./nix.sh flake check) and made sure it passes.Added or updated tests covering any composition function changes.This PR changes only tests.git commit -s.