ci(tests): add a plain-unittest leg to catch conftest-masked ordering defects - #984
Conversation
…rdering defects (#981) RegistrationGateTests bound Engine, EngineBase, Dumper, DumperBase, Reassembly, ReassemblyBase, TraceFlow, TraceFlowBase and Extractor once, at import time. A sibling module that purges pcapkit from sys.modules and reimports it (deliberately asymmetric; tests.conftest's autouse fixture reconciles it, but only under pytest) leaves those names stale while the registration hooks they exercise re-import their own collaborators and see the fresh generation -- so a dynamically-created subclass registers in one generation's registry while the test reads another's. Four subTests failed this way and pytest never reported it: pytest-subtests marks the parent test "passed" the moment only its subTests failed. RegistrationGateTests.setUp now re-resolves every base/public pair and the Extractor singleton via importlib.import_module, so the suite always compares like generation to like. util/run_unittest_leg.py and a new unittest-ordering job run each tests/ subdirectory, paired with the root-level modules, under unittest.TextTestRunner -- the runner this class of defect is invisible to. Scoped per directory for cost (the full suite OOMs at 29 GB in one process); tests/corekit and tests/vendor are excluded, see the job's own comment for why. Build: tests/test_base_class_contract.py passes alone and alongside its #981 reproduction; util/run_unittest_leg.py protocols/const/etc. all pass.
|
NEEDS CHANGES at
So pytest's green on #981 comes entirely from The mechanism itself is proven, not papered over. Independent
The
Four prose defects: One structural point I am passing on rather than acting on. The reviewer argues the reconciliation UNVERIFIED by the review: an uncontended |
…d timeout Problem: #981's job and script blamed pytest-subtests for masking the ordering defect, four times over, but that plugin is not installed and current pytest already reports subTest failures on its own -- the real masking comes from tests/conftest.py's autouse restore_module_table fixture, which plain unittest never loads. The job's 20-minute timeout also had no margin: a review pass measured contended protocols runs up to 1275s, over the cap, and the job carried no `nproc` step the other jobs have. The tests/vendor exclusion cited a nonexistent tracking issue, undercounted its excluded directories at two instead of three, and :data:`LEGS` was referenced twice though it lives in the YAML matrix, not this module. Fix: rewrote the four pytest-subtests passages (two per file) to name the autouse fixture as the real mechanism, and recorded `pytest --noconftest` as a cheaper alternative that was not evaluated at the time rather than inventing a reason for the choice after the fact. Raised timeout-minutes to 45, added the "Report available parallelism" step, and made the leg runner print wall-clock elapsed time unconditionally. Cited #985 for the vendor exclusion, named all three excluded directories, and fixed the LEGS/stats errors. Evidence: pytest-subtests absent (`pip show` warns not found); a synthetic subTest probe under this venv's pytest 9.1.1 reports SUBFAILED entries, not a passing parent; `pytest --noconftest` over tests/protocols/application/test_http_unit.py then tests/test_base_class_contract.py on the pre-#981-fix tree reproduced the same four subTest failures in 68.4s. cli and dumpkit legs still pass under the edited script (185 and 199 tests, 18.3s/39.2s elapsed). YAML reparsed: 5 top-level keys, 8 jobs, the other 7 jobs structurally unchanged. Refs #981, #985
|
All four findings applied at The mechanism claim is corrected rather than merely softened. Both files now name
The vendor exclusion now states its own cost. #985 cited three times, and the comment admits "the Prose: both Legs re-run through the edited script with the tree proven: UNVERIFIED: the inherited per-leg figures for |
|
NEEDS CHANGES at The causal claim is now demonstrated, by a control neither earlier round ran. Same pre-fix tree, same That third row is what makes I have already corrected the title myself. It read "...to catch pytest-subtests-masked ordering Three new false claims, all verified by me:
Confirmed fixed: timeout 45 (adequate even at 2.5× the baseline), Its judgement on whether the job earns its cost, which I think is the most useful thing here. For And a sharper framing of the exclusion cost: the job's coverage is anti-correlated with where the A worker is on items 1-3. UNVERIFIED by the review: |
…ering leg (#981) Round 2 fixed a false `pytest-subtests` attribution but introduced three new claims of the same class -- prose asserting more than the code or the cross-reference delivers. All three are comment/docstring only. * The elapsed-time print claimed a future timeout would "leave behind a measured number". It could not: there was no `try`/`finally` in the file, so a step killed at `timeout-minutes` printed nothing. The print now runs in a `finally`, and the comment claims only what that buys -- a failing leg, and an unwind such as the `KeyboardInterrupt` CPython raises for SIGINT. SIGKILL runs no Python and SIGTERM has no default handler, so the comment says outright that an Actions timeout is still not covered. * The workflow's `tests/integration` bullet said the exclusion was "enforced in the script itself", while the script's docstring said the legs are defined in the workflow and not in the module. The script enforces nothing: `is_unit_tier()` rejects all 12 of that directory's `test_*.py` files, so `leg_modules('integration')` returns an empty tuple and the leg would run the root modules alone. The bullet now says that, keeping the two true reasons. * The "every leg measured well under a gigabyte of peak RSS" claim cited a workflow comment that holds wall times and no RSS at all. Replaced in both files with figures actually measured here via `resource.getrusage(RUSAGE_CHILDREN)`: cli 152 MiB, dumpkit 237 MiB, interface 305 MiB, with no figure claimed for the larger legs. Also reconciled two true-but-inconsistent passages: both files opened on "three test modules" while their own `--noconftest` paragraph said two, and `dumpkit` was called "the smallest of the multi-file directories" when `interface` has the same two unit-tier modules, 201 tests against 199, and runs ~2s faster. Behaviour unchanged. `yaml.safe_load` gives the same 5 top-level keys, 8 jobs, `unittest-ordering.timeout-minutes: 45`, and a byte-identical per-job digest for all 8 jobs. `py_compile` clean; cli (185 tests, 18.3s) and dumpkit (199 tests, 39.2s) legs both pass against the edited script.
|
All three round-2 items applied at The elapsed-print comment no longer claims what it cannot do, and the control flow earns the narrower The cross-file contradiction is resolved, with a precision I had got slightly wrong. I called the RSS is now measured rather than asserted, via One change beyond my brief, and it was the right call. The YAML carried the same unmeasured RSS Both "three modules"/"two modules" contradictions fixed — it found the YAML carried it too, which I had One defect it found and correctly left alone: UNVERIFIED by it: that two modules suffice to reproduce #981 (it relied on my measurement, since proving it |
…#981) `util/run_unittest_leg.py`'s docstring referenced `:data:`ROOT_MODULES``, a name the module does not define -- the real names are the `root_modules()` function and the `_EXCLUDED_ROOT_MODULES` frozenset. Same defect class as the `:data:`LEGS`` references round 1 caught: a dangling Sphinx target renders as plain text and warns about nothing, since the docs build sets neither `-n` nor `-W` and does not take `util/` in at all. Re-pointed at `:func:`root_modules``, which is what the sentence means -- the set of modules that run every invocation. That is also more precise than the old target implied, since `root_modules()` excludes `tests.test_tier_guard_xdist` and a constant listing every module would not have. Fixing the reference exposed a worse claim in the same sentence: it had the ordering backwards. It read "ahead of the directory's own modules -- not interleaved, and not after", while `build_suite()` queues `leg_modules()` first and `root_modules()` second. Measured on both the `dumpkit` and `cli` legs: the root-level modules occupy the last five positions, after the directory's own. Three other places in the tree already said so correctly -- `build_suite()`'s own docstring, the `_EXCLUDED_ROOT_MODULES` comment, and the workflow's "directory first and root second" -- as did the next sentence but one of this very paragraph, so the docstring was contradicting itself. Now reads "after the directory's own modules -- not before, and not interleaved", which is the order #981 reproduces in and the only order that reproduces it. Audited every remaining cross-reference in the file by resolving each target: 33 role occurrences, 25 of them resolvable roles, all 25 now resolve. Comment and docstring only; the workflow file is untouched this round. `py_compile` clean, and the `cli` leg still passes (185 tests, 18.3s).
|
The dangling cross-reference turned out to be hiding a worse defect: the docstring stated the leg's The ordering claim. Line 68 read "Every It was self-contradicting inside its own paragraph: two sentences later the same text says "Running the It only surfaced because the fix could not be applied without reading what the sentence asserted. Worth The reference itself is now Dangling-target audit: none beyond that one. It resolved every role rather than eyeballing them — One contextual finding I confirmed:
|
|
NEEDS CHANGES at The From And Why this one survived three rounds of careful verification, which is the part worth keeping. Rounds 1-3 Everything else checks out, including two bullets nobody had ever verified. The One unclaimed gap it names fairly: On the exclusion's missing trigger I agree it is a follow-up, not a blocker. UNVERIFIED by it: whether the job-timeout path funnels through that same |
#981) Problem: round 2's `finally`-block comment in util/run_unittest_leg.py drew an accurate pair of Python-signal facts (SIGKILL runs no Python code; SIGTERM has no default handler) into a false conclusion -- that a GitHub Actions `timeout-minutes` expiry is therefore not covered and a leg killed at the cap prints nothing. Neither signal is what the runner sends first. Fix: `actions/runner`'s `src/Runner.Sdk/ProcessInvoker.cs` sends SIGINT with a 7.5s grace period before SIGTERM (2.5s) and only then SIGKILL (`CancelAndKillProcessTree`, `_sigintTimeout`/`_sigtermTimeout`); `run:` steps select that ladder by passing `killProcessOnCancel: false` (`Handlers/ScriptHandler.cs`). Traced the job-level path too, not just the generic step-cancellation one the previous round verified: the backend's cancellation reaches the worker as a `CancelRequest` (`Runner.Worker/Worker.cs`), cancelling the token `JobRunner`/`StepsRunner` thread into the step's own `ExecutionContext.CancellationToken` -- the same token `ScriptHandler.cs` hands `ProcessInvoker.ExecuteAsync`. So a `timeout-minutes` expiry normally does still print, via the `KeyboardInterrupt` CPython raises for SIGINT, which `unittest.case._Outcome.testPartExecutor` re-raises rather than swallowing. Narrowed the real gap to a test blocked inside a C extension, which can defer signal delivery past SIGTERM/SIGKILL. No mirror of this claim exists in .github/workflows/unit-tests.yml -- round 2's diff there touched three other passages, not this one -- so that file is unchanged this round. Evidence: unittest.case._Outcome.testPartExecutor in this venv (Python 3.14.7) contains `except KeyboardInterrupt: raise`, verbatim. A real SIGINT sent to a child running the `cli` leg 3s in printed "tests/cli + 5 root module(s): interrupted before a result was available, 2.9s elapsed" and exited via signal 2, 0.148s after the signal. `py_compile` clean; `cli` leg still passes (185 tests, 18.4s elapsed) against the edited script. Refs #981
|
Round 3's finding is fixed at I told it the SIGKILL/SIGTERM passage was mirrored in the workflow comment. It is not, and never was. It It also closed the caveat I told it to keep, properly, by tracing the job-timeout path to primary source:
And it demonstrated the corrected claim instead of only asserting it. It started the The comment now names its sources inline — One open thread it names honestly, and I would rather leave it stated than closed by assertion: the
|
|
NEEDS CHANGES at Every round, mine included, reasoned about signal sending and signal handling. None asked which process
The review reproduced that end to end: at So round 2's "prints nothing" was right for the wrong reason, and round 4 replaced it with a densely-cited What the citations did and did not buy. Every path in the five-file chain checks out, and the boolean is I am having the paragraph deleted, not corrected a fifth time, and I think that is clearly right rather Its verdict on the artefact, which I agree with: merge after the paragraph is deleted, not after a fifth One residual it spotted: the text says the line buys "the number from a run that finished … close to the UNVERIFIED: |
#981) * The comment on `main`'s `finally` block has been wrong four times running, each fix introducing the next error: `pytest-subtests` masking the defect (not installed), the print surviving a kill (no `try`/`finally` existed yet), a timed-out leg printing nothing, and now a timeout expiry normally reaching the line. * The fourth is wrong too. Every round reasoned about which signal is sent and how Python handles it; none asked which process is signalled. `ProcessInvoker.SendSignal` is `kill()` on one positive pid -- no `killpg`/`setsid` on the Unix path -- and that pid is the step's shell, since the step is a bare `run:` and Python is bash's child. Non-interactive bash does not forward SIGINT to a foreground child (measured: child alive 2.5s after `kill -INT` of the bash pid), so both grace windows expire with the leg running and SIGKILL lands on bash. * So delete the paragraph rather than rewrite it a fifth time: a CI timeout says nothing about whether this line prints, and the claim was never load-bearing. What remains needs no CI justification -- the `finally` is there so a failing leg, and one unwinding on an exception, still report elapsed time. The residual "number from a run that finished" clause goes with it; a finished run prints by the normal path regardless. Comment-only: tokenising before and after, excluding comments and docstrings, gives an identical 552-token sequence. `py_compile` clean, workflow file byte-identical, `cli` leg still prints its elapsed line.
|
GOOD TO GO at 25 comment lines became 5, and the surviving sentence claims only what the code does: Gone: the I re-derived the comment-only proof independently rather than taking the worker's: tokenising the file with It declined to attempt a fifth rewrite and gave a better reason than mine. My argument was that four One thing it was properly honest about: it did not reproduce my Labelled Still open on this thread, neither a blocker: |
Please follow the guide below
You will be asked some questions, please read them carefully and answer honestly
Put an
xinto all the boxes [ ] relevant to your pull request (like that [x])Use Preview tab to see how your pull request will actually look like
Searched for similar pull requests
Followed the coding style (
make pylint,make mypy,make isort)make testpasses, and a test case covers the changeAdded a changelog entry under
docs/source/changelog/and regeneratedCHANGELOG.md, if the change is user-visible — N/A, centralised in docs(changelog): shared 1.5.0 changelog — long-lived, merges last (#610, #616, #617, #618, #620) #657What is the purpose of your pull request?
Tick the commit type your subject line carries.
fix— corrects a defectfeat— adds a featureperf— changes performance, not behaviourrefactor— changes neither behaviour nor performancetest— tests onlydocs— documentation onlyci— workflows or build toolingchore— anything elseDescription of your pull request and other information
Fixes the test half of #981 and adds the CI leg it calls for.
The defect.
RegistrationGateTestsboundEngine,EngineBase,Dumper,DumperBase,Reassembly,ReassemblyBase,TraceFlow,TraceFlowBaseandExtractoronce, at import time. A sibling module that purgespcapkitfromsys.modulesand reimports it (deliberately asymmetric; pytest's own autouse fixture reconciles it, plainunittestdoes not) leaves those names stale while the registration hooks they exercise see the fresh generation, so a dynamically-created subclass registers in one generation's registry while the test reads another's.setUpnow re-resolves everything viaimportlib.import_moduleinstead.Before:
tests.protocols.application.test_http_unit+tests.test_base_class_contract+tests.const.test_const_str_payload_870_unitunder plainunittest→Ran 73 tests ... FAILED (failures=1, errors=3). After:OK.The CI leg.
util/run_unittest_leg.py+ the newunittest-orderingjob run eachtests/subdirectory (paired with the root-level modules) underunittest.TextTestRunner, which has nopytest-subteststo swallow a subTest failure into a passing parent.tests/corekitandtests/vendorare excluded; see the job's own comment for the known, pre-existing reasons (one already documented, one newly found while sizing this leg and left for separate attention).tests/utilities/test_decorators.py's own leakedsys.tracebacklimitis a second instance of the same reporting gap, left untouched here — already being fixed as a side effect of #719.