.github/workflows/unit-tests.yml:96 sets timeout-minutes: 30 on the integration job. Measured, that is
below the job's worst-case runtime, so integration checks are being killed mid-suite on PRs that have nothing
wrong with them.
The measurement
Durations of integration jobs that succeeded, sampled across branches and Python versions from recent
Unit Tests runs:
| duration |
job |
| 29 min |
Integration Python 3.10 |
| 28 min |
Integration Python 3.15 |
| 28 min |
Integration Python 3.12 |
| 27 min |
Integration Python 3.14 |
| 26 min |
Integration Python 3.14 |
| 26 min |
Integration Python 3.10 |
| 21 min |
Integration Python 3.15 |
| 21 min |
Integration Python 3.14 |
| 20 min |
Integration Python 3.13 |
The slowest passing run is 29 minutes against a 30-minute ceiling — one minute of margin.
Two jobs already killed
| PR |
job id |
started |
killed |
duration |
step reached |
| #691 |
107202753543 |
14:12:13Z |
14:42:32Z |
30 min 19 s |
Run full test suite, steps 1-5 all green |
| #695 |
107202791532 |
14:31:48Z |
15:02:05Z |
30 min 17 s |
Run full test suite, steps 1-5 all green |
Both are Integration Python 3.12, but that is coincidence rather than a property of 3.12 — the table above
shows 3.10, 3.14 and 3.15 all landing at 26-29 minutes. Whichever job draws a slow runner tips over.
Why this blocks merges rather than just being noisy
Integration Python 3.10 through 3.14 are five of the fifteen required status checks in ruleset 23497679.
A timed-out job is not re-run automatically, so any PR that loses one cannot merge until someone re-runs it
by hand — and re-running into a busy queue can hit the same wall.
A timeout is reported as cancelled, not failure. That makes it easy to miss: a conclusion-based failure
tally returns zero while statusCheckRollup.state reads FAILURE, and the usual explanation for a cancelled
job — a force-push superseding the run — does not apply. Anything automating over these checks needs to treat
cancelled as "needs attention" rather than folding it in with superseded runs.
Suggested fix
Raise timeout-minutes on the integration job from 30 to 45, roughly 1.55x the observed maximum. That
keeps a real runaway bounded while leaving headroom for a slow runner. 45 is a suggestion, not a measurement —
the right number is a judgement about CI budget.
Worth considering alongside it, though out of scope here: the integration job runs the full suite per Python
version with no parallelism inside the job, which is why it sits near half an hour. Sharding it or using
pytest-xdist would cut wall-clock more durably than raising the ceiling, at the cost of a more complex
workflow.
Related sites with the same setting, unchanged by this and listed for completeness: unit-tests.yml:54,
unit-tests.yml:207, lint.yml:72 (all 30), and unit-tests.yml:179 (5).
.github/workflows/unit-tests.yml:96setstimeout-minutes: 30on theintegrationjob. Measured, that isbelow the job's worst-case runtime, so integration checks are being killed mid-suite on PRs that have nothing
wrong with them.
The measurement
Durations of integration jobs that succeeded, sampled across branches and Python versions from recent
Unit Testsruns:Integration Python 3.10Integration Python 3.15Integration Python 3.12Integration Python 3.14Integration Python 3.14Integration Python 3.10Integration Python 3.15Integration Python 3.14Integration Python 3.13The slowest passing run is 29 minutes against a 30-minute ceiling — one minute of margin.
Two jobs already killed
10720275354314:12:13Z14:42:32ZRun full test suite, steps 1-5 all green10720279153214:31:48Z15:02:05ZRun full test suite, steps 1-5 all greenBoth are
Integration Python 3.12, but that is coincidence rather than a property of 3.12 — the table aboveshows 3.10, 3.14 and 3.15 all landing at 26-29 minutes. Whichever job draws a slow runner tips over.
Why this blocks merges rather than just being noisy
Integration Python 3.10through3.14are five of the fifteen required status checks in ruleset23497679.A timed-out job is not re-run automatically, so any PR that loses one cannot merge until someone re-runs it
by hand — and re-running into a busy queue can hit the same wall.
A timeout is reported as
cancelled, notfailure. That makes it easy to miss: a conclusion-based failuretally returns zero while
statusCheckRollup.statereadsFAILURE, and the usual explanation for acancelledjob — a force-push superseding the run — does not apply. Anything automating over these checks needs to treat
cancelledas "needs attention" rather than folding it in with superseded runs.Suggested fix
Raise
timeout-minuteson theintegrationjob from30to 45, roughly 1.55x the observed maximum. Thatkeeps a real runaway bounded while leaving headroom for a slow runner.
45is a suggestion, not a measurement —the right number is a judgement about CI budget.
Worth considering alongside it, though out of scope here: the integration job runs the full suite per Python
version with no parallelism inside the job, which is why it sits near half an hour. Sharding it or using
pytest-xdistwould cut wall-clock more durably than raising the ceiling, at the cost of a more complexworkflow.
Related sites with the same setting, unchanged by this and listed for completeness:
unit-tests.yml:54,unit-tests.yml:207,lint.yml:72(all30), andunit-tests.yml:179(5).