You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
CI: cut job count and duplicated work — contention, not job duration, is what makes CI slow #715
Splitting the structural half of #713 out, so the one-line ceiling raise there can land and close on its own
instead of waiting on measurement.
The finding that reframed this
Raising timeout-minutes does not make CI faster — it stops jobs dying. The actual cost is runner
contention, measured across five real runs:
run
condition
sum(execution)
wall-clock to done
35890190308
main push, quiet queue
272 min
31 min
35867536092
PR, busy queue
292 min
178 min
35867531818
PR, busy queue
290 min
161 min
35867526185
PR, busy queue
284 min
103 min
35867523310
PR, busy queue
283 min
115 min
~285 minutes of real execution every time; wall-clock to green swings 6× purely on how contended the runner
pool was, with individual jobs staggered 82-161 minutes apart waiting for a runner. The fix is therefore to remove jobs and duplicated work, not to give each job a longer leash.
Per-job execution for reference: unit 14.4-27.6 min, integration 17.4-29.4 min, gate ~27.2 min, all six Compat Python jobs 1.4 min combined (they run compileall plus an import smoke test, no pytest), changelog ~6-10 s, lint 4.0 min, CodeQL 2.4 min. A PR push is 22 jobs, 12 of which run pytest; a push
to main is 32 jobs. The repository is public, so Linux minutes are free — the cost is queue depth and
wall-clock, not billing.
Already shipped — listed so nobody proposes them again
gate-only (e87ec884e, ci: run the unit tier once per event, not once per caller #392): unit-tests.yml:10 defines the input, jobs gate on it at :51, :93, :205, and all four callers pass gate-only: true (deploy-pages.yml:39, cron-vendor.yml:29, cron-conda.yml:31, create-release.yml:87). Reusable-workflow callers already run only the single gate job, not the 12-job matrix.
Item 1 — partition the integration job to its own files
unit-tests.yml:156 runs a bare python -m pytest -q, i.e. the whole suite, while the test job at :84-89
runs pytest -q --ignore=tests/integration --ignore-glob='*_runtime.py' --ignore-glob='*_regression.py'. So each
of the six Integration Python 3.1x jobs re-runs every one of the ~118 files the matching unit job just ran,
adding only ~30 files of its own.
Change :156 to run only tests/integration/ plus the two globs the unit job excludes.
Expected saving: large but not yet quantified — it depends on whether integration's ~30 unique files are
cheap or expensive relative to the ~118 shared ones. A per-directory profile is the input needed before landing
this, and that measurement is in progress.
Risk to check first: whether any unit-scoped test implicitly depends on fixtures examples/generators/make_samples.py generates only inside the integration and gate jobs. If so, those tests
pass today only because integration re-runs them after fixture generation, and partitioning would break them
silently.
Item 2 — stop the gate re-running on every push to main
The gate job (Gate (full suite, Python 3.14), :203) is byte-for-byte identical work to Integration Python 3.14 from the same commit: same interpreter, same bare pytest -q, same tree. On a push to main it fires up to four times, once per gate-only caller, on top of the matrix that already ran.
Replace the push-triggered gate-only: true calls in deploy-pages.yml, cron-vendor.yml and cron-conda.yml with a workflow_run trigger on Unit Tests completion, gated on conclusion == 'success' and
checked out at head_sha. create-release.yml:6-9 already uses exactly this pattern, chaining off Vendor Update's completion, so there is a working template in-tree.
Saving: ~108 min of duplicate execution per main push; main-push job count 32 → ~20-23. Signal lost: none.
Item 3 — move Python 3.15 off the PR-blocking matrix
Ruleset 23497679 requires 15 contexts — Python, Integration Python and Compat Python for 3.10-3.14
only. 3.15 is verified absent from all fifteen, so it cannot block a merge, yet it costs 3 jobs per PR
push (22 → 19) and its legs are among the longest-running.
Remove the 3.15 legs from unit-tests.yml's test and integration matrices and from python-compatibility.yml, and add a nightly or weekly schedule instead.
Signal lost: none for merge-gating. Only cost is detection latency — a 3.15 regression surfaces on the next
scheduled run rather than in the PR.
Item 4 — pytest-xdist, undecided
pytest-xdist is absent from pyproject.toml and all three pytest invocations are bare pytest -q, so there is
no parallelism anywhere. -n auto would dwarf items 1-3 if it works.
The reason it may not: this suite mutates global state that parallel workers would race on — the registries
under pcapkit/foundation/registry/**, aenum.extend_enum adding members to shared enum classes at runtime
(measured at 16.7% of extraction self-time in #575), the sys.modules swapping that #674/#687/#688 already
had to fix once, and __warningregistry__ ordering that warning-count assertions depend on. Subtest-heavy
sweeps are also indivisible units — a file reporting thousands of subtests lands in one worker and caps the
achievable speedup.
A feasibility assessment is in progress. Do not add -n until it reports; a suite that passes serially and
fails under -n is worse than a slow one.
Ordering
Item 1 is the largest win for the PR queue specifically and should go first once its measurement lands. Item 2 is
the largest absolute win and can land immediately — it needs no measurement and loses no signal. Item 3 is the
cheapest and safest. Item 4 is gated.
Each item is independent and should be its own PR; they touch different files and carry different risk.
Splitting the structural half of #713 out, so the one-line ceiling raise there can land and close on its own
instead of waiting on measurement.
The finding that reframed this
Raising
timeout-minutesdoes not make CI faster — it stops jobs dying. The actual cost is runnercontention, measured across five real runs:
3589019030835867536092358675318183586752618535867523310~285 minutes of real execution every time; wall-clock to green swings 6× purely on how contended the runner
pool was, with individual jobs staggered 82-161 minutes apart waiting for a runner. The fix is therefore to
remove jobs and duplicated work, not to give each job a longer leash.
Per-job execution for reference: unit 14.4-27.6 min, integration 17.4-29.4 min, gate ~27.2 min, all six
Compat Pythonjobs 1.4 min combined (they runcompileallplus an import smoke test, no pytest),changelog~6-10 s,lint4.0 min, CodeQL 2.4 min. A PR push is 22 jobs, 12 of which run pytest; a pushto
mainis 32 jobs. The repository is public, so Linux minutes are free — the cost is queue depth andwall-clock, not billing.
Already shipped — listed so nobody proposes them again
gate-only(e87ec884e, ci: run the unit tier once per event, not once per caller #392):unit-tests.yml:10defines the input, jobs gate on it at:51,:93,:205, and all four callers passgate-only: true(deploy-pages.yml:39,cron-vendor.yml:29,cron-conda.yml:31,create-release.yml:87). Reusable-workflow callers already run only the singlegatejob, not the 12-job matrix.concurrency:withcancel-in-progress(eabf89417, ci: stop dropping a commit's validation when the next one lands #394): present on all eight workflows.unit-tests.yml:38-40keys ongithub.refforpull_requestwithcancel-in-progress: true.Item 1 — partition the integration job to its own files
unit-tests.yml:156runs a barepython -m pytest -q, i.e. the whole suite, while thetestjob at:84-89runs
pytest -q --ignore=tests/integration --ignore-glob='*_runtime.py' --ignore-glob='*_regression.py'. So eachof the six
Integration Python 3.1xjobs re-runs every one of the ~118 files the matching unit job just ran,adding only ~30 files of its own.
Change
:156to run onlytests/integration/plus the two globs the unit job excludes.Expected saving: large but not yet quantified — it depends on whether integration's ~30 unique files are
cheap or expensive relative to the ~118 shared ones. A per-directory profile is the input needed before landing
this, and that measurement is in progress.
Risk to check first: whether any unit-scoped test implicitly depends on fixtures
examples/generators/make_samples.pygenerates only inside the integration and gate jobs. If so, those testspass today only because integration re-runs them after fixture generation, and partitioning would break them
silently.
Item 2 — stop the gate re-running on every push to
mainThe
gatejob (Gate (full suite, Python 3.14),:203) is byte-for-byte identical work toIntegration Python 3.14from the same commit: same interpreter, same barepytest -q, same tree. On a push tomainit fires up to four times, once pergate-onlycaller, on top of the matrix that already ran.Replace the
push-triggeredgate-only: truecalls indeploy-pages.yml,cron-vendor.ymlandcron-conda.ymlwith aworkflow_runtrigger onUnit Testscompletion, gated onconclusion == 'success'andchecked out at
head_sha.create-release.yml:6-9already uses exactly this pattern, chaining offVendor Update's completion, so there is a working template in-tree.Saving: ~108 min of duplicate execution per main push; main-push job count 32 → ~20-23.
Signal lost: none.
Item 3 — move Python 3.15 off the PR-blocking matrix
Ruleset
23497679requires 15 contexts —Python,Integration PythonandCompat Pythonfor 3.10-3.14only. 3.15 is verified absent from all fifteen, so it cannot block a merge, yet it costs 3 jobs per PR
push (22 → 19) and its legs are among the longest-running.
Remove the 3.15 legs from
unit-tests.yml'stestandintegrationmatrices and frompython-compatibility.yml, and add a nightly or weekly schedule instead.Signal lost: none for merge-gating. Only cost is detection latency — a 3.15 regression surfaces on the next
scheduled run rather than in the PR.
Item 4 —
pytest-xdist, undecidedpytest-xdistis absent frompyproject.tomland all three pytest invocations are barepytest -q, so there isno parallelism anywhere.
-n autowould dwarf items 1-3 if it works.The reason it may not: this suite mutates global state that parallel workers would race on — the registries
under
pcapkit/foundation/registry/**,aenum.extend_enumadding members to shared enum classes at runtime(measured at 16.7% of extraction self-time in #575), the
sys.modulesswapping that #674/#687/#688 alreadyhad to fix once, and
__warningregistry__ordering that warning-count assertions depend on. Subtest-heavysweeps are also indivisible units — a file reporting thousands of subtests lands in one worker and caps the
achievable speedup.
A feasibility assessment is in progress. Do not add
-nuntil it reports; a suite that passes serially andfails under
-nis worse than a slow one.Ordering
Item 1 is the largest win for the PR queue specifically and should go first once its measurement lands. Item 2 is
the largest absolute win and can land immediately — it needs no measurement and loses no signal. Item 3 is the
cheapest and safest. Item 4 is gated.
Each item is independent and should be its own PR; they touch different files and carry different risk.