Skip to content

CI: Plain unittest ordering (protocols) doubled and now cancels at its 45-minute cap on main #1029

Description

@JarryShaw

Plain unittest ordering (protocols) has roughly doubled and now hits its 45-minute cap on main itself, so the leg is cancelled on every run rather than reporting a result. It is not a required check — the required contexts are the five Compat Python 3.x legs plus Required checks passed — so merges are not blocked, but the leg currently tells us nothing.

The break is sharp and sits between 07:48Z and 10:49Z on 2026-10-05. Durations of that one job across consecutive runs:

run created branch head result duration
03:05Z fix/1021-... 9f8b774ac success 22.5m
03:44Z main 5fd0a1b11 success 27.0m
04:05Z main fa3e861f5 success 20.8m
04:32Z fix/1021-... 1f24b9aa9 success 21.5m
06:36Z docs/1024-... 8ae0d5eed success 22.7m
07:48Z docs/1024-... 983fb92a0 success 22.6m
10:49Z main cacbf2e70 cancelled 45.3m
10:50Z main aa5f01213 cancelled 55.0m
11:19Z main 51100da7e cancelled 45.4m
11:27Z fix/1021-... 39a4c4d8f cancelled 45.4m
11:37Z fix/1026-... 467eced4d cancelled 47.1m

It is the leg, not the runners. On main 51100da7e (bad) against main fa3e861f5 (good), every sibling leg is flat and only protocols moved:

leg good 04:05Z bad 11:19Z
protocols 20.8m 45.4m (cancelled)
const 5.6m 6.2m
foundation 5.5m 5.5m
toolkit 2.3m 1.8m
the other six 1.3–1.9m 1.1–1.7m

foundation being unchanged matters: it is the leg that exercises Extractor.

On cause, I have been wrong in both directions and am not asserting one now. I first attributed the slowdown to #1020 (which was the only slow branch at the time), then withdrew that after a profiling pass found identical cProfile call counts — 3,303,238 on both sides — and showed the new registrars run zero times per extract. That profiling was honest but narrow: it covered single extract calls and two test modules, never the whole leg in one process, because this host OOM-killed a 56 GB tests/protocols run. The before/after boundary does bracket #1020 and #1027 merging, so #1020 is back in scope — but a cumulative, whole-leg effect is precisely what was never measured, and the local evidence against a per-extract cost still stands.

What would settle it: run the full protocols selection on main fa3e861f5 and on main 51100da7e, in memory-bounded chunks summed to a total, and compare. If the totals match, the cause is CI-side and the fix is the leg's budget or its sharding; if post-#1020 is ~2x, it is cumulative and belongs in the code.

Either way the leg needs to stop silently cancelling — a timeout-minutes cut produces no summary line at all, so a real failure inside it would currently be invisible.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    ciPull requests that change CI or workflow configuration (ci: subject prefix)

    Projects

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions