Skip to content

ci: stop dropping a commit's validation when the next one lands - #394

Merged
JarryShaw merged 3 commits into
mainfrom
fix/ci-validate-every-commit
Sep 15, 2026
Merged

JarryShaw merged 3 commits into
mainfrom
fix/ci-validate-every-commit

Conversation

@JarryShaw

Copy link
Copy Markdown
Owner

You reported that main's pushes are no longer getting the full unit test and compatibility run. They are — but not reliably, and when they are lost it is silent. Here is what is actually happening.

The triggers are fine; the concurrency groups are not

on: push: branches: [main] is correct, and a quiet push to main does run everything. Verified on a71dd366d:

Unit Tests           -> 12 jobs (Python 3.10-3.15 + Integration 3.10-3.15), all success
Python Compatibility -> 6 jobs (Compat Python 3.10-3.15), all success

But e0f0160a3 — the #393 merge, one commit earlier — got nothing:

run 34970682137  event=push  head=e0f0160a3  conclusion=cancelled
  jobs: (none)
  created 12:44:53Z   cancelled 12:55:43Z

Zero jobs, and cancelled one second after the next push's run was created at 12:55:42Z. Its CodeQL, Conda Update and Vendor Update runs died the same way.

The cause is the concurrency groups I added in #392. Every workflow keyed its group on github.ref, so all of main's pushes shared one group — and GitHub cancels a pending run whenever a newer one joins the group. cancel-in-progress: false only protects a run that has already started; it does nothing for one still queued. So e0f0160a3's run sat pending behind e87ec884e's still-running one, and a71dd366d evicted it.

Net effect: merge two PRs a minute apart and the first commit reaches main validated by nothing. Exactly the hole you were worried about.

Fix

Non-pull-request events key on github.run_id, so each run gets its own group — nothing coalesces, nothing is dropped. Pull requests still key on the ref and still cancel in progress, which is where the saving belongs.

This does not undo #392. That PR's fan-out reduction came from removing the duplicate callers, not from cancellation, so the job counts per event are unchanged.

The workflows differ on purpose:

workflow group key (non-PR) cancel why
unit-tests run_id PR only every commit validated
python-compatibility run_id PR only every commit validated
codeql run_id PR only a commit with no analysis on record is a security gap
create-release run_id never dropped pending run = a release that never shipped
cron-conda run_id never publishes to Anaconda; value tied to when it ran
cron-vendor run_id never crawls IANA; two runs at one SHA are both meaningful
deploy-pages ref (unchanged) PR only publishes latest state — coalescing is correct here

deploy-pages is deliberately left alone: it publishes a latest state rather than a verdict on a commit, so dropping the pending build and letting the newer one publish is right, and deploying the older commit's docs after the newer ones would be actively wrong. Pages also permits only one live deployment.

One trap worth recording: github.run_id inside a called workflow is the caller's run id, so unit-tests keeps its literal unit-tests- prefix. Without it a callee would share its caller's group and — now that pushes do not cancel — sit pending behind its own caller forever.

Validation

All seven files parse; zero duplicate keys at any nesting level (the #375 bug class, checked with a recursive walk of the YAML node tree rather than safe_load, which silently keeps the last of a duplicate pair). The expression resolves as intended under GitHub's value-returning &&/||: true && ref → ref on a pull request; false && ref → false, then false || run_id → run_id on every other event.

Separate finding, and it needs your decision — not fixed here

Making the runs happen is only half of "validate no bad changes are being shipped". They also have to matter, and right now they do not:

$ gh api repos/JarryShaw/PyPCAPKit/branches/main/protection/required_status_checks
strict (branch must be up to date) = true
required contexts = []
required checks   = []

No check is required to merge. A PR with red unit tests is mergeable today; branch protection only insists the branch be current. So even with this PR merged, nothing stops a failing commit from landing — it just gets a red mark afterwards.

Fixing that means adding required contexts, which is a repo-settings change rather than a code change, so I have left it to you. Note the check names changed in #392, so the list would be the current ones — Python 3.10…3.15, Integration Python 3.10…3.15, Compat Python 3.10…3.15. Worth choosing deliberately: requiring a context that never reports on some event makes PRs permanently unmergeable, so the safe set is the ones that run on every pull request.

Merging two PRs a minute apart left the first commit with no unit tests, no
compatibility run and no CodeQL analysis. Not a flake, and not the triggers --
`on: push: branches: [main]` is correct and the full twelve-job matrix does run
on a quiet push. The loss came from the concurrency groups added in #392.

Every workflow keyed its group on `github.ref`, so all of main's pushes shared
one group. GitHub cancels a *pending* run whenever a newer one joins the group:
`cancel-in-progress: false` protects a run that has already started, never one
still queued. Measured on run 34970682137 for e0f0160 (the #393 merge), which
finished `cancelled` with **zero jobs**, one second after the next push's run was
created. That commit is on main having been validated by nothing at all, along
with its CodeQL, Conda and Vendor runs.

Non-pull-request events now key on `github.run_id`, giving each run its own
group, so nothing coalesces and nothing is dropped. Pull requests still key on
the ref and still cancel in progress, which is where the saving belongs and is
untouched -- this does not undo #392's fan-out reduction, which came from
removing the duplicate callers rather than from cancellation.

Per workflow, and they differ on purpose:

* unit-tests, python-compatibility, codeql -- per-run for pushes, so every
  commit on main is validated. unit-tests keeps its literal `unit-tests-` prefix
  because `github.run_id` inside a called workflow is the *caller's* run id, so
  without the prefix a callee would sit pending behind its own caller and
  deadlock.
* create-release, cron-conda, cron-vendor -- per-run, never cancelled. These tag,
  upload and publish, so a dropped pending run is a release or a registry update
  that never happened. The crons additionally derive their value from *when* they
  ran rather than from which commit, so two scheduled runs at one SHA are both
  meaningful.
* deploy-pages -- deliberately left keyed on the ref. It publishes a latest state
  rather than a verdict on a commit, so coalescing is correct, and publishing the
  older commit's docs after the newer ones would be wrong. Pages also allows only
  one live deployment.

Validated: all seven files parse, zero duplicate keys at any nesting level, and
the expression `${{ github.event_name == 'pull_request' && github.ref ||
github.run_id }}` resolves to the ref on a pull request and to the run id on every
other event, per GitHub's value-returning `&&`/`||`.

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Changes recommended

create-release.yml now uses a per-run concurrency group, which does not serialize release runs and can reintroduce overlapping tag/release races described in the workflow’s own comments.

Get a fresh assessment by requesting another Copilot review.

Pull request overview

This PR adjusts GitHub Actions concurrency group keys so that non-PR events (notably push to main) don’t silently lose an entire queued validation run when a newer commit lands shortly after.

Changes:

  • Update validation workflows (Unit Tests, Python Compatibility, CodeQL) to key non-PR concurrency groups on github.run_id, while keeping PR runs keyed on github.ref and cancellable.
  • Switch maintenance/publishing workflows (Conda Update, Vendor Update, Create Release) to run_id-keyed groups with no cancellation to avoid dropped pending runs.
  • Add explanatory comments, and explicitly retain ref-keyed coalescing for Pages deployments.
File summaries
File Description
.github/workflows/unit-tests.yml Use run_id for non-PR concurrency to prevent dropped main-push validation runs.
.github/workflows/python-compatibility.yml Same concurrency strategy as unit tests for compatibility checks.
.github/workflows/codeql-analysis.yml Same concurrency strategy to avoid missing security scans on main pushes.
.github/workflows/create-release.yml Changes concurrency group to run_id with added rationale comment.
.github/workflows/cron-conda.yml Changes concurrency group to run_id with added rationale comment.
.github/workflows/cron-vendor.yml Changes concurrency group to run_id with added rationale comment.
.github/workflows/deploy-pages.yml Keeps ref-keyed concurrency and documents why coalescing is desired for Pages.
Review details
  • Files reviewed: 7/7 changed files
  • Comments generated: 3
  • Review effort level: Lite

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread .github/workflows/create-release.yml
Comment thread .github/workflows/cron-conda.yml Outdated
Comment thread .github/workflows/cron-vendor.yml Outdated
Copilot was right, and about create-release it was right in the way that matters.
The first pass moved all six non-Pages workflows to per-run groups, which fixed
dropped validation runs but silently removed serialisation from the paths that
publish. Those are not the same trade:

* a dropped *validation* run means a commit nobody ever checked, and there is no
  second chance at it;
* a dropped *publishing* run means a publish that did not happen twice.

So per-run is now only on unit-tests, python-compatibility and codeql, and
create-release, cron-conda and cron-vendor go back to ref-keyed.

create-release is the sharp case: unlike the two crons it has **no** job-level
concurrency anywhere, so that group was the only thing stopping two runs tagging
and uploading to PyPI and Anaconda at once. Per-run would have allowed exactly
the race the file's own comment warned about.

Also corrected the cron comments, which claimed the workflow-level group queued
superseded runs. It does not, and it is not what protects the repository writes
either -- the `vendor-update` and `conda-update` jobs each carry their own
`repository-maintenance` group, shared between the two workflows, and that is
what actually serialises the commits. The comments now say so, so the next reader
does not credit the wrong mechanism.

Validated: all seven parse, zero duplicate keys at any nesting level.
@JarryShaw
JarryShaw requested a lite review from Copilot September 15, 2026 16:06
@JarryShaw

Copy link
Copy Markdown
Owner Author

Copilot's call on create-release was correct, and it caught a real regression I introduced. Fixed in b6c584522.

The first pass moved all six non-Pages workflows to per-run groups. That fixed the dropped-validation problem but quietly removed serialisation from the paths that publish — and those two are not the same trade:

  • a dropped validation run means a commit nobody ever checked, and there is no second chance at it;
  • a dropped publishing run means a publish that did not happen twice.

So per-run now applies only to unit-tests, python-compatibility and codeql. create-release, cron-conda and cron-vendor are back to ref-keyed.

create-release is the sharp case, exactly as Copilot said. I checked the job-level concurrency across all four:

workflow job-level concurrency
cron-conda conda-update → repository-maintenance
cron-vendor vendor-update → repository-maintenance
deploy-pages deploy-pages → ci-${{ github.ref }}
create-release none, on any of its six jobs

So for create-release the workflow-level group was the only thing stopping two runs tagging and uploading to PyPI and Anaconda concurrently. Per-run would have permitted precisely the race the file's own comment warns about.

And Copilot was right about the comments on both crons. They claimed the workflow-level group "queue[s] superseded runs" — it does not, and it is also not what protects the repository writes. That is the repository-maintenance job-level group, shared between the two workflows so neither can commit while the other is committing. The comments now say so, so the next reader credits the right mechanism instead of the group name.

Final split, for the record:

workflow key rationale
unit-tests, python-compatibility, codeql run_id (non-PR) a dropped run is an unchecked commit
create-release ref only serialisation it has; concurrent publish is worse
cron-conda, cron-vendor ref dropped duplicate is harmless; writes serialised at job level
deploy-pages ref publishes latest state; coalescing is correct

All seven files parse with zero duplicate keys at any nesting level.

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot encountered an error and was unable to review this pull request. You can try again by re-requesting a review.

@JarryShaw
JarryShaw merged commit eabf894 into main Sep 15, 2026
24 checks passed
@JarryShaw
JarryShaw deleted the fix/ci-validate-every-commit branch September 17, 2026 01:08
@JarryShaw JarryShaw added the ci Pull requests that change CI or workflow configuration (ci: subject prefix) label Sep 22, 2026
@JarryShaw JarryShaw added this to the 1.5 milestone Oct 6, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ci Pull requests that change CI or workflow configuration (ci: subject prefix)

Projects

Status: Done

Development

Successfully merging this pull request may close these issues.

2 participants