Skip to content

fix(release-please): serialize runs per repo with a concurrency group - #39

Closed
kojiromike wants to merge 3 commits into
mainfrom
fix/release-please-concurrency
Closed

kojiromike wants to merge 3 commits into
mainfrom
fix/release-please-concurrency

Conversation

@kojiromike

@kojiromike kojiromike commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

GitHub sometimes starts this workflow twice for a single push to main. Both runs build the same release commit, and the slower one fails with Error updating ref heads/release-please--branches--main--.... That leaves a red run on main even though the release PR is correct. Example: chart-oce-openemr run 36780271346 vs. its twin 36780270472.

This change puts the reusable job in a repo-wide release-please-reusable concurrency group without cancel-in-progress. The second run then waits and finds nothing left to update.

  • Per repo, not per ref: release-please always targets the default branch, whichever ref triggered the run.
  • Reserved group name: a caller that declares a group with the same name would deadlock with this job. No current caller declares any concurrency group.
  • One pending run per group: a third overlapping run cancels the one already waiting, and the canceled run shows as a failed check on its commit. No release work is lost, because the replacement reads the branch's current HEAD rather than the commit that triggered it. This will happen in practice: replaying each caller's last 100 runs through a one-running, one-pending queue predicts 37 canceled runs out of 1,354, across 9 repos, where merges land seconds apart. Those same runs had 28 failures from all causes. queue: max raises the limit to 100 pending runs, but actionlint rejects the key today (release-please: use queue: max on the concurrency group once actionlint accepts it #40).
  • 30-minute job timeout: a hung run now holds the group, so it must not inherit GitHub's six-hour default. The longest job in the last 100 Release Please runs of any caller took 277 seconds.

After the next release, Dependabot opens a pin bump in the 21 callers whose Dependabot config covers github-actions. The other 22 (the tfm-* modules, four chart-oce-* charts and two oce-py-* tools, pinned at @1.0.0 or @1.0.2) have no such config and need a manual bump.

GitHub occasionally delivers a single push to main as two workflow runs.
Both runs build the identical release commit, and whichever loses the race
fails with "Error updating ref heads/release-please--branches--main--...",
leaving a red run on main even though the release PR is correct
(chart-oce-openemr run 36780271346).

Put the job in a repo-wide concurrency group without cancel-in-progress so
the second run queues, then no-ops against the already-updated PR. Group
per repo rather than per ref because release-please always targets the
default branch, and use a name callers won't reuse to avoid a caller/callee
deadlock.

Assisted-by: Claude Code
The concurrency group's comment and README paragraph said overlapping
runs queue. GitHub keeps only one pending run per group, so a third
overlapping run cancels the one already waiting, and the canceled run
shows as a failed check on its commit. No release work is lost, because
release-please reads the default branch's current HEAD rather than the
triggering commit, but callers should expect the canceled run.

`queue: max` raises the limit to 100 pending runs. actionlint rejects
the key today, so that change is tracked in #40.

Also drop the claim that the group name is distinct from any name a
caller might declare. Nothing enforces that; the README warning is the
only guard.

Assisted-by: Claude Code
The job had no timeout, so it inherited GitHub's 360-minute default.
That was harmless while runs were independent. Now that the job holds a
repo-wide concurrency group, a hung run would stall every later release
PR update and release in that repo for up to six hours.

Cap the job at 30 minutes. The longest job in the last 100 Release
Please runs of any caller took 277 seconds, so the cap leaves about six
times that.

Assisted-by: Claude Code
@kojiromike

Copy link
Copy Markdown
Contributor Author

Self-reviewed: 5 local passes (3 author, 2 independent) at b9ab5ef, converged. Complex change (shared workflow that every caller picks up).

When (UTC) Head Complexity Passes Result
2026-10-01 03:07 b9ab5ef complex: concurrency group and job timeout on a shared reusable workflow 5 (3 author, 2 independent) converged

@kojiromike kojiromike closed this Oct 1, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant