push_repo_memory retry never refreshes its base: git ls-remote origin runs without credentials
Version: v0.89.21 (actions/setup/js/push_repo_memory.cjs), engine claude, repository visibility internal.
Summary
After #62414 removed the push_repo_memory concurrency group, concurrent memory writes are supposed to converge through the retry loop in pushRepoMemoryChangesWithRetry. On a private or internal repository they don't. On every retry, the base refresh fails:
ls-remote on retry failed, keeping existing baseRef: The process '/usr/bin/git' failed with exit code 128
So reconcileRepoMemoryRetry (the authenticated fetch, then rebase -X theirs --onto <remoteHead>) never runs. Each retry re-pushes on the same stale base until all 11 attempts are used up and the job fails.
Cause
- Line 133 points
origin at a URL with no credentials:
execGitSync(["remote", "set-url", "origin", originUrlForPush || `https://${serverHost}/${targetRepo}.git`], ...);
- The job's checkout step uses
persist-credentials: false, so no auth header is left on origin either.
- Line 175 then refreshes the head via that unauthenticated remote:
await execGetExecOutput("git", ["ls-remote", "origin", `refs/heads/${branchName}`], { cwd: workspaceDir });
On a non-public repository this exits 128, which the catch treats as "keep existing baseRef".
The authenticated URL (repoUrlWithToken) is already built at line 128, but it's only used inside reconcileRepoMemoryRetry, which this failure prevents from being reached.
Reproduce
- On a private or internal repository, run several workflows that share one
repo-memory branch and finish within a few seconds of each other. We fanned out ~60 workflows against the same branch-name.
- The first push wins. The others log, on every attempt:
! [rejected] repo-context/<x> -> repo-context/<x> (fetch first)
##[warning]Push failed (attempt N/11), retrying in …ms
ls-remote on retry failed, keeping existing baseRef: … exit code 128
- They then end in one of two ways:
- Different files changed.
pushSignedCommits tries its own rebase onto the GraphQL parent, which was never fetched locally:
fatal: Not a valid commit name <GraphQL parent sha>
##[warning]pushSignedCommits: replay parent <a> does not match GraphQL parent <b>; rebasing commit range before signed replay …
pushSignedCommits: processing commit 1/243 … # replaying other runs' commits
##[error]Failed to push changes after 11 attempts
We also saw this in a run whose own memory patch was 0 bytes.
- The same
.md file changed. pushSignedCommits' rebase has no -X theirs, so the documented "non-JSONL conflicts keep this run's files" behaviour never applies:
CONFLICT (content): Merge conflict in <file>.md
ERR_SYSTEM: pushSignedCommits: failed to rebase commit range onto current GraphQL parent
In our run, 12 of 62 workflows failed in push_repo_memory this way. Every other job in those runs succeeded.
Expected
A rejected push refreshes the remote head, rebases onto it, and succeeds on a later attempt, as the comment at lines 135–138 describes.
Suggested fix
Refresh the head through the authenticated URL already in scope, e.g.:
const { stdout: lsOut } = await execGetExecOutput("git", ["ls-remote", repoUrlWithToken, `refs/heads/${branchName}`], { cwd: workspaceDir });
(Or pass the git auth env / extraheader that pushSignedCommits uses.) Also, a failed ls-remote probably shouldn't fail silently into "keep the stale base": on a private repository that path can't succeed, so retrying on it only spends the attempts.
It may also be worth making pushSignedCommits' internal rebase fetch the GraphQL parent before rebasing onto it, since it currently assumes that commit already exists locally.
Workaround
We have each workflow write only its own memory file, so concurrent pushes don't touch the same file. That avoids the CONFLICT path but not the stale-base path.
push_repo_memoryretry never refreshes its base:git ls-remote originruns without credentialsVersion: v0.89.21 (
actions/setup/js/push_repo_memory.cjs), engineclaude, repository visibilityinternal.Summary
After #62414 removed the
push_repo_memoryconcurrency group, concurrent memory writes are supposed to converge through the retry loop inpushRepoMemoryChangesWithRetry. On a private or internal repository they don't. On every retry, the base refresh fails:So
reconcileRepoMemoryRetry(the authenticated fetch, thenrebase -X theirs --onto <remoteHead>) never runs. Each retry re-pushes on the same stale base until all 11 attempts are used up and the job fails.Cause
originat a URL with no credentials:persist-credentials: false, so no auth header is left onorigineither.catchtreats as "keep existing baseRef".The authenticated URL (
repoUrlWithToken) is already built at line 128, but it's only used insidereconcileRepoMemoryRetry, which this failure prevents from being reached.Reproduce
repo-memorybranch and finish within a few seconds of each other. We fanned out ~60 workflows against the samebranch-name.pushSignedCommitstries its own rebase onto the GraphQL parent, which was never fetched locally:.mdfile changed.pushSignedCommits' rebase has no-X theirs, so the documented "non-JSONL conflicts keep this run's files" behaviour never applies:In our run, 12 of 62 workflows failed in
push_repo_memorythis way. Every other job in those runs succeeded.Expected
A rejected push refreshes the remote head, rebases onto it, and succeeds on a later attempt, as the comment at lines 135–138 describes.
Suggested fix
Refresh the head through the authenticated URL already in scope, e.g.:
(Or pass the git auth env / extraheader that
pushSignedCommitsuses.) Also, a failedls-remoteprobably shouldn't fail silently into "keep the stale base": on a private repository that path can't succeed, so retrying on it only spends the attempts.It may also be worth making
pushSignedCommits' internal rebase fetch the GraphQL parent before rebasing onto it, since it currently assumes that commit already exists locally.Workaround
We have each workflow write only its own memory file, so concurrent pushes don't touch the same file. That avoids the
CONFLICTpath but not the stale-base path.