Skip to content

Copilot CLI 1.0.85 (gh-aw v0.89.17) returns instant HTTP 400 on first request, breaking agent runs #62363

Description

@fr4nc1sc0-r4m0n

Motivation

We upgraded the gh-aw compiler pin in elastic/oblt-aw from v0.88.7 to v0.89.17 (elastic/oblt-aw#2090), which also bumped the pinned Copilot CLI engine from 1.0.80 to 1.0.85. The very next live run of an agentic workflow on the new pin failed instantly, and it kept failing on rerun, so we had to revert (elastic/oblt-aw#2100) to restore the workflow.

Summary

After upgrading to gh-aw v0.89.17 / Copilot CLI 1.0.85, our status-triggered agentic workflow (gh-aw-estc-pr-buildkite-detective, model gpt-5.3-codex) fails on the very first request with an HTTP 400 from api.githubcopilot.com/responses, before the agent does any work or consumes any tokens.

Evidence

  • Run: https://github.com/elastic/oblt-aw/actions/runs/35592278916 — failed on first trigger right after the upgrade landed.
  • Rerun of the same run: identical failure, different resume session ID (34b9f80c-a70a-42a7-993b-d59f88da5e21 -> ca37fb75-46d8-4d3e-a39e-ef399a4d27de), confirming this is not stale/resumed session state (per http_400_response_error.md's hypothesis Add workflow: githubnext/agentics/weekly-research #3) but a fresh first request failing both times.
  • Downstream impact: our E2E harness caught this as a missing PR comment (agent_comment_present: comment=None) in run https://github.com/elastic/oblt-aw/actions/runs/35592128904.
  • copilot-harness log excerpt from agent-stdio.log:
    [copilot-harness] attempt 1: process started (pid=390)
    400 Bad Request
    Changes    +0 -0
    Duration   0s
    Resume     copilot --resume=ca37fb75-46d8-4d3e-a39e-ef399a4d27de
    [copilot-harness] attempt 1: process exit event exitCode=1
    [copilot-harness] attempt 1 failed: exitCode=1 failureClass=http_400_response_error ... isHTTP400ResponseError=true ... tokenCount=0 attemptDurationMs=1497 retriesRemaining=3
    [copilot-harness] attempt 1: HTTP 400 response error — not retrying (persistent request validation/state failure)
    
  • The upstream UPSTREAM_ERROR_RESPONSE captured by our API proxy shows the 400 body is GitHub's generic edge/WAF "Bad request" HTML page (title: Bad request - GitHub), not a structured Copilot API error payload — suggesting the request itself is being rejected before it reaches the model backend.
  • We diffed the compiled .lock.yml between the two compiler versions: the only functional change affecting this workflow is the pinned setup-cli/setup action moving from v0.88.7 to v0.89.17 (installing Copilot CLI 1.0.85 instead of 1.0.80). Our workflow's frontmatter, allowed tools, --add-dir args, and prompt content are unchanged.

What we did

Reverted the compiler/engine pin back to v0.88.7 / Copilot CLI 1.0.80 in elastic/oblt-aw#2100 to restore the workflow to a working state, pending investigation here.

Ask

Could you help root-cause why Copilot CLI 1.0.85 produces a malformed/rejected /responses request on its very first call for this workflow shape (model gpt-5.3-codex, tools: github, public-code-search(search_code), safeoutputs, shell, web_fetch, write)? Happy to share the full agent-stdio.log, token-tracker-audit.jsonl, and prompt file if useful.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Labels

No labels
No labels

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions