Skip to content

[copilot-opt] Copilot session logs (events.jsonl) missing from optimization audit data pipeline #50061

Description

@github-actions

Problem

No events.jsonl or per-session conversation logs were available under /tmp/gh-aw/agent/session-data/logs/ for any of the 50 sessions in the last-14-day analysis window, blocking tool-latency, payload-size, validation-timing, context-size, and prompt-drift analysis required by Phase 2 of the optimization audit.

Evidence

  • Analysis window: 2026-08-03T04:29:10Z to 2026-08-03T18:18:02Z (session data only covered the current day; the intended 14-day window could not be populated)
  • Sessions analyzed: 50 (metadata only: id, timestamps, conclusion, head_branch)
  • Key metrics and examples:
    • find /tmp/gh-aw/agent/session-data/logs -type f | wc -l returned 0 files
    • sessions-list.json only exposes coarse fields (conclusion, created_at, updated_at, head_branch, html_url) — no tool call records, no MCP latency, no payload sizes
    • Duration proxy (updated_at - created_at) was the only performance signal available; longest observed run was 1462s ("Running Copilot cloud agent"), but this cannot be decomposed into tool-level latency
    • Without events.jsonl, none of the required Phase 2 checks (slow MCP calls, oversized responses, late validation failures, context size, model-switch patterns, prompt drift) could be executed with real evidence

Proposed Change

Update the shared/copilot-session-data-fetch.md setup component to also download/extract per-run events.jsonl (or the raw conversation transcript) for each session in scope, using the GitHub Actions API/artifact download (e.g., gh run download or gh api .../logs) before the copilot-opt analysis step runs. Add a fallback that fails loudly (non-silently) when logs are missing so data-quality gaps are visible in the audit output rather than silently degrading analysis.

Expected Impact

  • Restores the ability to detect slow tool calls, oversized payloads, and prompt drift with quantitative evidence instead of proxy metrics
  • Prevents future optimization audits from silently operating on incomplete data
  • Establishes a clear failure signal so this fetch step itself becomes debuggable

Notes

  • Distinct root cause category: instruction/context reduction or restructuring (data pipeline gap upstream of context/tool analysis)
  • Data quality caveats: this entire audit relied on session metadata and PR data only; no direct tool-call telemetry was available

Generated by ⚡ Copilot Opt · auto · 55 AIC · ⌖ 4.37 AIC · ⊞ 8K · ◷

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions