Problem
No events.jsonl or per-session conversation logs were available under /tmp/gh-aw/agent/session-data/logs/ for any of the 50 sessions in the last-14-day analysis window, blocking tool-latency, payload-size, validation-timing, context-size, and prompt-drift analysis required by Phase 2 of the optimization audit.
Evidence
- Analysis window: 2026-08-03T04:29:10Z to 2026-08-03T18:18:02Z (session data only covered the current day; the intended 14-day window could not be populated)
- Sessions analyzed: 50 (metadata only: id, timestamps, conclusion, head_branch)
- Key metrics and examples:
find /tmp/gh-aw/agent/session-data/logs -type f | wc -l returned 0 files
sessions-list.json only exposes coarse fields (conclusion, created_at, updated_at, head_branch, html_url) — no tool call records, no MCP latency, no payload sizes
- Duration proxy (
updated_at - created_at) was the only performance signal available; longest observed run was 1462s ("Running Copilot cloud agent"), but this cannot be decomposed into tool-level latency
- Without
events.jsonl, none of the required Phase 2 checks (slow MCP calls, oversized responses, late validation failures, context size, model-switch patterns, prompt drift) could be executed with real evidence
Proposed Change
Update the shared/copilot-session-data-fetch.md setup component to also download/extract per-run events.jsonl (or the raw conversation transcript) for each session in scope, using the GitHub Actions API/artifact download (e.g., gh run download or gh api .../logs) before the copilot-opt analysis step runs. Add a fallback that fails loudly (non-silently) when logs are missing so data-quality gaps are visible in the audit output rather than silently degrading analysis.
Expected Impact
- Restores the ability to detect slow tool calls, oversized payloads, and prompt drift with quantitative evidence instead of proxy metrics
- Prevents future optimization audits from silently operating on incomplete data
- Establishes a clear failure signal so this fetch step itself becomes debuggable
Notes
- Distinct root cause category: instruction/context reduction or restructuring (data pipeline gap upstream of context/tool analysis)
- Data quality caveats: this entire audit relied on session metadata and PR data only; no direct tool-call telemetry was available
Generated by ⚡ Copilot Opt · auto · 55 AIC · ⌖ 4.37 AIC · ⊞ 8K · ◷
Problem
No
events.jsonlor per-session conversation logs were available under/tmp/gh-aw/agent/session-data/logs/for any of the 50 sessions in the last-14-day analysis window, blocking tool-latency, payload-size, validation-timing, context-size, and prompt-drift analysis required by Phase 2 of the optimization audit.Evidence
find /tmp/gh-aw/agent/session-data/logs -type f | wc -lreturned 0 filessessions-list.jsononly exposes coarse fields (conclusion,created_at,updated_at,head_branch,html_url) — no tool call records, no MCP latency, no payload sizesupdated_at - created_at) was the only performance signal available; longest observed run was 1462s ("Running Copilot cloud agent"), but this cannot be decomposed into tool-level latencyevents.jsonl, none of the required Phase 2 checks (slow MCP calls, oversized responses, late validation failures, context size, model-switch patterns, prompt drift) could be executed with real evidenceProposed Change
Update the
shared/copilot-session-data-fetch.mdsetup component to also download/extract per-runevents.jsonl(or the raw conversation transcript) for each session in scope, using the GitHub Actions API/artifact download (e.g.,gh run downloadorgh api .../logs) before the copilot-opt analysis step runs. Add a fallback that fails loudly (non-silently) when logs are missing so data-quality gaps are visible in the audit output rather than silently degrading analysis.Expected Impact
Notes