Target Workflow: Smoke Claude (.github/workflows/smoke-claude.md)
Source report: #8703
Estimated cost per run: n/a (estimated_cost not populated in this artifact)
Total tokens per run: ~5.0K (avg of 6 successful runs; 2 failed runs recorded 0)
Cache read rate: unavailable (token_usage_summary is null for all 9 runs — instrumentation gap)
Cache write rate: unavailable (same gap)
LLM turns: target is 1 turn per run (per the prompt's own comment); working_set.invocations shows 5–6 per successful run
Current Configuration
| Setting |
Value |
| Tools loaded |
bash: [bash] only; github: false (already minimal — 1 tool) |
| Tools actually used |
bash (via harness), safeoutputs.add_comment, safeoutputs.add_labels, safeoutputs.noop |
| Network groups |
network.allowDomains — 33 infra/CA domains from the defaults preset (archive.ubuntu.com, ocsp., crl., packages.microsoft.com, etc.) — not user-added, comes from the shared AWF default allowlist |
| Pre-agent steps |
Yes — 5 steps already precompute the smoke test file, gh pr list data, GitHub.com reachability check, and a final final-result.json the agent just reads |
| Prompt size |
~1.6K chars (short); includes a ~500-char HTML comment explaining a load-bearing ${{ github.run_id }} self-import workaround |
max-turns |
8 (target behavior is 1 turn) |
| Model |
claude-haiku-4-5 (already the cheapest Claude tier) |
This workflow is already well-optimized at the structural level: it uses the cheapest model, restricts tools to a single bash capability with GitHub tools fully disabled (github: false), and does all deterministic work (file creation, gh pr list, reachability curl, JSON result assembly) in pre-agent steps: rather than inside the LLM loop. The remaining token cost (~5K/run) is largely fixed overhead (system prompt + MCP config + tool schema for the single safeoutputs MCP server), not agent inefficiency. Recommendations below focus on the actual waste patterns visible in the last 7 days of runs.
Recommendations
1. Fix the 25% driver_exit failure rate (highest-impact, non-token waste)
Estimated savings: ~9–10 Actions-minutes/week with zero useful output eliminated; indirectly protects against retry-driven token/API spend
2 of 8 runs in the report window failed with failure_kind: driver_exit and token_usage: 0, meaning the Claude harness crashed before recording any usage — pure waste (4–5 action-minutes each, no smoke-test signal):
- Schedule-triggered failure: run 35169394899 — 1 error, 13
github_api_calls (vs. 2 for normal pull_request runs)
- PR-triggered failure (attempt 2): run 35125325883 — 2 errors, 2
github_api_calls
Action: pull the raw driver/harness logs for these two runs (.github/aw/logs/run-35169394899, .github/aw/logs/run-35125325883) to find the crash root cause (likely MCP gateway startup timing or claude_harness.cjs exit before the prompt was consumed). Since one of the two failures is schedule-triggered and shows 6x more GitHub API calls than normal, check whether the gh pr list --state merged --limit 2 pre-step is retrying/paginating unexpectedly when there's no PR context — this is worth instrumenting with set -x in the "Pre-fetch GitHub API data" step.
2. Reduce max-turns from 8 to 2
Estimated savings: caps worst-case runaway cost; no change to the normal-path token count (~5K/run already achieved in 1 turn)
The prompt's own post-step explicitly targets 1 turn (echo "::notice::Smoke test completed in ${TURN_COUNT} turns (target: 1)"), but max-turns: 8 in the frontmatter allows up to 8x that ceiling if the agent misbehaves (e.g., malformed JSON, retries, or schema-probe calls to safeoutputs). Since the prompt already warns against probing (do NOT pipe JSON via stdin ... that sends empty arguments and the call is rejected as a schema probe, wasting a turn), tightening the ceiling to 2 turns is a low-risk safety net:
3. Restore cache visibility instrumentation
Estimated savings: not directly token savings, but currently blocks any Anthropic cache write/read cost analysis for this workflow
token_usage_summary is null for all 9 runs in the log artifact, so no cache read/write ratio can be computed. Given the prompt is short and largely static across runs (aside from the run ID and PR number), it is a strong candidate for prefix caching — but this can't be confirmed without the missing field. File a follow-up to check why the api-proxy sidecar isn't surfacing per-model cache metrics for the claude engine (this is called out as a "persistent instrumentation gap" in report #8703 across two consecutive periods).
4. Trim the load-bearing comment block (minor)
Estimated savings: ~100–150 tokens/run (small, since it's outside the prompt cache boundary if unchanged between runs, but adds parse overhead)
The prompt body opens with a ~500-character HTML comment justifying why ${{ github.run_id }} must remain in the prompt. Consider moving this rationale to a code comment in the workflow source's YAML frontmatter (which is stripped from the rendered prompt) instead of the markdown body, so it's preserved for maintainers without being sent to the model on every run.
Cache Analysis (Anthropic-Specific)
Cache read/write breakdown could not be computed — token_usage_summary was null for all 9 runs (see Recommendation #3). Token usage was flat and highly consistent across the 6 successful runs (4,917–5,079 tokens, σ ≈ 70 tokens), which is consistent with effective prefix caching already occurring, but this cannot be confirmed from the available fields.
Cache write amortization: Unknown — instrumentation gap.
Cache cost vs benefit: Unknown — instrumentation gap.
Expected Impact
| Metric |
Current |
Projected |
Savings |
| Total tokens/run |
~5.0K |
~5.0K |
~0% (already lean) |
| Failed-run waste |
2/8 runs (25%), ~4–5 min each, 0 tokens |
0/8 runs (target) |
-100% wasted Actions minutes |
max-turns ceiling |
8 |
2 |
-75% worst-case turn budget |
| Cache visibility |
none (null) |
populated |
enables future cache-cost audits |
Implementation Checklist
Generated by Daily Claude Token Optimization Advisor · copilot · auto · 55.3 AIC · ⊞ 10.1K · ◷
Target Workflow:
Smoke Claude(.github/workflows/smoke-claude.md)Source report: #8703
Estimated cost per run: n/a (
estimated_costnot populated in this artifact)Total tokens per run: ~5.0K (avg of 6 successful runs; 2 failed runs recorded 0)
Cache read rate: unavailable (
token_usage_summaryisnullfor all 9 runs — instrumentation gap)Cache write rate: unavailable (same gap)
LLM turns: target is 1 turn per run (per the prompt's own comment);
working_set.invocationsshows 5–6 per successful runCurrent Configuration
bash: [bash]only;github: false(already minimal — 1 tool)bash(via harness),safeoutputs.add_comment,safeoutputs.add_labels,safeoutputs.noopnetwork.allowDomains— 33 infra/CA domains from thedefaultspreset (archive.ubuntu.com, ocsp., crl., packages.microsoft.com, etc.) — not user-added, comes from the shared AWF default allowlistgh pr listdata, GitHub.com reachability check, and a finalfinal-result.jsonthe agent just reads${{ github.run_id }}self-import workaroundmax-turnsclaude-haiku-4-5(already the cheapest Claude tier)This workflow is already well-optimized at the structural level: it uses the cheapest model, restricts tools to a single
bashcapability with GitHub tools fully disabled (github: false), and does all deterministic work (file creation,gh pr list, reachability curl, JSON result assembly) in pre-agentsteps:rather than inside the LLM loop. The remaining token cost (~5K/run) is largely fixed overhead (system prompt + MCP config + tool schema for the singlesafeoutputsMCP server), not agent inefficiency. Recommendations below focus on the actual waste patterns visible in the last 7 days of runs.Recommendations
1. Fix the 25%
driver_exitfailure rate (highest-impact, non-token waste)Estimated savings: ~9–10 Actions-minutes/week with zero useful output eliminated; indirectly protects against retry-driven token/API spend
2 of 8 runs in the report window failed with
failure_kind: driver_exitandtoken_usage: 0, meaning the Claude harness crashed before recording any usage — pure waste (4–5 action-minutes each, no smoke-test signal):github_api_calls(vs. 2 for normal pull_request runs)github_api_callsAction: pull the raw driver/harness logs for these two runs (
.github/aw/logs/run-35169394899,.github/aw/logs/run-35125325883) to find the crash root cause (likely MCP gateway startup timing orclaude_harness.cjsexit before the prompt was consumed). Since one of the two failures is schedule-triggered and shows 6x more GitHub API calls than normal, check whether thegh pr list --state merged --limit 2pre-step is retrying/paginating unexpectedly when there's no PR context — this is worth instrumenting withset -xin the "Pre-fetch GitHub API data" step.2. Reduce
max-turnsfrom 8 to 2Estimated savings: caps worst-case runaway cost; no change to the normal-path token count (~5K/run already achieved in 1 turn)
The prompt's own post-step explicitly targets 1 turn (
echo "::notice::Smoke test completed in ${TURN_COUNT} turns (target: 1)"), butmax-turns: 8in the frontmatter allows up to 8x that ceiling if the agent misbehaves (e.g., malformed JSON, retries, or schema-probe calls tosafeoutputs). Since the prompt already warns against probing (do NOT pipe JSON via stdin ... that sends empty arguments and the call is rejected as a schema probe, wasting a turn), tightening the ceiling to 2 turns is a low-risk safety net:3. Restore cache visibility instrumentation
Estimated savings: not directly token savings, but currently blocks any Anthropic cache write/read cost analysis for this workflow
token_usage_summaryisnullfor all 9 runs in the log artifact, so no cache read/write ratio can be computed. Given the prompt is short and largely static across runs (aside from the run ID and PR number), it is a strong candidate for prefix caching — but this can't be confirmed without the missing field. File a follow-up to check why the api-proxy sidecar isn't surfacing per-model cache metrics for theclaudeengine (this is called out as a "persistent instrumentation gap" in report #8703 across two consecutive periods).4. Trim the load-bearing comment block (minor)
Estimated savings: ~100–150 tokens/run (small, since it's outside the prompt cache boundary if unchanged between runs, but adds parse overhead)
The prompt body opens with a ~500-character HTML comment justifying why
${{ github.run_id }}must remain in the prompt. Consider moving this rationale to a code comment in the workflow source's YAML frontmatter (which is stripped from the rendered prompt) instead of the markdown body, so it's preserved for maintainers without being sent to the model on every run.Cache Analysis (Anthropic-Specific)
Cache read/write breakdown could not be computed —
token_usage_summarywasnullfor all 9 runs (see Recommendation #3). Token usage was flat and highly consistent across the 6 successful runs (4,917–5,079 tokens, σ ≈ 70 tokens), which is consistent with effective prefix caching already occurring, but this cannot be confirmed from the available fields.Cache write amortization: Unknown — instrumentation gap.
Cache cost vs benefit: Unknown — instrumentation gap.
Expected Impact
max-turnsceilingnull)Implementation Checklist
driver_exitroot causeset -x(or equivalent tracing) to the "Pre-fetch GitHub API data" step to explain the 13 vs 2github_api_callsdiscrepancy on schedule triggersmax-turns: 8tomax-turns: 2in.github/workflows/smoke-claude.mdtoken_usage_summary(cache metrics) isnullfor the Claude engine's api-proxy integrationgh aw compile .github/workflows/smoke-claude.mdnpx tsx scripts/ci/postprocess-smoke-workflows.ts