Skip to content

⚡ Claude Token Optimization2026-09-17 — Smoke Claude #8705

Description

@github-actions

Target Workflow: Smoke Claude (.github/workflows/smoke-claude.md)

Source report: #8703
Estimated cost per run: n/a (estimated_cost not populated in this artifact)
Total tokens per run: ~5.0K (avg of 6 successful runs; 2 failed runs recorded 0)
Cache read rate: unavailable (token_usage_summary is null for all 9 runs — instrumentation gap)
Cache write rate: unavailable (same gap)
LLM turns: target is 1 turn per run (per the prompt's own comment); working_set.invocations shows 5–6 per successful run

Current Configuration

Setting Value
Tools loaded bash: [bash] only; github: false (already minimal — 1 tool)
Tools actually used bash (via harness), safeoutputs.add_comment, safeoutputs.add_labels, safeoutputs.noop
Network groups network.allowDomains — 33 infra/CA domains from the defaults preset (archive.ubuntu.com, ocsp., crl., packages.microsoft.com, etc.) — not user-added, comes from the shared AWF default allowlist
Pre-agent steps Yes — 5 steps already precompute the smoke test file, gh pr list data, GitHub.com reachability check, and a final final-result.json the agent just reads
Prompt size ~1.6K chars (short); includes a ~500-char HTML comment explaining a load-bearing ${{ github.run_id }} self-import workaround
max-turns 8 (target behavior is 1 turn)
Model claude-haiku-4-5 (already the cheapest Claude tier)

This workflow is already well-optimized at the structural level: it uses the cheapest model, restricts tools to a single bash capability with GitHub tools fully disabled (github: false), and does all deterministic work (file creation, gh pr list, reachability curl, JSON result assembly) in pre-agent steps: rather than inside the LLM loop. The remaining token cost (~5K/run) is largely fixed overhead (system prompt + MCP config + tool schema for the single safeoutputs MCP server), not agent inefficiency. Recommendations below focus on the actual waste patterns visible in the last 7 days of runs.

Recommendations

1. Fix the 25% driver_exit failure rate (highest-impact, non-token waste)

Estimated savings: ~9–10 Actions-minutes/week with zero useful output eliminated; indirectly protects against retry-driven token/API spend

2 of 8 runs in the report window failed with failure_kind: driver_exit and token_usage: 0, meaning the Claude harness crashed before recording any usage — pure waste (4–5 action-minutes each, no smoke-test signal):

  • Schedule-triggered failure: run 35169394899 — 1 error, 13 github_api_calls (vs. 2 for normal pull_request runs)
  • PR-triggered failure (attempt 2): run 35125325883 — 2 errors, 2 github_api_calls

Action: pull the raw driver/harness logs for these two runs (.github/aw/logs/run-35169394899, .github/aw/logs/run-35125325883) to find the crash root cause (likely MCP gateway startup timing or claude_harness.cjs exit before the prompt was consumed). Since one of the two failures is schedule-triggered and shows 6x more GitHub API calls than normal, check whether the gh pr list --state merged --limit 2 pre-step is retrying/paginating unexpectedly when there's no PR context — this is worth instrumenting with set -x in the "Pre-fetch GitHub API data" step.

2. Reduce max-turns from 8 to 2

Estimated savings: caps worst-case runaway cost; no change to the normal-path token count (~5K/run already achieved in 1 turn)

The prompt's own post-step explicitly targets 1 turn (echo "::notice::Smoke test completed in ${TURN_COUNT} turns (target: 1)"), but max-turns: 8 in the frontmatter allows up to 8x that ceiling if the agent misbehaves (e.g., malformed JSON, retries, or schema-probe calls to safeoutputs). Since the prompt already warns against probing (do NOT pipe JSON via stdin ... that sends empty arguments and the call is rejected as a schema probe, wasting a turn), tightening the ceiling to 2 turns is a low-risk safety net:

max-turns: 2

3. Restore cache visibility instrumentation

Estimated savings: not directly token savings, but currently blocks any Anthropic cache write/read cost analysis for this workflow

token_usage_summary is null for all 9 runs in the log artifact, so no cache read/write ratio can be computed. Given the prompt is short and largely static across runs (aside from the run ID and PR number), it is a strong candidate for prefix caching — but this can't be confirmed without the missing field. File a follow-up to check why the api-proxy sidecar isn't surfacing per-model cache metrics for the claude engine (this is called out as a "persistent instrumentation gap" in report #8703 across two consecutive periods).

4. Trim the load-bearing comment block (minor)

Estimated savings: ~100–150 tokens/run (small, since it's outside the prompt cache boundary if unchanged between runs, but adds parse overhead)

The prompt body opens with a ~500-character HTML comment justifying why ${{ github.run_id }} must remain in the prompt. Consider moving this rationale to a code comment in the workflow source's YAML frontmatter (which is stripped from the rendered prompt) instead of the markdown body, so it's preserved for maintainers without being sent to the model on every run.

Cache Analysis (Anthropic-Specific)

Cache read/write breakdown could not be computed — token_usage_summary was null for all 9 runs (see Recommendation #3). Token usage was flat and highly consistent across the 6 successful runs (4,917–5,079 tokens, σ ≈ 70 tokens), which is consistent with effective prefix caching already occurring, but this cannot be confirmed from the available fields.

Cache write amortization: Unknown — instrumentation gap.
Cache cost vs benefit: Unknown — instrumentation gap.

Expected Impact

Metric Current Projected Savings
Total tokens/run ~5.0K ~5.0K ~0% (already lean)
Failed-run waste 2/8 runs (25%), ~4–5 min each, 0 tokens 0/8 runs (target) -100% wasted Actions minutes
max-turns ceiling 8 2 -75% worst-case turn budget
Cache visibility none (null) populated enables future cache-cost audits

Implementation Checklist

  • Pull harness/driver logs for runs 35169394899 and 35125325883 and identify the driver_exit root cause
  • Add set -x (or equivalent tracing) to the "Pre-fetch GitHub API data" step to explain the 13 vs 2 github_api_calls discrepancy on schedule triggers
  • Change max-turns: 8 to max-turns: 2 in .github/workflows/smoke-claude.md
  • File a follow-up issue to investigate why token_usage_summary (cache metrics) is null for the Claude engine's api-proxy integration
  • Recompile: gh aw compile .github/workflows/smoke-claude.md
  • Post-process: npx tsx scripts/ci/postprocess-smoke-workflows.ts
  • Verify CI passes on PR
  • Compare token usage and failure rate on new runs vs. this baseline (~5.0K tokens/run, 25% failure rate)

Generated by Daily Claude Token Optimization Advisor · copilot · auto · 55.3 AIC · ⊞ 10.1K · ◷

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions