Target Workflow: smoke-copilot-byok-aoai-apikey
Source report: #7408
Estimated cost per run: $0.00 (BYOK/Foundry billing not tracked by this metric, but token volume itself drives inference latency/cost on the Foundry side)
Total tokens per run: ~84K avg (76K, 155K, 105K across 3 successful runs; one driver_exit failure with no token data)
Cache hit rate: N/A (not reported per-run in the aggregated JSON)
LLM turns: N/A (report shows 0 — turns field not populated for this workflow's runs)
Current Configuration
| Setting |
Value |
| Tools loaded |
bash: ["*"], github: {toolsets: [pull_requests]} — already minimal |
| Tools actually used |
GitHub MCP list_pull_requests (per prompt), bash for file write/read + curl check |
| Network groups |
defaults, github — both used (github.com connectivity test + GitHub MCP) |
| Pre-agent steps |
Yes — activation.pre-steps pre-computes PR data, HTTP check, file I/O before the agent runs |
| Prompt size |
~3.4K chars body (short, single-page smoke-test prompt) |
| Model |
o4-mini-aw |
Recommendations
1. Switch model from o4-mini-aw to claude-haiku-4.5 (match sibling workflow)
Estimated savings: ~60-75K tokens/run (~75-90%)
The sibling workflow smoke-copilot-byok (.github/workflows/smoke-copilot-byok.md) runs the same prompt structure, same pre-steps pattern, same output contract but uses COPILOT_MODEL: claude-haiku-4.5 and averages only ~5K tokens/run across its runs (4.8K, 5.0K, 5.1K), vs this workflow's ~84K tokens/run using o4-mini-aw for an equivalent, deterministic, single-page smoke test. This is by far the largest lever available — the task itself (verify 4 pre-computed facts and post/noop a short comment) does not require a heavier reasoning model.
# .github/workflows/smoke-copilot-byok-aoai-apikey.md
env:
COPILOT_MODEL: claude-haiku-4.5 # was: o4-mini-aw
If o4-mini-aw was intentionally chosen to validate a specific model routed through the Azure OpenAI/Foundry BYOK path (rather than to test Copilot's own model selection), keep the model but confirm this deliberately — in that case, skip this recommendation and instead see #2 for reducing turns.
2. Investigate elevated turn count on o4-mini-aw for this simple task
Estimated savings: ~10-20K tokens/run (~15-25%), compounding with #1
Token totals climbing from 76K → 105K → 155K across recent runs (with one driver_exit failure) for a task that should be near-constant cost suggests the agent is re-reading/re-verifying more than needed, or retrying tool calls. Since the prompt already instructs "Keep all outputs extremely short and concise," add an explicit turn/step budget hint and confirm the MCP GitHub call for list_pull_requests isn't being retried on transient errors:
**IMPORTANT: Complete this task in a single tool-call pass — call `github-list_pull_requests` at most once, then produce the summary directly. Do not retry on partial data; fall back to the Pre-Fetched PR Data instead.**
3. Trim prompt duplication between pre-fetched data and inline instructions
Estimated savings: ~1-2K tokens/run (~2%)
The prompt repeats the BYOK explanation in both the workflow description: frontmatter and the markdown body's "Purpose" section (near-duplicate paragraphs). Since neither is user-facing (only injected into the agent context), collapse to a single one-paragraph purpose statement and drop the repeated mode-comparison sentence already covered by code comments in the workflow file.
Expected Impact
| Metric |
Current |
Projected |
Savings |
| Total tokens/run |
~84K |
~10-15K |
-85% |
| Cost/run |
$0.00 (untracked) |
$0.00 (untracked) |
n/a — Foundry-side cost still reduced via lower token volume |
| LLM turns |
unknown (elevated vs sibling) |
reduced to match sibling (~1-2 turns) |
-~50%+ |
| Session time |
4.8-6.3 min |
~2-3 min (est., matching sibling's faster runs) |
-~40% |
Implementation Checklist
Generated by Daily Copilot Token Optimization Advisor · auto · 36.1 AIC · ⊞ 10.6K · ◷
Target Workflow:
smoke-copilot-byok-aoai-apikeySource report: #7408
Estimated cost per run: $0.00 (BYOK/Foundry billing not tracked by this metric, but token volume itself drives inference latency/cost on the Foundry side)
Total tokens per run: ~84K avg (76K, 155K, 105K across 3 successful runs; one
driver_exitfailure with no token data)Cache hit rate: N/A (not reported per-run in the aggregated JSON)
LLM turns: N/A (report shows 0 — turns field not populated for this workflow's runs)
Current Configuration
bash: ["*"],github: {toolsets: [pull_requests]}— already minimallist_pull_requests(per prompt), bash for file write/read + curl checkdefaults,github— both used (github.com connectivity test + GitHub MCP)activation.pre-stepspre-computes PR data, HTTP check, file I/O before the agent runso4-mini-awRecommendations
1. Switch model from
o4-mini-awtoclaude-haiku-4.5(match sibling workflow)Estimated savings: ~60-75K tokens/run (~75-90%)
The sibling workflow
smoke-copilot-byok(.github/workflows/smoke-copilot-byok.md) runs the same prompt structure, same pre-steps pattern, same output contract but usesCOPILOT_MODEL: claude-haiku-4.5and averages only ~5K tokens/run across its runs (4.8K, 5.0K, 5.1K), vs this workflow's ~84K tokens/run usingo4-mini-awfor an equivalent, deterministic, single-page smoke test. This is by far the largest lever available — the task itself (verify 4 pre-computed facts and post/noop a short comment) does not require a heavier reasoning model.If
o4-mini-awwas intentionally chosen to validate a specific model routed through the Azure OpenAI/Foundry BYOK path (rather than to test Copilot's own model selection), keep the model but confirm this deliberately — in that case, skip this recommendation and instead see #2 for reducing turns.2. Investigate elevated turn count on
o4-mini-awfor this simple taskEstimated savings: ~10-20K tokens/run (~15-25%), compounding with #1
Token totals climbing from 76K → 105K → 155K across recent runs (with one
driver_exitfailure) for a task that should be near-constant cost suggests the agent is re-reading/re-verifying more than needed, or retrying tool calls. Since the prompt already instructs "Keep all outputs extremely short and concise," add an explicit turn/step budget hint and confirm the MCP GitHub call forlist_pull_requestsisn't being retried on transient errors:3. Trim prompt duplication between pre-fetched data and inline instructions
Estimated savings: ~1-2K tokens/run (~2%)
The prompt repeats the BYOK explanation in both the workflow
description:frontmatter and the markdown body's "Purpose" section (near-duplicate paragraphs). Since neither is user-facing (only injected into the agent context), collapse to a single one-paragraph purpose statement and drop the repeated mode-comparison sentence already covered by code comments in the workflow file.Expected Impact
Implementation Checklist
COPILOT_MODELfromo4-mini-awtoclaude-haiku-4.5in.github/workflows/smoke-copilot-byok-aoai-apikey.md(or confirm intentional model pinning and skip)description:and markdown bodygh aw compile .github/workflows/smoke-copilot-byok-aoai-apikey.mdnpx tsx scripts/ci/postprocess-smoke-workflows.ts