Target Workflow: smoke-copilot-byok-aoai-entra
Source report: #7591
Estimated cost per run: $0.00 (BYOK — cost not billed to Copilot metering, but token/turn volume still drives duration & Actions minutes)
Total tokens per run: ~193K avg (102K on schedule, 236K–241K on pull_request)
Cache hit rate: Not exposed in per-run data, but token variance (102K vs ~240K) and aic swings (4.89 vs 14.15–14.57) between runs of an otherwise-static prompt strongly suggest low prefix-cache reuse
LLM turns: Not recorded in report/log data (turns=0 for all workflows in this dataset — likely not captured by the analyzer for this engine)
Current Configuration
| Setting |
Value |
| Tools loaded |
bash: ["*"] (unrestricted shell), github: {toolsets: [pull_requests]} |
| Tools actually used |
bash for cat only (per prompt); GitHub MCP: only github-list_pull_requests is invoked, and the prompt explicitly allows falling back to pre-fetched data without calling it at all |
| Network groups |
defaults, github, login.microsoftonline.com (all required — Entra OIDC↔Azure AD token exchange needs the last one) |
| Pre-agent steps |
Yes — Pre-compute BYOK smoke test data already fetches PRs, tests github.com connectivity, and does file I/O deterministically before the agent runs |
| Post-agent steps |
Yes — Validate safe outputs were invoked, Verify BYOK mode was active |
| Prompt size |
~4.4K chars (body) + ~5.5K chars (frontmatter) |
Recommendations
1. Drop the redundant live GitHub MCP verification call
Estimated savings: ~15–25K tokens/run (~10–20%), plus 1 fewer github_api_calls round trip
The prompt (step "1. GitHub MCP Testing") tells the agent to call github-list_pull_requests live to "verify MCP connectivity" even though the pre-agent steps: block already fetched the same 2 merged PRs deterministically via gh pr list into /tmp/gh-aw/agent/smoke-pr-data.txt. Since the fallback path ("if unavailable, validate the pre-fetched data instead") is explicitly allowed and produces the same test outcome, the live call adds a full MCP tool invocation + response payload (~5–10K tokens for the tool schema + result) for no additional signal — the pre-agent step already validates GitHub API reachability via gh pr list.
Change the prompt section to:
### 1. GitHub PR Data Verification
Read `/tmp/gh-aw/agent/smoke-context.txt` once. The pre-agent step already confirmed
GitHub API connectivity by fetching the 2 most recent merged PRs via `gh pr list`.
Mark ✅ if `recent_prs:` in the context file contains PR data, ❌ if it says "(PR fetch failed)".
Do not call any GitHub MCP tool for this check — the pre-fetched data is authoritative.
And remove tools: { github: {...} } entirely if no other step in the prompt needs a live GitHub MCP call (double-check the rest of the prompt doesn't rely on it — currently it doesn't). This also removes the GitHub MCP toolset schema (~8–10 tool definitions for pull_requests toolset) from every turn.
2. Restrict bash from "*" to the specific commands actually used
Estimated savings: ~1–3K tokens/run (schema size) + tightens the security surface
The prompt only ever needs cat (to verify the smoke-test file) plus whatever the pre-agent step already ran outside the agent. Change:
This is a smaller win than #1 but is free to combine, and reduces the effective tool-call action space the model has to reason over each turn.
3. Trim prompt verbosity in the body (static content resent every run)
Estimated savings: ~1–2K tokens/run
The ## Purpose and ### 4. BYOK Inference Test sections repeat the full auth-chain explanation (OIDC → Azure AD → Foundry → sidecar) in two places (frontmatter description: field and the body). Condense to one canonical explanation and have the body reference it briefly, e.g.:
## Purpose
Validates direct BYOK mode against Azure OpenAI (Foundry) via Microsoft Entra
(GitHub OIDC federated credential), routed through the api-proxy sidecar.
See workflow `description:` frontmatter for the full auth chain.
4. Reorder pre-computed context file to improve prefix-cache reuse
Estimated savings: Indirect — improves cache hit rate, reducing effective input-token cost across repeated 12h schedule runs
smoke-context.txt currently opens with event=, item_number=, http_code=, file_path= (which includes ${GITHUB_RUN_ID}, unique per run) before the largely-stable recent_prs: block. Since the agent reads this file with cat (its content becomes part of the transcript), placing the run-unique file_path/file_content lines after the more stable event/recent_prs data — or better, moving the run-ID-bearing line to the very end of the file — maximizes the stable prefix shared across the twice-daily schedule runs, improving cache hits when the model reasons over repeated content.
5. Confirm threat-detection: enabled: false stays disabled (already optimal)
No change needed — this is already turned off, avoiding an extra classification LLM turn. Documenting here so it isn't re-enabled by a future template sync.
Expected Impact
| Metric |
Current |
Projected |
Savings |
| Total tokens/run |
~193K avg |
~150–165K avg |
~15–20% |
| Cost/run |
$0.00 (BYOK, untracked) |
$0.00 |
n/a |
| GitHub API calls/run |
2–12 |
1–11 (drop 1 live MCP call on PR runs) |
-1 call |
| Session time |
5.2–7.1 min |
~4.5–6 min (est., fewer round trips) |
~10–15% |
Implementation Checklist
Generated by Daily Copilot Token Optimization Advisor · auto · 35 AIC · ⊞ 10.7K · ◷
Target Workflow:
smoke-copilot-byok-aoai-entraSource report: #7591
Estimated cost per run: $0.00 (BYOK — cost not billed to Copilot metering, but token/turn volume still drives duration & Actions minutes)
Total tokens per run: ~193K avg (102K on
schedule, 236K–241K onpull_request)Cache hit rate: Not exposed in per-run data, but token variance (102K vs ~240K) and
aicswings (4.89 vs 14.15–14.57) between runs of an otherwise-static prompt strongly suggest low prefix-cache reuseLLM turns: Not recorded in report/log data (turns=0 for all workflows in this dataset — likely not captured by the analyzer for this engine)
Current Configuration
bash: ["*"](unrestricted shell),github: {toolsets: [pull_requests]}bashforcatonly (per prompt); GitHub MCP: onlygithub-list_pull_requestsis invoked, and the prompt explicitly allows falling back to pre-fetched data without calling it at alldefaults,github,login.microsoftonline.com(all required — Entra OIDC↔Azure AD token exchange needs the last one)Pre-compute BYOK smoke test dataalready fetches PRs, tests github.com connectivity, and does file I/O deterministically before the agent runsValidate safe outputs were invoked,Verify BYOK mode was activeRecommendations
1. Drop the redundant live GitHub MCP verification call
Estimated savings: ~15–25K tokens/run (~10–20%), plus 1 fewer
github_api_callsround tripThe prompt (step "1. GitHub MCP Testing") tells the agent to call
github-list_pull_requestslive to "verify MCP connectivity" even though the pre-agentsteps:block already fetched the same 2 merged PRs deterministically viagh pr listinto/tmp/gh-aw/agent/smoke-pr-data.txt. Since the fallback path ("if unavailable, validate the pre-fetched data instead") is explicitly allowed and produces the same test outcome, the live call adds a full MCP tool invocation + response payload (~5–10K tokens for the tool schema + result) for no additional signal — the pre-agent step already validates GitHub API reachability viagh pr list.Change the prompt section to:
And remove
tools: { github: {...} }entirely if no other step in the prompt needs a live GitHub MCP call (double-check the rest of the prompt doesn't rely on it — currently it doesn't). This also removes the GitHub MCP toolset schema (~8–10 tool definitions forpull_requeststoolset) from every turn.2. Restrict
bashfrom"*"to the specific commands actually usedEstimated savings: ~1–3K tokens/run (schema size) + tightens the security surface
The prompt only ever needs
cat(to verify the smoke-test file) plus whatever the pre-agent step already ran outside the agent. Change:This is a smaller win than #1 but is free to combine, and reduces the effective tool-call action space the model has to reason over each turn.
3. Trim prompt verbosity in the body (static content resent every run)
Estimated savings: ~1–2K tokens/run
The
## Purposeand### 4. BYOK Inference Testsections repeat the full auth-chain explanation (OIDC → Azure AD → Foundry → sidecar) in two places (frontmatterdescription:field and the body). Condense to one canonical explanation and have the body reference it briefly, e.g.:4. Reorder pre-computed context file to improve prefix-cache reuse
Estimated savings: Indirect — improves cache hit rate, reducing effective input-token cost across repeated 12h
schedulerunssmoke-context.txtcurrently opens withevent=,item_number=,http_code=,file_path=(which includes${GITHUB_RUN_ID}, unique per run) before the largely-stablerecent_prs:block. Since the agent reads this file withcat(its content becomes part of the transcript), placing the run-uniquefile_path/file_contentlines after the more stableevent/recent_prsdata — or better, moving the run-ID-bearing line to the very end of the file — maximizes the stable prefix shared across the twice-dailyscheduleruns, improving cache hits when the model reasons over repeated content.5. Confirm
threat-detection: enabled: falsestays disabled (already optimal)No change needed — this is already turned off, avoiding an extra classification LLM turn. Documenting here so it isn't re-enabled by a future template sync.
Expected Impact
Implementation Checklist
github-list_pull_requestsverification step from the prompt body; rely solely on pre-fetchedsmoke-context.txtdatatools: { github: { toolsets: [pull_requests] } }from frontmatter if no other prompt step needs a live GitHub MCP callbash:tool list from["*"]to["cat *"]smoke-context.txtwrites in the pre-agent step so run-unique content (file_path,file_content) is appended lastgh aw compile .github/workflows/smoke-copilot-byok-aoai-entra.mdnpx tsx scripts/ci/postprocess-smoke-workflows.ts