Skip to content

⚡ Copilot Token Optimization2026-08-21 — smoke-copilot-byok-aoai-entra #7592

Description

@github-actions

Target Workflow: smoke-copilot-byok-aoai-entra

Source report: #7591
Estimated cost per run: $0.00 (BYOK — cost not billed to Copilot metering, but token/turn volume still drives duration & Actions minutes)
Total tokens per run: ~193K avg (102K on schedule, 236K–241K on pull_request)
Cache hit rate: Not exposed in per-run data, but token variance (102K vs ~240K) and aic swings (4.89 vs 14.15–14.57) between runs of an otherwise-static prompt strongly suggest low prefix-cache reuse
LLM turns: Not recorded in report/log data (turns=0 for all workflows in this dataset — likely not captured by the analyzer for this engine)

Current Configuration

Setting Value
Tools loaded bash: ["*"] (unrestricted shell), github: {toolsets: [pull_requests]}
Tools actually used bash for cat only (per prompt); GitHub MCP: only github-list_pull_requests is invoked, and the prompt explicitly allows falling back to pre-fetched data without calling it at all
Network groups defaults, github, login.microsoftonline.com (all required — Entra OIDC↔Azure AD token exchange needs the last one)
Pre-agent steps Yes — Pre-compute BYOK smoke test data already fetches PRs, tests github.com connectivity, and does file I/O deterministically before the agent runs
Post-agent steps Yes — Validate safe outputs were invoked, Verify BYOK mode was active
Prompt size ~4.4K chars (body) + ~5.5K chars (frontmatter)

Recommendations

1. Drop the redundant live GitHub MCP verification call

Estimated savings: ~15–25K tokens/run (~10–20%), plus 1 fewer github_api_calls round trip

The prompt (step "1. GitHub MCP Testing") tells the agent to call github-list_pull_requests live to "verify MCP connectivity" even though the pre-agent steps: block already fetched the same 2 merged PRs deterministically via gh pr list into /tmp/gh-aw/agent/smoke-pr-data.txt. Since the fallback path ("if unavailable, validate the pre-fetched data instead") is explicitly allowed and produces the same test outcome, the live call adds a full MCP tool invocation + response payload (~5–10K tokens for the tool schema + result) for no additional signal — the pre-agent step already validates GitHub API reachability via gh pr list.

Change the prompt section to:

### 1. GitHub PR Data Verification
Read `/tmp/gh-aw/agent/smoke-context.txt` once. The pre-agent step already confirmed
GitHub API connectivity by fetching the 2 most recent merged PRs via `gh pr list`.
Mark ✅ if `recent_prs:` in the context file contains PR data, ❌ if it says "(PR fetch failed)".
Do not call any GitHub MCP tool for this check — the pre-fetched data is authoritative.

And remove tools: { github: {...} } entirely if no other step in the prompt needs a live GitHub MCP call (double-check the rest of the prompt doesn't rely on it — currently it doesn't). This also removes the GitHub MCP toolset schema (~8–10 tool definitions for pull_requests toolset) from every turn.

2. Restrict bash from "*" to the specific commands actually used

Estimated savings: ~1–3K tokens/run (schema size) + tightens the security surface

The prompt only ever needs cat (to verify the smoke-test file) plus whatever the pre-agent step already ran outside the agent. Change:

tools:
  bash:
    - "cat *"

This is a smaller win than #1 but is free to combine, and reduces the effective tool-call action space the model has to reason over each turn.

3. Trim prompt verbosity in the body (static content resent every run)

Estimated savings: ~1–2K tokens/run

The ## Purpose and ### 4. BYOK Inference Test sections repeat the full auth-chain explanation (OIDC → Azure AD → Foundry → sidecar) in two places (frontmatter description: field and the body). Condense to one canonical explanation and have the body reference it briefly, e.g.:

## Purpose
Validates direct BYOK mode against Azure OpenAI (Foundry) via Microsoft Entra
(GitHub OIDC federated credential), routed through the api-proxy sidecar.
See workflow `description:` frontmatter for the full auth chain.

4. Reorder pre-computed context file to improve prefix-cache reuse

Estimated savings: Indirect — improves cache hit rate, reducing effective input-token cost across repeated 12h schedule runs

smoke-context.txt currently opens with event=, item_number=, http_code=, file_path= (which includes ${GITHUB_RUN_ID}, unique per run) before the largely-stable recent_prs: block. Since the agent reads this file with cat (its content becomes part of the transcript), placing the run-unique file_path/file_content lines after the more stable event/recent_prs data — or better, moving the run-ID-bearing line to the very end of the file — maximizes the stable prefix shared across the twice-daily schedule runs, improving cache hits when the model reasons over repeated content.

5. Confirm threat-detection: enabled: false stays disabled (already optimal)

No change needed — this is already turned off, avoiding an extra classification LLM turn. Documenting here so it isn't re-enabled by a future template sync.

Expected Impact

Metric Current Projected Savings
Total tokens/run ~193K avg ~150–165K avg ~15–20%
Cost/run $0.00 (BYOK, untracked) $0.00 n/a
GitHub API calls/run 2–12 1–11 (drop 1 live MCP call on PR runs) -1 call
Session time 5.2–7.1 min ~4.5–6 min (est., fewer round trips) ~10–15%

Implementation Checklist

  • Remove the live github-list_pull_requests verification step from the prompt body; rely solely on pre-fetched smoke-context.txt data
  • Remove tools: { github: { toolsets: [pull_requests] } } from frontmatter if no other prompt step needs a live GitHub MCP call
  • Restrict bash: tool list from ["*"] to ["cat *"]
  • Condense duplicated BYOK auth-chain explanation in the prompt body
  • Reorder smoke-context.txt writes in the pre-agent step so run-unique content (file_path, file_content) is appended last
  • Recompile: gh aw compile .github/workflows/smoke-copilot-byok-aoai-entra.md
  • Post-process: npx tsx scripts/ci/postprocess-smoke-workflows.ts
  • Verify CI passes on PR
  • Compare token usage on new run vs baseline (~193K avg) via the next token usage report

Generated by Daily Copilot Token Optimization Advisor · auto · 35 AIC · ⊞ 10.7K · ◷

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions