Skip to content

feat(usage): fast-tier pricing and cost breakdown (upstream G7) - #1001

Merged
rynfar merged 3 commits into
pylonfrom
upstream/2026-10-03-g7-usage
Oct 3, 2026
Merged

rynfar merged 3 commits into
pylonfrom
upstream/2026-10-03-g7-usage

Conversation

@rynfar

@rynfar rynfar commented Oct 3, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

Ports upstream group G7: Codex Fast/Ultrafast pricing, plus a cost breakdown by token type, speed and model. The change is adapted onto Pylon's diverged apps/server/src/usage. All Pylon readers (OpenCode Go/local, Antigravity, cliproxy), price overrides, settings and the v3 scan-cache migration are kept.

Sources

Upstream SHA Outcome How
56914128c1 fix(usage): Codex Fast and Ultrafast now cost what they bill Adapted Cherry-picked with -x. UsageRecord.fast becomes speed: standard/fast/ultrafast. The Codex thread_settings_applied.service_tier value is tracked. Rates read LiteLLM *_priority / *_ultrafast tiers alongside the Claude provider_specific_entry.fast multiple. The scan cache moves to v5 in its own file, usage-scan-cache-v5.json, and the legacy file is read once.
e8545b293b feat(usage): show cost by token type, speed, and model detail Adapted Cherry-picked with -x. Optional bucket fields are added. mergeUsage gains categoryCost, speedCost, ModelTotals.tokens and unpricedTokens. Web gets share bars, a ranked model table, a per-model dialog and a Set price entry point. Mobile gets a Cost section.

Pylon adaptations

  • Readers: Pylon's OpenCode and Antigravity readers record speed: "standard". The upstream cursorUsageReader.ts and the rateModel / Cursor pricing tests were dropped because Pylon has no Cursor usage reader. Cursor summary rows were not adopted for the same reason.
  • Scan cache migration: Pylon's v3 migration (rows without a speed column, with a persisted fr pending-rescan flag) is kept and extended:
    • v3 and v4 Codex entries keep their history.
    • Each such entry is marked needsFastRescan, so a live rollout is fully re-parsed before it is served warm again. Upstream used a size: -1 sentinel instead.
    • v4 rows load as-is, because speed index 0/1 matches the old fast flag.
    • Tests cover v3→v5 and v4→v5, and that the pending rescan survives a round trip.
  • Pricing source: Pylon's provider model-manifest.json carries no prices, so the Fast/Ultrafast premiums come from the LiteLLM rate table, as upstream does. Custom price overrides apply at every speed (unchanged).
  • Web page: Pylon's environment filter (approximate environments, source warnings) is kept. The Model prices dialog moves from the filter up to the page, so the model dialog's Set price can open it with that model pre-filled.
  • Mobile: the share-segment logic is extracted into usageCostMix.ts with unit tests. Upstream had it inline.
  • Docs: docs/user/usage.md keeps Pylon's provider list and wording. Only the new breakdown and Set-price sentences were added.
  • Compatibility: USAGE_CONTRACT_VERSION stays at 8. All new bucket fields are Schema.optional. Older clients ignore them, and the merge treats buckets from older servers as unsplit standard cost.

Review fixes (c1230f7ab9)

  • P1, model dialog showed whole-page totals: the dialog narrowed only the provider-wide summary.buckets list, but v7+ servers also send per-source buckets, and the merge reads those instead. New narrowUsageSummary (in packages/shared/usageMerge.ts) filters both. Regression tests use per-source buckets in usageMerge.test.ts and state/usage.test.tsx: models a ($1) and b ($5) filtered to a now give $1 and [a].
  • P3, duplicate Set-price row: whether to show the prefilled row is now decided on every render (withoutSupersededPrefill). An untouched prefill yields to an existing override once prices load. Typed rates are never dropped. Tested in usagePriceTable.test.ts.
  • P3, speed bar copy: the web page, the model dialog and mobile now show a footnote that servers predating speed tracking count all cost as Standard. The docs say the same.

Verification

  • vp test run src/usage (apps/server): 11 files, 143 tests passed.
  • vp test run src/usageMerge.test.ts (packages/shared): 46 passed.
  • vp test run src/state/usage.test.tsx src/components/usage (apps/web): 11 files, 89 passed.
  • vp test run src/features/usage (apps/mobile): 5 files, 12 passed.
  • Typechecks, each confirmed to run tsc --noEmit: @t3tools/contracts, @t3tools/shared, @t3tools/web, @t3tools/mobile and t3 all passed.
  • vp lint and vp fmt --check on the changed files: formatting is clean. The only lint output is 3 warnings that already exist on pylon (an unused layerTest in UsageService.ts, and unused useRef/useState in state/usage.ts).

Not verified

  • No browser or simulator run, so the visual layout of the share bars, model dialog and mobile Cost section is unchecked.
  • No real Codex ultrafast rollout or live LiteLLM document. Pricing math is covered only by synthetic rate tables.
  • An older Pylon server sharing the same state dir keeps writing the legacy usage-scan-cache.json. A v5 server reads that file only when its own file is missing, which is the intended upstream design.

Part of upstream cycle #996.

🤖 Generated with Claude Code

t3dotgg and others added 2 commits October 3, 2026 13:32
Records carry a billing speed (standard/fast/ultrafast) instead of a fast
flag. Codex rollouts now attribute the thread service tier, and rates read
LiteLLM priority/ultrafast tiers alongside the Claude fast multiple. The scan
cache moves to v5 in its own file; v3/v4 Codex history is kept and live
rollouts are fully re-parsed via Pylon's pending-rescan flag.

Pylon adaptations: OpenCode and Antigravity readers record standard speed;
no Cursor reader or rateModel (not in Pylon); v3 migration retained.

(cherry picked from commit 56914128c1ff9dcf6585c380b7b25845d6ca5f98)

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Usage buckets gain optional categoryCostUsd, fastCostUsd, ultrafastCostUsd
and speedPremiumUsd (contract version unchanged; older clients ignore them,
older servers merge as unsplit standard cost). The web page adds cost/token
share bars, a cost-by-speed bar with the premium, a ranked model table, and a
per-model dialog with Set price; mobile adds a Cost section.

Pylon adaptations: no Cursor rateModel or summary rows; Pylon's environment
filter (approximate environments, source warnings) is kept and the price
dialog lifts to the page; mobile segment logic is extracted to usageCostMix
with tests; docs keep Pylon's provider list.

(cherry picked from commit e8545b293b79cdd73d30ef921db149ef1b5abee0)

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@github-actions github-actions Bot added vouch:trusted PR author is trusted by repo permissions or the VOUCHED list. size:XXL labels Oct 3, 2026
@github-actions

github-actions Bot commented Oct 3, 2026 •

Copy link
Copy Markdown
Contributor

Thread transfer impact

✅ Thread transfer remains within every enforced ceiling.

Provider Metric Main baseline This PR Impact PR ceiling
Codex Total thread wire 4.9 KiB 4.9 KiB 0 B (0.0%) 6.8 KiB ✅
Codex Thread snapshot wire 3.7 KiB 3.7 KiB 0 B (0.0%) 4.9 KiB ✅
Codex Live turn WebSocket wire 1.2 KiB 1.2 KiB 0 B (0.0%) 2.0 KiB ✅
Codex Live turn WebSocket decoded 20.4 KiB 20.4 KiB 0 B (0.0%) 29.3 KiB ✅
Codex Live turn messages 2 2 0 (0.0%) 8 ✅
Claude Total thread wire 4.9 KiB 4.9 KiB 0 B (0.0%) 6.8 KiB ✅
Claude Thread snapshot wire 3.7 KiB 3.7 KiB 0 B (0.0%) 4.9 KiB ✅
Claude Live turn WebSocket wire 1.2 KiB 1.2 KiB 0 B (0.0%) 2.0 KiB ✅
Claude Live turn WebSocket decoded 20.8 KiB 20.8 KiB 0 B (0.0%) 29.3 KiB ✅
Claude Live turn messages 2 2 0 (0.0%) 8 ✅

Baseline: 30ec43a · PR result: c1230f7 · Source CI: success

Scenario and decoded snapshot size

10 historical turns, 5 command tools per turn, 878.9 KiB retained MCP result per historical turn, and a 1.05 MiB retained result in the measured turn.

  • Codex decoded thread snapshot: 106.1 KiB
  • Claude decoded thread snapshot: 106.4 KiB

Updated in place by a trusted workflow. PR artifacts are strictly validated and never executed.

The per-model dialog filtered only provider-wide buckets, but v7+ servers
also send per-source buckets that the merge prefers, so it showed whole-page
totals. narrowUsageSummary filters both. The Set price prefill now yields to
an existing override after prices load instead of deciding once at mount,
and the speed bar notes that older servers count all cost as Standard.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:XXL vouch:trusted PR author is trusted by repo permissions or the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants