fix(router): don't re-select the same failed deployment within a stage - #105
Merged
Conversation
Within a routing stage, selectDeploymentChoice now takes the just-failed deployment ID and excludes it from weighted selection when alternatives remain. Previously the retry loop could re-select the same dead endpoint up to stage.Retries times before the circuit breaker tripped and the stage advanced, burning latency and API quota. The Chat and StreamChat paths now track a stage-scoped recentlyFailed id and re-evaluate selection each attempt; a single-deployment stage still retries the one option to trip its breaker and move on. Adds TestDeploymentRouterRetriesPreferDifferentEndpoint.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Within a routing stage, the retry loop re-computed the weighted deployment
selection from the same candidate list each attempt, so the just-failed
endpoint could be re-selected up to stage.Retries times before the circuit
breaker tripped and the stage advanced. That burned latency and API quota on
known-dead deployments when healthy alternatives were available.
selectDeploymentChoice now takes the just-failed deployment ID and excludes
it from the weighted pool when alternatives remain (single-deployment stages
still retry the one option to detect failure / trip the breaker and advance,
preserving the half-open semantics existing tests rely on). Both Chat and
StreamChat track a stage-scoped recentlyFailed id and exclude it on each
attempt.
Adds TestDeploymentRouterRetriesPreferDifferentEndpoint: a two-deployment
stage (anthropic-direct fails 503, anthropic-vertex healthy, Retries:3) ->
direct is tried at most once, vertex is reached and returns success.
All router/..., retry/..., fallback/... tests pass; gofumpt + golangci-lint clean.