feat(proxy): POST /v1/chat/completions + bearer auth + Hub dispatch - #10
Conversation
Wires the proxy surface end-to-end: an authenticated request arrives,
gets its API key looked up in the snapshot, its Model resolved, its
Provider dispatched through the Hub to the registered Bridge, and the
response re-rendered into OpenAI's wire shape (or an SSE stream of
ChatCompletionChunks).
Proxy crate layout (each file <200 LOC):
- error.rs: ProxyError → OpenAI-style {error:{message,type}} envelope
with stable status/type tokens. Bridge errors inherit status via
BridgeError::http_status() so 4xx passes through and 5xx collapses
to 502.
- auth.rs: AuthenticatedKey FromRequestParts extractor reading
Authorization: Bearer <key> (with x-api-key fallback), looking the
key up in snapshot.apikeys.
- state.rs: ProxyState now holds Arc<Hub> alongside the SnapshotHandle.
- render.rs: own render types for chat.completion / .chunk — kept
distinct from the provider wire types so a future client-facing
schema change doesn't ripple into every upstream adapter.
- chat.rs: request-flow handler. Empty messages -> 400, unknown model
-> 404, forbidden model -> 403, unregistered provider -> 503.
Streaming runs through axum's Sse + 15s keep-alive and terminates
with [DONE] after the upstream closes.
Server bootstrap now constructs a Hub registering all four bridges
(OpenAi, Anthropic, Gemini, DeepSeek) and hands it to ProxyState
alongside the existing snapshot handle.
22 new proxy tests (15 unit + 7 integration via axum oneshot + wiremock):
happy path (200 OpenAI-shaped JSON), missing auth (401), unknown key
(401), forbidden model (403), unknown model (404), empty messages (400),
upstream 429 pass-through, no registered bridge (503), streaming SSE
with [DONE] sentinel. 164 tests pass workspace-wide.
There was a problem hiding this comment.
Pull request overview
This PR wires up the first end-to-end proxy request path for POST /v1/chat/completions, including bearer-key authentication, model resolution/authorization via snapshot, Hub-based provider dispatch, and OpenAI-shaped JSON/SSE responses.
Changes:
- Add
aisix-proxychat-completions handler with auth extractor, Hub dispatch, OpenAI-shaped rendering, and OpenAI-style error envelopes. - Extend
aisix-serverbootstrap to build/register aHubwith all provider bridges and pass it into proxy state. - Add unit + wiremock integration tests for non-streaming + streaming behavior and common error cases.
Reviewed changes
Copilot reviewed 9 out of 10 changed files in this pull request and generated 5 comments.
Show a summary per file
| File | Description |
|---|---|
| crates/aisix-server/src/main.rs | Builds a Hub at startup, registers provider bridges, and injects into ProxyState. |
| crates/aisix-server/Cargo.toml | Adds provider bridge crates as server dependencies for Hub registration. |
| crates/aisix-proxy/src/state.rs | Introduces ProxyState holding snapshot + Arc<Hub> + body limit config. |
| crates/aisix-proxy/src/render.rs | Renders normalized gateway chat types into OpenAI response / chunk shapes. |
| crates/aisix-proxy/src/lib.rs | Mounts /v1/chat/completions, updates /health, and adds extensive endpoint tests. |
| crates/aisix-proxy/src/error.rs | Defines ProxyError → OpenAI-style {error:{message,type}} envelope + status mapping. |
| crates/aisix-proxy/src/chat.rs | Implements request flow: validate, resolve model, authz, Hub dispatch, JSON/SSE responses. |
| crates/aisix-proxy/src/auth.rs | Adds AuthenticatedKey extractor parsing Bearer (or x-api-key) and snapshot lookup. |
| crates/aisix-proxy/Cargo.toml | Adds async-stream and dev deps for OpenAI bridge + wiremock tests. |
| Cargo.lock | Locks new deps introduced by proxy/server changes. |
💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.
| /// Build the proxy router. Mounts `/health` plus the | ||
| /// OpenAI-compatible chat-completions surface. | ||
| pub fn build_router(state: ProxyState) -> Router { | ||
| Router::new() | ||
| .route("/health", get(health)) | ||
| .route("/v1/chat/completions", post(chat::chat_completions)) | ||
| .with_state(state) |
There was a problem hiding this comment.
ProxyState::request_body_limit_bytes is configured and carried in state, but build_router doesn't apply any body-size limiting layer/extractor configuration, so requests can exceed the intended limit (or just use axum's default). Consider wiring state.request_body_limit_bytes into the router via DefaultBodyLimit::max(...) / RequestBodyLimitLayer (tower-http) so the config is actually enforced.
| while let Some(item) = upstream.next().await { | ||
| let ev = match item { | ||
| Ok(chunk) => { | ||
| let rendered = render_chunk(created, chunk); | ||
| match serde_json::to_string(&rendered) { | ||
| Ok(json) => Event::default().data(json), | ||
| Err(err) => Event::default() | ||
| .event("error") | ||
| .data(err.to_string()), | ||
| } | ||
| } | ||
| Err(err) => Event::default() | ||
| .event("error") | ||
| .data(err.to_string()), | ||
| }; | ||
| yield Ok::<_, Infallible>(ev); | ||
| } | ||
| // Emit the OpenAI-style [DONE] sentinel so clients that terminate | ||
| // on it behave correctly. | ||
| yield Ok::<_, Infallible>(Event::default().data("[DONE]")); | ||
| } |
There was a problem hiding this comment.
build_sse_stream emits the [DONE] sentinel unconditionally after draining upstream, even if the upstream stream yielded an Err. That can cause clients to treat an errored stream as a successful completion. Consider terminating the SSE stream immediately on the first upstream error (and ideally emitting an OpenAI-shaped {error:{...}} payload before closing) and only sending [DONE] on a clean upstream completion.
| if let Some(rest) = s.strip_prefix("Bearer ") { | ||
| let rest = rest.trim(); | ||
| if rest.is_empty() { | ||
| return Err(ProxyError::MissingAuth); | ||
| } | ||
| return Ok(rest.to_string()); |
There was a problem hiding this comment.
extract_bearer only accepts an Authorization scheme that exactly matches the case-sensitive prefix "Bearer ". Per HTTP auth scheme rules (RFC 9110), the scheme token is case-insensitive, and some clients send bearer. Consider parsing the scheme case-insensitively (e.g., split_once(' ') + eq_ignore_ascii_case("bearer")).
| if let Some(rest) = s.strip_prefix("Bearer ") { | |
| let rest = rest.trim(); | |
| if rest.is_empty() { | |
| return Err(ProxyError::MissingAuth); | |
| } | |
| return Ok(rest.to_string()); | |
| if let Some((scheme, rest)) = s.split_once(' ') { | |
| if scheme.eq_ignore_ascii_case("bearer") { | |
| let rest = rest.trim(); | |
| if rest.is_empty() { | |
| return Err(ProxyError::MissingAuth); | |
| } | |
| return Ok(rest.to_string()); | |
| } |
| //! OpenAI-compatible error envelope used by every proxy endpoint. | ||
| //! | ||
| //! OpenAI's clients expect this exact shape (spec §3): | ||
| //! | ||
| //! ```json | ||
| //! { | ||
| //! "error": { | ||
| //! "message": "…", | ||
| //! "type": "invalid_request_error", | ||
| //! "param": null, | ||
| //! "code": null | ||
| //! } | ||
| //! } | ||
| //! ``` | ||
| //! | ||
| //! `ProxyError` is the internal error taxonomy; it implements | ||
| //! `IntoResponse` so handlers can `?`-propagate without touching | ||
| //! JSON shape boilerplate. | ||
|
|
||
| use aisix_gateway::BridgeError; | ||
| use axum::http::StatusCode; | ||
| use axum::response::{IntoResponse, Response}; | ||
| use axum::Json; | ||
| use serde::Serialize; | ||
|
|
||
| #[derive(Debug, Serialize, Clone)] | ||
| pub struct ErrorEnvelope { | ||
| pub error: ErrorBody, | ||
| } | ||
|
|
||
| #[derive(Debug, Serialize, Clone)] | ||
| pub struct ErrorBody { | ||
| pub message: String, | ||
| #[serde(rename = "type")] | ||
| pub kind: &'static str, | ||
| #[serde(skip_serializing_if = "Option::is_none")] | ||
| pub param: Option<String>, | ||
| #[serde(skip_serializing_if = "Option::is_none")] | ||
| pub code: Option<String>, | ||
| } |
There was a problem hiding this comment.
The module-level docs show OpenAI errors with param: null / code: null and say clients expect that exact shape, but the implementation uses skip_serializing_if = Option::is_none which omits these fields entirely (and tests assert omission). Please either update the docs to match the actual wire format you're emitting, or change serialization to emit explicit null fields if that is the intended compatibility target.
| if req.is_streaming() { | ||
| let upstream = bridge.chat_stream(&req, &ctx).await?; | ||
| let model_name = req.model.clone(); | ||
| let sse_stream = build_sse_stream(upstream, model_name, now); | ||
| let response = | ||
| Sse::new(sse_stream).keep_alive(KeepAlive::new().interval(Duration::from_secs(15))); | ||
| return Ok(response.into_response()); | ||
| } | ||
|
|
||
| let upstream = bridge.chat(&req, &ctx).await?; | ||
| let rendered = render_response(now, upstream); | ||
| Ok(Json(rendered).into_response()) | ||
| } | ||
|
|
||
| fn created_ts() -> i64 { | ||
| SystemTime::now() | ||
| .duration_since(UNIX_EPOCH) | ||
| .map(|d| d.as_secs() as i64) | ||
| .unwrap_or(0) | ||
| } | ||
|
|
||
| fn build_sse_stream( | ||
| upstream: aisix_gateway::ChatChunkStream, | ||
| _model: String, | ||
| created: i64, | ||
| ) -> impl Stream<Item = Result<Event, Infallible>> { |
There was a problem hiding this comment.
chat_completions clones req.model into model_name and passes it into build_sse_stream, but build_sse_stream takes _model and never uses it. This is currently dead code / wasted allocation. Either remove the parameter + clone, or use it to force the rendered model field to the caller-facing model alias (instead of whatever the upstream returns) if that’s the intended behavior.
Summary
First end-to-end request path: client → auth → model resolve →
allowed-models check → Hub dispatch → upstream Bridge → OpenAI-shaped
response (or SSE stream of `ChatCompletionChunk`s).
Proxy crate layout (each file <200 LOC):
envelope. Bridge errors inherit status via `BridgeError::http_status()`
so 4xx passes through and 5xx collapses to 502.
`Authorization: Bearer ` (or `x-api-key` fallback), looks up
`snapshot.apikeys`.
`SnapshotHandle`.
kept separate from provider wire types so client schema changes don't
ripple into every adapter.
model → 404, forbidden model → 403, unregistered provider → 503.
Streaming runs through axum's `Sse` + 15s keep-alive and terminates
with `[DONE]`.
Server bootstrap now builds a Hub registering all four bridges (OpenAI,
Anthropic, Gemini, DeepSeek) and hands it to `ProxyState`.
Test plan