Skip to content

feat(proxy): POST /v1/chat/completions + bearer auth + Hub dispatch - #10

Merged
moonming merged 1 commit into
mainfrom
feat/proxy-chat-route
Apr 17, 2026
Merged

feat(proxy): POST /v1/chat/completions + bearer auth + Hub dispatch#10
moonming merged 1 commit into
mainfrom
feat/proxy-chat-route

Conversation

@moonming

Copy link
Copy Markdown
Collaborator

Summary

First end-to-end request path: client → auth → model resolve →
allowed-models check → Hub dispatch → upstream Bridge → OpenAI-shaped
response (or SSE stream of `ChatCompletionChunk`s).

Proxy crate layout (each file <200 LOC):

  • error.rs — `ProxyError` → OpenAI-style `{error:{message,type}}`
    envelope. Bridge errors inherit status via `BridgeError::http_status()`
    so 4xx passes through and 5xx collapses to 502.
  • auth.rs — `AuthenticatedKey` `FromRequestParts` extractor,
    `Authorization: Bearer ` (or `x-api-key` fallback), looks up
    `snapshot.apikeys`.
  • state.rs — `ProxyState` now holds `Arc` alongside the
    `SnapshotHandle`.
  • render.rs — own render types for `chat.completion` / `.chunk`,
    kept separate from provider wire types so client schema changes don't
    ripple into every adapter.
  • chat.rs — request-flow handler. Empty messages → 400, unknown
    model → 404, forbidden model → 403, unregistered provider → 503.
    Streaming runs through axum's `Sse` + 15s keep-alive and terminates
    with `[DONE]`.

Server bootstrap now builds a Hub registering all four bridges (OpenAI,
Anthropic, Gemini, DeepSeek) and hands it to `ProxyState`.

Test plan

  • 22 new proxy tests (15 unit + 7 wiremock-backed integration)
    • happy path → OpenAI-shaped JSON
    • missing auth → 401
    • unknown key → 401
    • forbidden model → 403
    • unknown model → 404
    • empty messages → 400
    • upstream 429 pass-through
    • provider not registered → 503
    • streaming SSE with `[DONE]` sentinel
  • `cargo test --workspace` — 164 tests pass
  • `cargo clippy --all-targets -- -D warnings` clean
  • `cargo fmt --check` clean
  • CI green across all 6 jobs

Wires the proxy surface end-to-end: an authenticated request arrives,
gets its API key looked up in the snapshot, its Model resolved, its
Provider dispatched through the Hub to the registered Bridge, and the
response re-rendered into OpenAI's wire shape (or an SSE stream of
ChatCompletionChunks).

Proxy crate layout (each file <200 LOC):
- error.rs: ProxyError → OpenAI-style {error:{message,type}} envelope
  with stable status/type tokens. Bridge errors inherit status via
  BridgeError::http_status() so 4xx passes through and 5xx collapses
  to 502.
- auth.rs: AuthenticatedKey FromRequestParts extractor reading
  Authorization: Bearer <key> (with x-api-key fallback), looking the
  key up in snapshot.apikeys.
- state.rs: ProxyState now holds Arc<Hub> alongside the SnapshotHandle.
- render.rs: own render types for chat.completion / .chunk — kept
  distinct from the provider wire types so a future client-facing
  schema change doesn't ripple into every upstream adapter.
- chat.rs: request-flow handler. Empty messages -> 400, unknown model
  -> 404, forbidden model -> 403, unregistered provider -> 503.
  Streaming runs through axum's Sse + 15s keep-alive and terminates
  with [DONE] after the upstream closes.

Server bootstrap now constructs a Hub registering all four bridges
(OpenAi, Anthropic, Gemini, DeepSeek) and hands it to ProxyState
alongside the existing snapshot handle.

22 new proxy tests (15 unit + 7 integration via axum oneshot + wiremock):
happy path (200 OpenAI-shaped JSON), missing auth (401), unknown key
(401), forbidden model (403), unknown model (404), empty messages (400),
upstream 429 pass-through, no registered bridge (503), streaming SSE
with [DONE] sentinel. 164 tests pass workspace-wide.
Copilot AI review requested due to automatic review settings April 17, 2026 07:01
@moonming
moonming merged commit 39ea4df into main Apr 17, 2026
9 checks passed
@moonming
moonming deleted the feat/proxy-chat-route branch April 17, 2026 07:06

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR wires up the first end-to-end proxy request path for POST /v1/chat/completions, including bearer-key authentication, model resolution/authorization via snapshot, Hub-based provider dispatch, and OpenAI-shaped JSON/SSE responses.

Changes:

  • Add aisix-proxy chat-completions handler with auth extractor, Hub dispatch, OpenAI-shaped rendering, and OpenAI-style error envelopes.
  • Extend aisix-server bootstrap to build/register a Hub with all provider bridges and pass it into proxy state.
  • Add unit + wiremock integration tests for non-streaming + streaming behavior and common error cases.

Reviewed changes

Copilot reviewed 9 out of 10 changed files in this pull request and generated 5 comments.

Show a summary per file
File Description
crates/aisix-server/src/main.rs Builds a Hub at startup, registers provider bridges, and injects into ProxyState.
crates/aisix-server/Cargo.toml Adds provider bridge crates as server dependencies for Hub registration.
crates/aisix-proxy/src/state.rs Introduces ProxyState holding snapshot + Arc<Hub> + body limit config.
crates/aisix-proxy/src/render.rs Renders normalized gateway chat types into OpenAI response / chunk shapes.
crates/aisix-proxy/src/lib.rs Mounts /v1/chat/completions, updates /health, and adds extensive endpoint tests.
crates/aisix-proxy/src/error.rs Defines ProxyError → OpenAI-style {error:{message,type}} envelope + status mapping.
crates/aisix-proxy/src/chat.rs Implements request flow: validate, resolve model, authz, Hub dispatch, JSON/SSE responses.
crates/aisix-proxy/src/auth.rs Adds AuthenticatedKey extractor parsing Bearer (or x-api-key) and snapshot lookup.
crates/aisix-proxy/Cargo.toml Adds async-stream and dev deps for OpenAI bridge + wiremock tests.
Cargo.lock Locks new deps introduced by proxy/server changes.

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment on lines +41 to 47
/// Build the proxy router. Mounts `/health` plus the
/// OpenAI-compatible chat-completions surface.
pub fn build_router(state: ProxyState) -> Router {
Router::new()
.route("/health", get(health))
.route("/v1/chat/completions", post(chat::chat_completions))
.with_state(state)

Copilot AI Apr 17, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

ProxyState::request_body_limit_bytes is configured and carried in state, but build_router doesn't apply any body-size limiting layer/extractor configuration, so requests can exceed the intended limit (or just use axum's default). Consider wiring state.request_body_limit_bytes into the router via DefaultBodyLimit::max(...) / RequestBodyLimitLayer (tower-http) so the config is actually enforced.

Copilot uses AI. Check for mistakes.
Comment on lines +97 to +117
while let Some(item) = upstream.next().await {
let ev = match item {
Ok(chunk) => {
let rendered = render_chunk(created, chunk);
match serde_json::to_string(&rendered) {
Ok(json) => Event::default().data(json),
Err(err) => Event::default()
.event("error")
.data(err.to_string()),
}
}
Err(err) => Event::default()
.event("error")
.data(err.to_string()),
};
yield Ok::<_, Infallible>(ev);
}
// Emit the OpenAI-style [DONE] sentinel so clients that terminate
// on it behave correctly.
yield Ok::<_, Infallible>(Event::default().data("[DONE]"));
}

Copilot AI Apr 17, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

build_sse_stream emits the [DONE] sentinel unconditionally after draining upstream, even if the upstream stream yielded an Err. That can cause clients to treat an errored stream as a successful completion. Consider terminating the SSE stream immediately on the first upstream error (and ideally emitting an OpenAI-shaped {error:{...}} payload before closing) and only sending [DONE] on a clean upstream completion.

Copilot uses AI. Check for mistakes.
Comment on lines +53 to +58
if let Some(rest) = s.strip_prefix("Bearer ") {
let rest = rest.trim();
if rest.is_empty() {
return Err(ProxyError::MissingAuth);
}
return Ok(rest.to_string());

Copilot AI Apr 17, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

extract_bearer only accepts an Authorization scheme that exactly matches the case-sensitive prefix "Bearer ". Per HTTP auth scheme rules (RFC 9110), the scheme token is case-insensitive, and some clients send bearer. Consider parsing the scheme case-insensitively (e.g., split_once(' ') + eq_ignore_ascii_case("bearer")).

Suggested change
if let Some(rest) = s.strip_prefix("Bearer ") {
let rest = rest.trim();
if rest.is_empty() {
return Err(ProxyError::MissingAuth);
}
return Ok(rest.to_string());
if let Some((scheme, rest)) = s.split_once(' ') {
if scheme.eq_ignore_ascii_case("bearer") {
let rest = rest.trim();
if rest.is_empty() {
return Err(ProxyError::MissingAuth);
}
return Ok(rest.to_string());
}

Copilot uses AI. Check for mistakes.
Comment on lines +1 to +40
//! OpenAI-compatible error envelope used by every proxy endpoint.
//!
//! OpenAI's clients expect this exact shape (spec §3):
//!
//! ```json
//! {
//! "error": {
//! "message": "…",
//! "type": "invalid_request_error",
//! "param": null,
//! "code": null
//! }
//! }
//! ```
//!
//! `ProxyError` is the internal error taxonomy; it implements
//! `IntoResponse` so handlers can `?`-propagate without touching
//! JSON shape boilerplate.

use aisix_gateway::BridgeError;
use axum::http::StatusCode;
use axum::response::{IntoResponse, Response};
use axum::Json;
use serde::Serialize;

#[derive(Debug, Serialize, Clone)]
pub struct ErrorEnvelope {
pub error: ErrorBody,
}

#[derive(Debug, Serialize, Clone)]
pub struct ErrorBody {
pub message: String,
#[serde(rename = "type")]
pub kind: &'static str,
#[serde(skip_serializing_if = "Option::is_none")]
pub param: Option<String>,
#[serde(skip_serializing_if = "Option::is_none")]
pub code: Option<String>,
}

Copilot AI Apr 17, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The module-level docs show OpenAI errors with param: null / code: null and say clients expect that exact shape, but the implementation uses skip_serializing_if = Option::is_none which omits these fields entirely (and tests assert omission). Please either update the docs to match the actual wire format you're emitting, or change serialization to emit explicit null fields if that is the intended compatibility target.

Copilot uses AI. Check for mistakes.
Comment on lines +69 to +94
if req.is_streaming() {
let upstream = bridge.chat_stream(&req, &ctx).await?;
let model_name = req.model.clone();
let sse_stream = build_sse_stream(upstream, model_name, now);
let response =
Sse::new(sse_stream).keep_alive(KeepAlive::new().interval(Duration::from_secs(15)));
return Ok(response.into_response());
}

let upstream = bridge.chat(&req, &ctx).await?;
let rendered = render_response(now, upstream);
Ok(Json(rendered).into_response())
}

fn created_ts() -> i64 {
SystemTime::now()
.duration_since(UNIX_EPOCH)
.map(|d| d.as_secs() as i64)
.unwrap_or(0)
}

fn build_sse_stream(
upstream: aisix_gateway::ChatChunkStream,
_model: String,
created: i64,
) -> impl Stream<Item = Result<Event, Infallible>> {

Copilot AI Apr 17, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

chat_completions clones req.model into model_name and passes it into build_sse_stream, but build_sse_stream takes _model and never uses it. This is currently dead code / wasted allocation. Either remove the parameter + clone, or use it to force the rendered model field to the caller-facing model alias (instead of whatever the upstream returns) if that’s the intended behavior.

Copilot uses AI. Check for mistakes.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants