Availability
The Edge Gateway operates as an independent, dedicated AI proxy. It processes AI requests, enforces policies, and reports analytics to the control plane. It is optimized for high performance and resilience in production.
Note: The Edge Gateway is sometimes referred to as the Microgateway in configuration files. However, Edge Gateway is the preferred terminology.
- Unified Access Point: Provides a single, consistent endpoint for applications to interact with various LLMs.
- Security Enforcement: Handles authentication, authorization, and applies security policies.
- Policy Management: Enforces rules related to budget limits, model access, and applies custom Filters.
- Observability: Logs detailed analytics data for each request, feeding the Analytics & Monitoring system.
- Vendor Abstraction: Hides the complexities of different LLM provider APIs, especially through the OpenAI-compatible endpoint.
Edge Gateway Variants
There are two Edge Gateway variants:
Both variants rely on the core proxy library and access control mechanisms. The embedded gateway is lightweight, while the Edge Gateway offers the full feature set. Tools, datasources, and OAuth state are synced to edge gateways through the hub-spoke configuration system.
Core Features
-
Request Routing: Incoming requests include an
llmSlugin their path (e.g.,/llm/call/{llmSlug}/...). The Proxy uses this slug (auto-generated from the LLM configuration name) to identify the target LLM Configuration and route the request accordingly. -
Authentication & Authorization:
- Validates the API key provided by the client application.
- Identifies the associated Application and User.
- Checks if the Application/team has permission to access the requested LLM Configuration based on RBAC rules.
- Applies the same access check to every authentication method: App API key,
Authorization: BearerApp secret, or an auth plugin. The check also applies to every entry point:/llm/,/ai/,/anthropic/,/v1, and/datasource/. An App can call only the LLMs and data sources that it has a grant for. It can also reach the fallbacks of a granted LLM’s failover waterfall, but only through that waterfall. A request for any other LLM, or for a slug that does not exist, returns403. - If an endpoint has auth plugins attached, only those plugins authenticate its requests on the Edge Gateway. The Edge Gateway refuses App keys on that endpoint.
- Policy Enforcement: Before forwarding the request to the backend LLM, the Proxy enforces policies defined in the LLM Configuration or globally:
- Analytics Logging: After receiving the response from the backend LLM (and potentially applying response Filters), the Proxy logs detailed information about the interaction (user, app, model, tokens used, cost, latency, etc.) to the Analytics database.
Endpoints
All LLM endpoints accept the same App API key. They all apply the same policies, budgets, Filters, and analytics. The endpoints differ in two ways: how the gateway selects the target LLM, and which request format they accept.
The AI Portal shows these endpoints on the detail page of each App, with the values for that App.
How the Gateway Builds the Vendor URL
The gateway sends each request to<LLM API endpoint> + <request path>, and removes any part that the two have in common. For example, the LLM endpoint can be https://api.anthropic.com, https://api.anthropic.com/v1, https://my-proxy/anthropic, or https://my-proxy/anthropic/v1. In all four cases, the gateway calls .../v1/messages. The gateway does not repeat a version segment (/v1/v1/...) and does not remove one. The same rule applies to OpenAI-compatible providers that use a path prefix, such as .../inference/v1 or .../openai/v1.
The AI Studio chat interface and agents do not use the gateway. They use the vendor’s client library directly. The Anthropic library expects an endpoint that ends in /v1, so AI Studio adds /v1 when the endpoint has no version segment. If you put your own proxy in front of a vendor, set the LLM endpoint to the path that your proxy expects immediately before the vendor’s versioned path.
Main Ingress
The Main Ingress is one OpenAI-compatible endpoint at the root of the gateway. It gives access to all the LLMs that the calling App can use. The client selects the LLM in each request, so it can change the model or the vendor without a change to its base URL or credential.- Model strings have a prefix. Send
{llmSlug}/{model}, for exampleopenai-prod/gpt-4o. The prefix selects the LLM. The gateway removes the prefix before it sends the request to the vendor. A model string without a prefix returns400. The gateway splits the string at the first slash, so model names that contain slashes (for example, fine-tune IDs) stay complete. - OpenAI format only. Requests and responses use the OpenAI Chat Completions schema for all vendors. To use a vendor’s SDK and native parameters, use the vendor-native endpoint.
- Streaming. The gateway streams the response when the request contains
stream: true. - Model discovery.
GET /v1/modelslists all the models that the authenticated App can call, with the prefix included. The gateway builds this list from the allowed models and the default model of each LLM. An LLM that has no allowed models and no default model is not in the list. - Legacy completions. The gateway also serves
/v1/completionsfor clients that use the older completions API.
/v1. You can change it, or turn off the Main Ingress. Refer to Configure the Main Ingress.
Vendor-Native Endpoint
This endpoint sends requests to the native API of one vendor. The gateway changes only the fields that it needs for analytics and budgets. It sends all of the path after the slug to the vendor without changes.- You get access to all vendor features. When a vendor adds a new parameter, it works immediately, without a gateway change.
- One URL serves streaming and non-streaming requests. The gateway examines the request and selects the correct path.
OpenAI-Compatible Endpoint
This endpoint accepts OpenAI-format requests. The gateway translates each request to the API of the target vendor, and translates the response back to OpenAI format. The URL selects one LLM, so themodel field takes a plain vendor model name. If you do not send model, the gateway uses the default model of the LLM.
- Use this endpoint when a client must use only one LLM. Also use it when a tool cannot send a model string with a slug prefix.
- Vendor features that the OpenAI schema cannot express are not available.
- Vertex and Hugging Face LLMs are not supported on this endpoint or on the Main Ingress. A request returns
400with the codeunsupported_vendor, and the gateway does not send the request to the vendor. Use/llm/rest/{llmSlug}or/llm/stream/{llmSlug}for these vendors. In a failover waterfall, the gateway skips such a fallback and tries the next one.
Anthropic Messages Endpoint
This endpoint accepts native Anthropic Messages API requests and translates them to Bedrock Converse. Clients that use the Anthropic API, such as Claude Code, can then use a Bedrock LLM through the gateway.- This endpoint supports Bedrock LLMs only. Other vendors already accept their own API on the vendor-native endpoint. A request to a non-Bedrock LLM returns
400. - Set the endpoint URL as
ANTHROPIC_BASE_URL, and set the App key asANTHROPIC_AUTH_TOKEN(orANTHROPIC_API_KEY). The client adds/v1/messagesto the URL.
Model Discovery in Claude Code
The gateway sends every request on this endpoint to the default model of the LLM. The model name that the client sends has no effect. To show this model in the Claude Code/model picker, the endpoint also answers GET /anthropic/{llmSlug}/v1/models. To use discovery, set this variable in the Claude Code environment:
<LLM name> — <model id>. For example: Bedrock Claude — eu.anthropic.claude-3-5-sonnet-20241022-v2:0. The ID is the Bedrock model ID from the LLM configuration.
- The models request uses the same authentication and App access check as
/v1/messages. It does not call a model, so it does not use budget, run filters, or record usage. - The list is empty (
"data":[]) when the LLM has no default model. It is also empty when the allowed models of the LLM do not include the default model. In both cases, Claude Code then shows its built-in list. - Use Claude Code v2.1.223 or later. Claude Code keeps a model ID only if the ID contains “claude” or “anthropic”. Versions v2.1.129 to v2.1.222 required the ID to start with one of these words. For this reason, they hid region-prefixed IDs such as
us.anthropic...oreu.anthropic.... - Claude Code does not show application inference profiles. An application inference profile ARN (
arn:aws:bedrock:...:application-inference-profile/...) does not contain “claude” or “anthropic”, so Claude Code drops it. Requests still work. Only the picker entry is missing. System inference profile ARNs contain “anthropic”, so Claude Code shows them. - Do not redirect this path. Claude Code treats any redirect as a failed discovery, including an
http://tohttps://redirect at your ingress. SetANTHROPIC_BASE_URLto the final URL.
Legacy Endpoints
/llm/rest/{llmSlug}/...accepts only non-streaming (synchronous) requests./llm/stream/{llmSlug}/...accepts only Server-Sent Events streaming requests.
/llm/call/{llmSlug}/....
Configure the Main Ingress
The Main Ingress uses the root of the gateway. This can cause a conflict when you embed the gateway in a host that already uses/v1. You can change the base path, or turn off the Main Ingress. These settings have no effect on the per-LLM endpoints (/ai/, /llm/, /anthropic/).
AI Studio (the hub) and the Edge Gateway (the edge) use different environment variables for the same setting. In a hub-and-spoke deployment, set both to the same value.
For example, a base path of
/ai-gateway/v1 moves the three routes to /ai-gateway/v1/chat/completions, /ai-gateway/v1/completions, and /ai-gateway/v1/models. The gateway normalizes the value when it loads:
- It adds a leading slash if there is none.
- It removes a trailing slash and any empty or dot segments.
- It refuses
/and uses the default, so that the Main Ingress cannot take the full gateway path.
LLM Slug
The{llmSlug} in the endpoint path is automatically generated from the LLM configuration name when you create it. For example, an LLM named “My OpenAI Config” would have a slug like my-openai-config.
Model Router
The Edge Gateway also routes requests through Model Routers. A Model Router selects an LLM from a pool of vendors by model name pattern. Clients call a router on the Main Ingress with the model string{routerSlug}/{model}, in the same way as an LLM. The embedded gateway in AI Studio does not route through Model Routers.
A Semantic Router selects the LLM from the text of the prompt. Clients call it on the Main Ingress with the model string {routerSlug}/auto.
Configuration Reference
To know more about configuring Edge Gateways, see the Configuration Reference for detailed documentation on all environment variables.Troubleshooting
Edge Shows "Disconnected"
Edge Shows "Disconnected"
- Check network connectivity between the edge and control plane
- Verify the edge gateway is running and healthy
- Check edge gateway logs for connection errors
- Ensure firewall rules allow gRPC traffic (default port 50051)
Edge Shows "Pending" After Push
Edge Shows "Pending" After Push
- Wait a few seconds for the heartbeat cycle to complete
- Check if the edge is connected (not disconnected)
- Verify the edge gateway logs for configuration load errors
- Check if the edge has sufficient permissions to fetch configuration
Checksum Mismatch Persists
Checksum Mismatch Persists
- Try pushing configuration again
- Check for configuration validation errors in edge logs
- Verify the edge and control plane are running compatible versions
- Check for database replication lag if using PostgreSQL replication