Route Requests to the Right Model with NeMo Switchyard
The NeMo Switchyard middleware integrates NVIDIA NeMo Switchyard, a decision service that asks which model should serve each request.
It rewrites the request's model field to whatever the decision service selects.
The decision service is a separate deployment you run and point this middleware at with url.
It scores signals already in the conversation, such as tool results, error severity, and repeated turns, and uses them to choose between an efficient model and a more capable one.
Use this middleware when a route serves agentic or multi-turn traffic where most requests don't need your most capable, most expensive model, but some do. Instead of a client pinning one model for an entire session, or an operator guessing at a fixed split, each request gets judged on its own signals and sent to the model that fits it.
Hub keeps the request path. The decision service only sees the request needed to make its choice, doesn't receive the credential Hub uses to reach the upstream model, and doesn't proxy traffic itself.
Key Features and Benefits
- Per-request model routing: send each request to the decision service and act on its answer, instead of pinning one model for an entire session.
- Credentials stay in Hub: Hub forwards only the session, agent, and correlation headers NeMo Switchyard reads, plus
User-Agent. The client'sAuthorization,X-Api-Key, andCookienever reach the decision service. - Fails to a named model, not a 5xx: configure
fallbackModelso an unreachable or erroring decision service doesn't take the route down. - Works across client formats: supports Chat Completions, the Responses API, and the Messages API.
Requirements
-
You must have AI Gateway enabled:
helm upgrade traefik traefik/traefik -n traefik --wait \--reset-then-reuse-values \--set hub.aigateway.enabled=true -
A NeMo Switchyard deployment (or another service that implements the same
/v1/decisionAPI) reachable from the gateway. -
Place
nemo-switchyardafter Chat Completion, Responses API, or Messages API in the route's middleware chain, whichever matches the client format. That middleware is what detects the client request formatnemo-switchyardreads. Placed earlier,nemo-switchyardsees no detected format and skips routing on every request.nemo-switchyardstill has the final say onmodel: it rewrites whatever that middleware set, including a model it pinned withallowModelOverride: false. -
The route's existing upstream
Servicemust be able to serve every model your decision service can select. See Match Upstream Models to Decision Service Targets. -
The client, or the client-format middleware pinning a
model, must send one of the decision service's own route names. For NeMo Switchyard, pinchat-completion'smodel(withallowModelOverride: false) to a[routes.<name>] idfrom your NeMo Switchyard deployment. See Route Discovery for Clients to let clients discover that name themselves.
The /v1/decision endpoint requires NeMo Switchyard v0.3.0 or later.
How It Works
- Reads the client request format already detected for the route (Chat Completions, Responses API, or Messages API). If the format isn't one of these three, the middleware skips routing and forwards the request unchanged.
- Calls the decision service at
<url>/v1/decisionwith the original request body and the mappedinput_format(openai_chat,openai_responses, oranthropic_messages). Only the session, agent, and correlation headers NeMo Switchyard reads are forwarded, such asX-Switchyard-Session-Id,X-Claude-Code-Session-Id,X-Session-Id, andX-Request-Id, plusUser-Agent. Every other client header, includingAuthorization,X-Api-Key, andCookie, stays in Hub. Headers configured inclientConfig.headersare added to this call. The request'smodelfield has to name a route configured in the decision service itself, such as NeMo Switchyard's own[routes.<name>] id, not a literal upstream model. The decision service uses it to pick which routing algorithm to run. - Reads the selected model from the decision service's response (
selected.model) and rewrites the request body'smodelfield to that value. It ignores any other field the response carries, such as a per-target endpoint or format (NeMo Switchyard'sllm_client) or generation parameters (extra_body), and keeps using this route's existing configuration for those. - Falls back on error: the middleware uses
fallbackModel, when configured, if the call fails, returns a non-200status, or returns noselected.model. Without afallbackModel, it returns502 Bad Gatewayto the client instead. - Forwards the (possibly rewritten) request to the next middleware in the chain. A guard or content-filtering middleware placed later in the chain, such as LLM Guard, keeps applying its own configured rules regardless of which model was selected.
Configuration Example
Pin chat-completion's model to the decision service's route id, so every request that reaches nemo-switchyard already names a route it can resolve:
apiVersion: traefik.io/v1alpha1
kind: Middleware
metadata:
name: chatcompletion
spec:
plugin:
chat-completion:
token: urn:k8s:secret:ai-keys:openai-token
model: switchyard
allowModelOverride: false
apiVersion: traefik.io/v1alpha1
kind: Middleware
metadata:
name: model-router
spec:
plugin:
nemo-switchyard:
url: http://switchyard.ai.svc.cluster.local:4000
fallbackModel: gpt-4o-mini
clientConfig:
timeoutSeconds: 1
maxRetries: 0
fallbackModel names a literal model on this route's existing upstream, not a decision-service route id: nemo-switchyard sends it straight to that upstream without calling the decision service again.
Reference model-router after chat-completion on the route:
apiVersion: traefik.io/v1alpha1
kind: IngressRoute
metadata:
name: openai
spec:
routes:
- kind: Rule
match: Host(`ai.example.com`)
middlewares:
- name: chatcompletion
- name: model-router # must be chained after the format middleware
services:
- name: chatgpt-external
port: 443
scheme: https
passHostHeader: false
Configuration Options
| Field | Description | Required | Default |
|---|---|---|---|
url | Base URL of the NeMo Switchyard (or compatible) decision service. Hub calls <url>/v1/decision. Must use the http or https scheme. | Yes | |
fallbackModel | Model name to use when the decision service is unreachable, returns an error, or returns no selected model. It doesn't apply when the decision service picks a model the provider doesn't serve. The provider's rejection reaches the client instead. | No | |
clientConfig | HTTP client configuration for calls to the decision service. | No | |
clientConfig.timeoutSeconds | Timeout for each attempt, in whole seconds. | No | 5 |
clientConfig.maxRetries | Retries after an attempt fails. Each retry waits longer than the last: 1 second, then 2, then 4. With the default values for clientConfig.timeoutSeconds and clientConfig.maxRetries, an unreachable decision service delays each request by up to 27 seconds before fallbackModel applies. | No | 3 |
clientConfig.headers | Custom headers to send with requests to the decision service. | No | |
clientConfig.tls | TLS configuration for secure connections. | No | |
clientConfig.tls.ca | PEM-encoded certificate authority certificate. | No | System CA bundle |
clientConfig.tls.cert | PEM-encoded client certificate for mutual TLS. | No | |
clientConfig.tls.key | PEM-encoded client private key for mutual TLS. | No | |
clientConfig.tls.insecureSkipVerify | Skip TLS certificate verification. | No | false |
Match Upstream Models to Decision Service Targets
Hub sends the (possibly rewritten) request to the same upstream Service this route already points to.
It doesn't resolve a different back end based on which model was picked, and it doesn't check that the selected model exists on that back end or validate the model name the decision service returns.
Configure every target your decision service can choose between so they resolve on the same provider or back end as this route's existing Service.
For example, make the efficient and capable targets both available through the same OpenRouter endpoint, rather than pointing them at different providers.
Only point url at a decision service you trust to return a valid, reachable model, since Hub forwards whatever name it selects.
Example NeMo Switchyard Configuration
This middleware calls an existing NeMo Switchyard deployment. It doesn't configure one.
The following routes.toml complements the Configuration Example.
It defines the [routes.stage] id (switchyard) that chat-completion pins to, and the efficient and capable targets that resolve to the same upstream this route already points to.
schema_version = 1
[llm_clients.upstream]
format = "openai_chat"
base_url = "https://api.openai.com/v1"
[targets.efficient]
id = "gpt-4o-mini"
llm_client = "upstream"
[targets.capable]
id = "gpt-4o"
llm_client = "upstream"
[routes.stage]
id = "switchyard"
type = "stage_router"
capable_target = "capable"
efficient_target = "efficient"
picker = "efficient_first"
confidence_threshold = 0.5
See NeMo Switchyard's own server configuration reference for every available field.
In this example, llm_client can be a placeholder with no api_key_env. NeMo Switchyard refuses to start when api_key_env names a variable that isn't set, so add it only with the optional classifier block.
NeMo Switchyard's schema requires an llm_client on every target, but this configuration, with no classifier block, scores tool-result and activity signals alone and makes no completion or classifier call during /v1/decision.
Adding NeMo Switchyard's optional classifier block to stage_router does call an LLM client for low-confidence turns.
Route Discovery for Clients
A client needs to know the decision service's route name to request it, but this route's own GET /v1/models reflects the upstream Service this route already points to, not the decision service.
Add a separate route on the same host that points a Service directly at your decision service, so GET /v1/models on that path returns its own route names instead.
NeMo Switchyard already answers GET /v1/models on its own. You don't need to add chat-completion or nemo-switchyard to this route, since the request carries no body for either middleware to act on.
apiVersion: traefik.io/v1alpha1
kind: IngressRoute
metadata:
name: switchyard-models
spec:
routes:
- kind: Rule
match: Host(`ai.example.com`) && Path(`/v1/models`)
services:
- name: switchyard-decision-service
port: 4000
switchyard-decision-service is a Service reaching the same host and port as this middleware's url, not the route's existing upstream Service.
Traefik matches this Host-and-Path route ahead of the plainer Host-only route in Configuration Example, since it ranks a more specific rule higher by default.
Set an explicit priority on both routes instead of relying on this if you want it guaranteed.
Troubleshooting
Problem: Requests fail with 502 Bad Gateway.
Cause: the decision service is unreachable, timed out, or returned an error or an empty selected.model, and no fallbackModel is configured.
Set fallbackModel to a known model name. Also check that url points to a reachable decision service.
Problem: Requests reach the upstream model unchanged, with no routing decision applied.
Cause: the route's client request format isn't Chat Completions, the Responses API, or the Messages API. The middleware skips routing for any other format and forwards the request as-is.
Problem: The upstream model provider rejects the request with an unknown-model error.
Cause: the decision service selected a model the route's upstream Service doesn't serve.
See Match Upstream Models to Decision Service Targets.
Configure every target the decision service can select so they all resolve on the same upstream this route already points to.
Problem: A client that lists available models through GET /v1/models doesn't see the models the decision service can select.
Cause: GET /v1/models and other bodyless requests pass through the AI Gateway middlewares unchanged and reach the route's existing upstream Service directly.
The response reflects that upstream's own model list, not the decision service's route names.
Add a separate route exposing the decision service's own /v1/models. See Route Discovery for Clients.
Related Content
- Configure the Chat Completion middleware to govern the model a route uses.
- Configure the Responses API middleware for OpenAI Responses API routes.
- Configure the Messages API middleware for Anthropic Messages API routes.
