Skip to main content

Route Requests to the Right Model with NeMo Switchyard

The NeMo Switchyard middleware integrates NVIDIA NeMo Switchyard, a decision service that asks which model should serve each request. It rewrites the request's model field to whatever the decision service selects. The decision service is a separate deployment you run and point this middleware at with url. It scores signals already in the conversation, such as tool results, error severity, and repeated turns, and uses them to choose between an efficient model and a more capable one.

Use this middleware when a route serves agentic or multi-turn traffic where most requests don't need your most capable, most expensive model, but some do. Instead of a client pinning one model for an entire session, or an operator guessing at a fixed split, each request gets judged on its own signals and sent to the model that fits it.

Hub keeps the request path. The decision service only sees the request needed to make its choice, doesn't receive the credential Hub uses to reach the upstream model, and doesn't proxy traffic itself.

Key Features and Benefits​

  • Per-request model routing: send each request to the decision service and act on its answer, instead of pinning one model for an entire session.
  • Credentials stay in Hub: Hub forwards only the session, agent, and correlation headers NeMo Switchyard reads, plus User-Agent. The client's Authorization, X-Api-Key, and Cookie never reach the decision service.
  • Fails to a named model, not a 5xx: configure fallbackModel so an unreachable or erroring decision service doesn't take the route down.
  • Works across client formats: supports Chat Completions, the Responses API, and the Messages API.

Requirements​

  • You must have AI Gateway enabled:

    helm upgrade traefik traefik/traefik -n traefik --wait \
    --reset-then-reuse-values \
    --set hub.aigateway.enabled=true
  • A NeMo Switchyard deployment (or another service that implements the same /v1/decision API) reachable from the gateway.

  • Place nemo-switchyard after Chat Completion, Responses API, or Messages API in the route's middleware chain, whichever matches the client format. That middleware is what detects the client request format nemo-switchyard reads. Placed earlier, nemo-switchyard sees no detected format and skips routing on every request. nemo-switchyard still has the final say on model: it rewrites whatever that middleware set, including a model it pinned with allowModelOverride: false.

  • The route's existing upstream Service must be able to serve every model your decision service can select. See Match Upstream Models to Decision Service Targets.

  • The client, or the client-format middleware pinning a model, must send one of the decision service's own route names. For NeMo Switchyard, pin chat-completion's model (with allowModelOverride: false) to a [routes.<name>] id from your NeMo Switchyard deployment. See Route Discovery for Clients to let clients discover that name themselves.

NeMo Switchyard version

The /v1/decision endpoint requires NeMo Switchyard v0.3.0 or later.

How It Works​

  1. Reads the client request format already detected for the route (Chat Completions, Responses API, or Messages API). If the format isn't one of these three, the middleware skips routing and forwards the request unchanged.
  2. Calls the decision service at <url>/v1/decision with the original request body and the mapped input_format (openai_chat, openai_responses, or anthropic_messages). Only the session, agent, and correlation headers NeMo Switchyard reads are forwarded, such as X-Switchyard-Session-Id, X-Claude-Code-Session-Id, X-Session-Id, and X-Request-Id, plus User-Agent. Every other client header, including Authorization, X-Api-Key, and Cookie, stays in Hub. Headers configured in clientConfig.headers are added to this call. The request's model field has to name a route configured in the decision service itself, such as NeMo Switchyard's own [routes.<name>] id, not a literal upstream model. The decision service uses it to pick which routing algorithm to run.
  3. Reads the selected model from the decision service's response (selected.model) and rewrites the request body's model field to that value. It ignores any other field the response carries, such as a per-target endpoint or format (NeMo Switchyard's llm_client) or generation parameters (extra_body), and keeps using this route's existing configuration for those.
  4. Falls back on error: the middleware uses fallbackModel, when configured, if the call fails, returns a non-200 status, or returns no selected.model. Without a fallbackModel, it returns 502 Bad Gateway to the client instead.
  5. Forwards the (possibly rewritten) request to the next middleware in the chain. A guard or content-filtering middleware placed later in the chain, such as LLM Guard, keeps applying its own configured rules regardless of which model was selected.

Configuration Example​

Pin chat-completion's model to the decision service's route id, so every request that reaches nemo-switchyard already names a route it can resolve:

Chat Completion, pinned to the NeMo Switchyard route id
apiVersion: traefik.io/v1alpha1
kind: Middleware
metadata:
name: chatcompletion
spec:
plugin:
chat-completion:
token: urn:k8s:secret:ai-keys:openai-token
model: switchyard
allowModelOverride: false
apiVersion: traefik.io/v1alpha1
kind: Middleware
metadata:
name: model-router
spec:
plugin:
nemo-switchyard:
url: http://switchyard.ai.svc.cluster.local:4000
fallbackModel: gpt-4o-mini
clientConfig:
timeoutSeconds: 1
maxRetries: 0

fallbackModel names a literal model on this route's existing upstream, not a decision-service route id: nemo-switchyard sends it straight to that upstream without calling the decision service again.

Reference model-router after chat-completion on the route:

Route with chat-completion, then model-router
apiVersion: traefik.io/v1alpha1
kind: IngressRoute
metadata:
name: openai
spec:
routes:
- kind: Rule
match: Host(`ai.example.com`)
middlewares:
- name: chatcompletion
- name: model-router # must be chained after the format middleware
services:
- name: chatgpt-external
port: 443
scheme: https
passHostHeader: false

Configuration Options​

FieldDescriptionRequiredDefault
urlBase URL of the NeMo Switchyard (or compatible) decision service. Hub calls <url>/v1/decision. Must use the http or https scheme.Yes
fallbackModelModel name to use when the decision service is unreachable, returns an error, or returns no selected model. It doesn't apply when the decision service picks a model the provider doesn't serve. The provider's rejection reaches the client instead.No
clientConfigHTTP client configuration for calls to the decision service.No
clientConfig.timeoutSecondsTimeout for each attempt, in whole seconds.No5
clientConfig.maxRetriesRetries after an attempt fails. Each retry waits longer than the last: 1 second, then 2, then 4. With the default values for clientConfig.timeoutSeconds and clientConfig.maxRetries, an unreachable decision service delays each request by up to 27 seconds before fallbackModel applies.No3
clientConfig.headersCustom headers to send with requests to the decision service.No
clientConfig.tlsTLS configuration for secure connections.No
clientConfig.tls.caPEM-encoded certificate authority certificate.NoSystem CA bundle
clientConfig.tls.certPEM-encoded client certificate for mutual TLS.No
clientConfig.tls.keyPEM-encoded client private key for mutual TLS.No
clientConfig.tls.insecureSkipVerifySkip TLS certificate verification.Nofalse

Match Upstream Models to Decision Service Targets​

Hub sends the (possibly rewritten) request to the same upstream Service this route already points to. It doesn't resolve a different back end based on which model was picked, and it doesn't check that the selected model exists on that back end or validate the model name the decision service returns.

Configure every target your decision service can choose between so they resolve on the same provider or back end as this route's existing Service. For example, make the efficient and capable targets both available through the same OpenRouter endpoint, rather than pointing them at different providers. Only point url at a decision service you trust to return a valid, reachable model, since Hub forwards whatever name it selects.

Example NeMo Switchyard Configuration​

This middleware calls an existing NeMo Switchyard deployment. It doesn't configure one. The following routes.toml complements the Configuration Example. It defines the [routes.stage] id (switchyard) that chat-completion pins to, and the efficient and capable targets that resolve to the same upstream this route already points to.

routes.toml, complementing the Configuration Example
schema_version = 1

[llm_clients.upstream]
format = "openai_chat"
base_url = "https://api.openai.com/v1"

[targets.efficient]
id = "gpt-4o-mini"
llm_client = "upstream"

[targets.capable]
id = "gpt-4o"
llm_client = "upstream"

[routes.stage]
id = "switchyard"
type = "stage_router"
capable_target = "capable"
efficient_target = "efficient"
picker = "efficient_first"
confidence_threshold = 0.5

See NeMo Switchyard's own server configuration reference for every available field.

note

In this example, llm_client can be a placeholder with no api_key_env. NeMo Switchyard refuses to start when api_key_env names a variable that isn't set, so add it only with the optional classifier block. NeMo Switchyard's schema requires an llm_client on every target, but this configuration, with no classifier block, scores tool-result and activity signals alone and makes no completion or classifier call during /v1/decision. Adding NeMo Switchyard's optional classifier block to stage_router does call an LLM client for low-confidence turns.

Route Discovery for Clients​

A client needs to know the decision service's route name to request it, but this route's own GET /v1/models reflects the upstream Service this route already points to, not the decision service. Add a separate route on the same host that points a Service directly at your decision service, so GET /v1/models on that path returns its own route names instead. NeMo Switchyard already answers GET /v1/models on its own. You don't need to add chat-completion or nemo-switchyard to this route, since the request carries no body for either middleware to act on.

Separate route exposing NeMo Switchyard's own model discovery
apiVersion: traefik.io/v1alpha1
kind: IngressRoute
metadata:
name: switchyard-models
spec:
routes:
- kind: Rule
match: Host(`ai.example.com`) && Path(`/v1/models`)
services:
- name: switchyard-decision-service
port: 4000

switchyard-decision-service is a Service reaching the same host and port as this middleware's url, not the route's existing upstream Service. Traefik matches this Host-and-Path route ahead of the plainer Host-only route in Configuration Example, since it ranks a more specific rule higher by default. Set an explicit priority on both routes instead of relying on this if you want it guaranteed.

Troubleshooting​

Problem: Requests fail with 502 Bad Gateway.

Cause: the decision service is unreachable, timed out, or returned an error or an empty selected.model, and no fallbackModel is configured. Set fallbackModel to a known model name. Also check that url points to a reachable decision service.

Problem: Requests reach the upstream model unchanged, with no routing decision applied.

Cause: the route's client request format isn't Chat Completions, the Responses API, or the Messages API. The middleware skips routing for any other format and forwards the request as-is.

Problem: The upstream model provider rejects the request with an unknown-model error.

Cause: the decision service selected a model the route's upstream Service doesn't serve. See Match Upstream Models to Decision Service Targets. Configure every target the decision service can select so they all resolve on the same upstream this route already points to.

Problem: A client that lists available models through GET /v1/models doesn't see the models the decision service can select.

Cause: GET /v1/models and other bodyless requests pass through the AI Gateway middlewares unchanged and reach the route's existing upstream Service directly. The response reflects that upstream's own model list, not the decision service's route names. Add a separate route exposing the decision service's own /v1/models. See Route Discovery for Clients.

  • Configure the Chat Completion middleware to govern the model a route uses.
  • Configure the Responses API middleware for OpenAI Responses API routes.
  • Configure the Messages API middleware for Anthropic Messages API routes.