Skip to main content

Messages API

The Messages API middleware promotes any route to an Anthropic Messages API endpoint. It adds GenAI metrics, central governance of model/parameters, and (optionally) lets clients override them.

Key Features and Benefits​

  • One-line enablement: attach middleware, your route becomes an Anthropic Messages API endpoint.
  • Governance: lock or allow overrides for model, temperature, topP, topK, and more.
  • Metrics: emits OpenTelemetry GenAI spans and counters, including Anthropic-specific cache_creation_input_tokens and cache_read_input_tokens.
  • Prompt caching: attach an Anthropic ephemeral cache_control marker to the last cacheable block for reduced inference costs.
  • Works for local or cloud models: All you need is a Kubernetes Service pointing at the upstream host.

Requirements​

  • You must have AI Gateway enabled:

    helm upgrade traefik traefik/traefik -n traefik --wait \
    --reset-then-reuse-values \
    --set hub.aigateway.enabled=true
  • If routing to a cloud LLM provider (for example, Anthropic), define a Kubernetes ExternalName service.

How It Works​

  1. Intercepts the request and validates it against the Anthropic Messages API schema.
  2. Injects authentication by setting X-Api-Key or Authorization: Bearer headers from the configured credentials.
  3. Applies governance by rewriting model or param fields if overrides are denied.
  4. Marks the request format as messagesAPI in the request context so other AI middlewares in the chain (Content Guard, LLM Guard, Semantic Cache, Token Rate Limit) automatically adapt to the Anthropic payload format.
  5. Starts a GenAI span and records the input tokens.
  6. Forwards the (possibly rewritten) request to the upstream LLM.
  7. Records usage metrics from the response (model, input/output tokens, cache tokens, latency).

Configuration Example​

apiVersion: traefik.io/v1alpha1
kind: Middleware
metadata:
name: messagesapi
spec:
plugin:
messages-api:
apiKey: urn:k8s:secret:ai-keys:anthropic-api-key
model: claude-opus-4-5
allowModelOverride: false
allowParamsOverride: true
systemPrompt: "You are a helpful assistant. Answer only questions related to our product."
allowInlineSystemMessages: false # requires Claude Opus 4.8+
params:
temperature: 1
topP: 0.9
topK: 50
maxTokens: 1024

Configuration Options​

FieldDescriptionRequiredDefault
apiKeyURN of a Kubernetes Secret holding the Anthropic API key. Sets the X-Api-Key request header (for example, urn:k8s:secret:<secretname>:<key>)No
authTokenURN of a Kubernetes Secret holding a bearer token. Sets the Authorization: Bearer header (for example, urn:k8s:secret:<secretname>:<key>)No
modelDefault model to use (for example, claude-opus-4-5, claude-3-7-sonnet-latest). Required when allowModelOverride is false. Starting in v3.21.0, must be unset when it is trueConditional
allowModelOverridetrue = clients may set the model field; false = middleware rewrites to modelNoauto (true if model empty, else false)
allowParamsOverridetrue = clients may override params; false = middleware enforces paramsNoauto (true if params empty, else false)
systemPromptGateway-enforced system prompt. Replaces the client's top-level system field. Set it to an empty string ("") to strip any client-supplied system prompt. Not subject to allowParamsOverride.No
allowInlineSystemMessagesWhen false, removes any entry in the messages[] array whose role is system. Requires Claude Opus 4.8 or later. Not subject to allowParamsOverride.Notrue
paramsBlock containing default generation parametersNo
params.temperatureSampling temperature. Higher values make output more randomNo
params.topPNucleus sampling parameter. An alternative to temperature samplingNo
params.topKOnly sample from the top K options for each subsequent tokenNo
params.maxTokensMaximum number of tokens to generate in the responseNo
params.stopSequencesCustom sequences that will cause the model to stop generatingNo
params.serviceTierService tier for routing. Must be auto or standard_onlyNo
params.inferenceGeoGeographic region for inference routingNo
params.cacheControlAttach an Anthropic ephemeral cache_control marker to the last cacheable request blockNo
params.cacheControl.ttlCache lifetime. Must be 5m or 1hNo
params.outputConfigOutput configuration optionsNo
params.outputConfig.effortControls the reasoning effort applied to the generationNo
params.outputConfig.format.schemaJSON Schema for structured output formatNo
metricsOpt-in configuration that adds extra API Management labels to the emitted GenAI metrics. The labels are sourced from the Hub-* request headers set by API Management, so they are only populated when the route is governed by API ManagementNo

The params.* fields map to the Anthropic Messages API request parameters. For accepted values and detailed semantics (for example, valid params.outputConfig.effort levels, params.inferenceGeo regions, and params.stopSequences), see the Anthropic Messages API reference.

Parameter Override Behavior​

The middleware supports two modes for handling parameters:

Mode 1: Allow Parameters Override (allowParamsOverride: true)​

When enabled (default when no params are set), the middleware acts as a default value provider:

  • If a client provides a value for a parameter, the client's value is used.
  • If a client doesn't provide a value, the configured default is applied.
spec:
plugin:
messages-api:
model: claude-opus-4-5
allowParamsOverride: true # Clients can override
params:
temperature: 0.7
maxTokens: 1000

Mode 2: Force Parameters (allowParamsOverride: false)​

When turned off, the middleware enforces the configured values:

  • All configured parameters override client values.
  • Clients cannot change these settings.
  • Useful for strict governance and cost control.
spec:
plugin:
messages-api:
model: claude-opus-4-5
allowParamsOverride: false # Enforce configured values
params:
temperature: 0.5
maxTokens: 500

Model Override Behavior​

The allowModelOverride setting controls whether clients can specify their own model:

  • allowModelOverride: true: Clients choose the model. Starting in v3.21.0, leave model unset, because the middleware refuses model together with allowModelOverride: true.
  • allowModelOverride: false (default when model is set): The configured model is always used, even if the client specifies a different one.
spec:
plugin:
messages-api:
allowModelOverride: true # Clients choose the model

Model Access Policies​

Model access policies restrict which models a caller can reach, and control what a model discovery call (GET /v1/models) returns to them. Use them when a route serves more than one caller identity: reserving an expensive model for one team, hiding an entire model family from everyone, or publishing a short, renamed catalog to a chat client such as Open WebUI instead of the provider's full list.

policies and listPolicies govern which models a caller can reach. Both evaluate the caller's JWT claims, using the same expression language as MCP TBAC. modelAliases renames a model unconditionally, without evaluating any claims:

  • policies decide whether a request naming a given model reaches the provider. Both policies and defaultAction require allowModelOverride: true. Without it, the middleware rejects the configuration and every request on the route returns 404. allowModelOverride defaults to true only when model is unset, and the middleware rejects model together with allowModelOverride: true. To use policies, leave model unset. Each request then names its own model in the model field.
  • listPolicies decide what a GET /v1/models call returns: which models are hidden from the list, and which extra models are added. They run independent of allowModelOverride. They only run against GET /v1/models. A direct GET /v1/models/{id} request bypasses them and returns that model whether or not it's hidden or denied from the list.
  • modelAliases maps an upstream model to a different name shown in the list and accepted from callers instead of the upstream identifier. Unlike policies and listPolicies, an alias has no match condition. It applies to every caller.

See Restrict and Rename AI Models by Identity for a complete walkthrough that combines all three.

spec:
plugin:
messages-api:
allowModelOverride: true
policies:
- match: Prefix(`req.model`, `claude-opus-4-5`) && Contains(`jwt.groups`, `admin`)
action: allow
- match: Prefix(`req.model`, `claude-opus-4-5`)
action: deny
defaultAction: allow
listPolicies:
- match: Prefix(`res.model.id`, `claude-haiku-`)
action: hide
modelAliases:
- model: claude-opus-4-5-20251101
exposedAs: acme-flagship

With this configuration, only callers whose JWT carries the admin group can reach claude-opus-4-5 and its dated variants, whether they request it by name or by its alias acme-flagship. Every other caller's request for either name is denied by the second policies rule. A request for any other model is allowed by defaultAction: allow.

Every caller's model list hides claude-haiku-* models and shows claude-opus-4-5-20251101 under the name acme-flagship instead of its real identifier.

defaultAction defaults to deny

When defaultAction is left unset, the middleware denies every model no rule explicitly allows, including models no rule addresses at all. A policies list containing only allow rules becomes an allowlist unless you also set defaultAction: allow, as the example in this section does to let every other model through.

Match every name the provider accepts​

Write a rule for every name the provider accepts for a model. policies compare the model name as a string, and a provider can accept more than one name for the same model. Anthropic accepts both claude-opus-4-5 and its dated ID claude-opus-4-5-20251101. Under defaultAction: allow, a deny rule on one name leaves the other reachable. The example in this section uses Prefix for this reason.

Choose the match function by the names you need to cover:

  • Equals compares the exact string, so it covers only the one name you write.
  • OneOf compares against a list of exact names, for example OneOf(`req.model`, `claude-opus-4-5`, `claude-opus-4-5-20251101`). It covers those names and nothing else, so a snapshot released later isn't covered.
  • Prefix covers every name that starts with the text you give. Use it only when no other model shares the prefix.

To avoid listing names, set defaultAction: deny and add allow rules for only the models you want.

Fields​

FieldDescriptionRequiredDefault
policiesRequest-time rules deciding whether a call may target a given model. Requires allowModelOverride: true. The middleware rejects the configuration otherwiseNo
policies[].matchExpression matching the request. The only fields the match expression can reference are the model the request targets as req.model, after modelAliases resolve, and the caller's JWT claims as jwt (for example, Prefix(`req.model`,`claude-opus-4-5`) && Contains(`jwt.groups`,`admin`)). Any other field name is accepted without an error and never matches, so a deny rule that misspells one denies nothing. An omitted match matches every requestNomatches everything
policies[].actionallow or denyYes
defaultActionOutcome for a request no policies rule matches. Requires allowModelOverride: trueNodeny
listPoliciesRules deciding what a GET /v1/models response shows for the calling identityNo
listPolicies[].matchExpression matching a model in the response. The only fields the match expression can reference are the model as res.model (the matching entry in the response body's data[] array) and the caller's JWT claims as jwt (for example, Prefix(`res.model.id`,`claude-opus-4-5`) && Contains(`jwt.groups`,`admin`)). Any other field name is accepted without an error and never matches, so a hide rule that misspells one hides nothing. An omitted match matches every modelNomatches everything
listPolicies[].actionshow, hide, or appendYes
listPolicies[].modelsModels to add to the response. Required when action is appendConditional
listPolicies[].models[].idIdentifier of the appended modelYes
listPolicies[].models[].displayNameDisplay name, returned as display_name in the Anthropic format onlyNo
listPolicies[].models[].ownedByOwner, returned as owned_by in the OpenAI format onlyNo
listPolicies[].models[].createdCreation time in Unix seconds, returned as created (OpenAI format) or created_at in RFC 3339 (Anthropic format)No
listDefaultActionOutcome for a model no listPolicies rule matches: show or hideNoshow
modelAliasesUpstream-model-to-alias mappings, applied unconditionally to every callerNo
modelAliases[].modelThe upstream model identifier being aliasedYes
modelAliases[].exposedAsThe name shown in the list and accepted from callers instead of modelYes
statusCodeOnDenyHTTP status code returned when a policies rule denies a requestNo403

For policies and for listPolicies show and hide rules, the first matching rule decides, so list the more specific rules first. append rules are independent of that order.

When listPolicies hide every model for a caller, for example under listDefaultAction: hide when no show rule matches that caller, GET /v1/models still returns 200 with an empty data array. A model that policies deny for a caller still appears in the list unless a listPolicies rule hides it.

Anthropic's GET /v1/models is paginated. listPolicies filter each page as it comes back and leave the pagination cursors (first_id, last_id, has_more) untouched, so a page can hold fewer models than its limit, and a cursor can name a hidden model. Appended models are added to the last page only.

Adding a model not in the upstream list​

Add a listPolicies rule with action: append to add an entry the upstream response doesn't list on its own, for example a fine-tuned model reachable through this route but absent from the provider's own catalog:

listPolicies:
- match: Contains(`jwt.groups`, `admin`)
action: append
models:
- id: claude-opus-4-5-ft-support
displayName: Support Fine-Tune

With this configuration, only callers whose JWT carries the admin group see claude-opus-4-5-ft-support in GET /v1/models. An append policy's match can only check jwt, since there's no existing list entry yet for it to evaluate res.model against.

The upstream has to serve GET /v1/models. append adds to that list, and an error response passes through with nothing appended.

Appending a model to the list doesn't grant access to it. If allowModelOverride is true, add a matching policies rule so a request naming claude-opus-4-5-ft-support is allowed. Otherwise defaultAction governs it like any other name.

How an alias applies end to end​

Once modelAliases maps a model, the alias replaces the identifier everywhere a caller interacts with it, for every caller:

  • The GET /v1/models response shows the alias, not the upstream identifier.
  • listPolicies match against the upstream identifier, before aliases rename it. Write res.model.id rules for the upstream name, not the alias.
  • A request naming the alias resolves to the upstream identifier before policies evaluates it. See Order of evaluation.
  • A request naming the upstream identifier directly resolves to itself. policies see the same upstream name whether the caller sent the alias or the identifier. You can't allow an alias while refusing its upstream identifier. A caller who knows claude-opus-4-5-20251101 can use it wherever acme-flagship is allowed.
  • The response reports the alias back to the caller, but observability reports the real upstream identifier in traefik.hub.gen_ai.upstream.model. See AI Gateway Observability.
  • A Token Rate Limit with costLimit prices an alias as its upstream model only when it runs after this middleware. Placed before it, the quota sees the alias, which has no price, so every aliased request is refused with 422, or counted as zero when allowMissingPricing is true.
  • Each alias must be unique in both directions: one exposed name maps to exactly one upstream model, and one upstream model maps to exactly one exposed name. A modelAliases list that breaks either rule fails validation before the middleware starts.

Order of evaluation​

Traefik Hub resolves any modelAliases mapping first, then evaluates policies against the resolved upstream model name.

Write policies rules against the upstream identifier. A request naming the alias resolves to the upstream identifier before policies sees it, so one rule governs both names:

match: Prefix(`req.model`, `claude-opus-4-5`) && Contains(`jwt.groups`, `admin`)

This rule matches a request for claude-opus-4-5, claude-opus-4-5-20251101 and acme-flagship alike. acme-flagship resolves to claude-opus-4-5-20251101 before the match runs, and Prefix covers both real names. See Step 3 of Restrict and Rename AI Models by Identity for a complete walkthrough that applies this pattern.

System Prompt Governance​

The middleware provides two optional and independent controls over system prompts. When set, they are always enforced and independent from the allowParamsOverride field. These controls protect against prompt injection, where a client attempts to override the AI's behaviour by injecting unauthorized instructions into the request.

systemPrompt​

The systemPrompt field controls the top-level system property of every Anthropic request:

systemPrompt valueEffect on the request
not setClient's system field is forwarded unchanged
"You are a helpful assistant."Client's system field is replaced with the configured value
"" (empty string)Client's system field is removed entirely

For example, inject a gateway system prompt as follows:

spec:
plugin:
messages-api:
model: claude-opus-4-5
systemPrompt: "You are a helpful assistant. Answer only questions related to our product."
Note

Traefik Hub wraps the configured value in an Anthropic text content block before forwarding: [{"type": "text", "text": "…"}]. This is valid for all Anthropic models and behaves identically to a plain string system prompt. For instance, if you use systemPrompt: "You are a helpful assistant." in your middleware configuration, Traefik sends "system": [{ "type": "text", "text": "You are a helpful assistant." }]

Alternatively, to strip any client-supplied system prompt, use the following format for the systemPrompt field:

spec:
plugin:
messages-api:
model: claude-opus-4-5
systemPrompt: ""

allowInlineSystemMessages​

Some clients may attempt to pass system instructions as {"role": "system", ...} entries inside the messages[] array. Setting allowInlineSystemMessages: false removes all such entries before the request reaches the upstream model.

note

This control requires Claude Opus 4.8 or later.

For example:

spec:
plugin:
messages-api:
model: claude-opus-4-5
allowInlineSystemMessages: false

Combining both controls​

The two fields compose independently. A fully locked-down configuration that injects a gateway prompt and blocks all client-side system content:

spec:
plugin:
messages-api:
model: claude-opus-4-5
systemPrompt: "You are a customer support agent for Acme Corp. Do not discuss competitors."
allowInlineSystemMessages: false

With this configuration:

  • The client's top-level system field is replaced with the gateway prompt.
  • Any {"role":"system"} entries in the client's messages[] array are stripped.

Prompt Caching​

Anthropic supports prompt caching to reduce costs on repeated calls with the same long context. Use params.cacheControl to instruct the middleware to attach an ephemeral cache_control marker to the last eligible request block:

spec:
plugin:
messages-api:
model: claude-opus-4-5
params:
cacheControl:
ttl: 5m # Supported values: "5m" or "1h"
Caching Costs

Anthropic charges a small creation fee for cached content but reduces costs significantly on cache hits. See Anthropic's prompt caching guide for details on supported models and minimum cacheable token counts.

Request Body Size Limits​

The middleware enforces a maximum request body size based on the AI Gateway configuration:

helm upgrade traefik traefik/traefik -n traefik --wait \
--set hub.aigateway.enabled=true \
--set hub.aigateway.maxRequestBodySize=10485760 # 10MB

Requests exceeding this size will receive a 413 Request Entity Too Large response.

Streaming Support​

The middleware fully supports streaming responses. When a client sets "stream": true in the request, the response is streamed back as server-sent events (SSE).

Metrics and Streaming

Token usage metrics are recorded for streamed responses too. They are extracted from the usage fields in the message_start and message_delta events.

Metrics​

The middleware emits OpenTelemetry GenAI metrics when metrics are enabled. When prompt caching is active, the Anthropic cache_creation_input_tokens and cache_read_input_tokens usage fields are added to the input token count reported by gen_ai.client.token.usage.

Working with Other AI Middlewares​

When the Messages API middleware is present in the middleware chain, the AI Gateway automatically propagates the Anthropic Messages API format to the other AI middlewares in the chain. Content Guard, LLM Guard, Semantic Cache, and Token Rate Limit all detect the format and adapt their behavior without any additional configuration.

For detailed examples of combining these middlewares, see the Using AI Middlewares with Messages API guide.

Middleware Order Matters

Place the Messages API middleware first in the chain among the AI kind. It sets the request format in the context, which the other AI middlewares rely on to detect the Anthropic format. Guards placed before it cannot detect the format. See Middleware Order for details.

Troubleshooting​

Request Entity Too Large

Problem: Receiving 413 Request Entity Too Large errors.

Solution: Increase the AI Gateway's max request body size:

helm upgrade traefik traefik/traefik -n traefik --wait \
--reset-then-reuse-values \
--set hub.aigateway.maxRequestBodySize=20971520 # 20MB
Model Override Denied

Problem: Client's model selection is being overridden.

Solution: Remove model and enable model override in the middleware:

spec:
plugin:
messages-api:
allowModelOverride: true # Clients choose the model
Model Denied by Policy

Problem: Receiving a 403 (or the configured statusCodeOnDeny) response with a message like The model "X" is not allowed.

Solution: Check the policies list for a rule that allows the caller's JWT claims to reach that model, and check defaultAction. It denies any model no rule matches. Remember policies require allowModelOverride: true. Without it, the middleware rejects the configuration and every request on the route returns 404.

Authentication Errors

Problem: Receiving 401 Unauthorized responses from Anthropic.

Solution: Ensure the API key secret exists and is referenced correctly:

# Check secret exists
kubectl get secret ai-keys -n apps

# Verify secret content
kubectl get secret ai-keys -n apps -o yaml

Reference in middleware using the URN format:

apiKey: urn:k8s:secret:ai-keys:anthropic-api-key

For workspace API keys or alternative authentication methods, use authToken instead:

authToken: urn:k8s:secret:ai-keys:anthropic-bearer-token
Prompt Caching Not Working

Problem: cache_read_input_tokens is always 0 in metrics.

Solution: Ensure the cacheControl.ttl is set and the request payload exceeds Anthropic's minimum cacheable token count for your model. Verify your model supports prompt caching by checking the Anthropic prompt caching documentation.