Skip to main content

Azure Content Safety

The Azure Content Safety middleware scans requests and responses with Azure AI Content Safety's text moderation API, natively from the gateway.

Azure analyzes text against four moderation categories: Hate, SelfHarm, Sexual, and Violence. It reports a severity for each. You define blockConditions that compare that severity against a threshold you choose, so the moderation policy lives in Hub configuration rather than in a hand-built request template.

For prompt injection and jailbreak detection instead of content moderation, use the separate Azure Prompt Shield middleware. It calls a different Azure Content Safety API and is configured independently.

Key Features​

  • Native Azure moderation: calls Azure's analyzeText API directly, with no manual JSON template.
  • Per-severity block conditions: block only when a category's severity is at or above a threshold you set, not on any detection at all.
  • Request and response scanning: enable either direction independently, or both.
  • Blocklist support: match a custom blocklist alongside the moderation categories.
  • Every client request format: works with custom, ccr, responsesAPI, messagesAPI, and mcp traffic.

Requirements​

  • AI Gateway must be enabled:

    helm upgrade traefik traefik/traefik -n traefik --wait \
    --reset-then-reuse-values \
    --set hub.aigateway.enabled=true
  • An Azure AI Content Safety resource, its endpoint, and an API key.

How It Works​

  1. Reads the client request according to clientRequestFormat, extracting the fields to scan. Either the format's predefined fields, or the paths you list in jsonQueries for a custom format.
Automatic format detection

When a Chat Completion, Responses API, Messages API, Bedrock Mantle, Google Agent Platform, or MCP middleware precedes this guard in the chain, it marks the client's request format in the request context. The guard then uses that format automatically, and no clientRequestFormat value is required. Left unset with no such middleware in the chain, it falls back to custom.

  1. Calls Azure's analyzeText API with the extracted text, the configured categories filter, and outputType (severity scale).
  2. Blocks immediately on any blocklist match, if blocklistNames is set. This check runs before block conditions are evaluated. It happens no matter how haltOnBlocklistHit is set. That setting only controls whether Azure also computes category severities when a blocklist match already occurred. It does not control whether Hub blocks.
  3. Otherwise, evaluates blockConditions in the order they are written, against the categories Azure reports. The content is blocked on the first matching condition, a severity at or above threshold for one of categories. A condition with no categories of its own falls back to the top-level categories filter, or every reported category when that is also empty.
  4. Returns a deny response on a match, formatted for the active clientRequestFormat; otherwise the request or response continues unchanged.

List narrower conditions first. The deny response is the same whichever condition matches, but the logged reason names the first one: list a categories: [Violence] condition before a broader condition with no categories of its own, so a Violence match logs that specific reason instead of the broader condition's, if both would otherwise match.

Configuration Examples​

apiVersion: traefik.io/v1alpha1
kind: Middleware
metadata:
name: azure-content-safety
spec:
plugin:
azure-content-safety:
baseURL: https://my-resource.cognitiveservices.azure.com
apiKey: urn:k8s:secret:azure-content-safety-credentials:api-key
categories:
- Hate
- Violence
request:
blockConditions:
- threshold: 4
response:
blockConditions:
- threshold: 4

Reference​

General​

ParameterDescriptionRequiredDefault
baseURLAzure Content Safety resource endpointYes
apiKeyAPI key for the resourceYes
apiVersionAzure Content Safety API versionNo2024-09-01
categoriesCategory filter sent to Azure, and the default for any blockConditions entry with no categories of its own. One or more of Hate, SelfHarm, Sexual, Violence. Empty matches every categoryNo(every category)
outputTypeSeverity scale returned by Azure: FourSeverityLevels or EightSeverityLevelsNoFourSeverityLevels
blocklistNamesNames of Azure blocklists to match against, in addition to the moderation categoriesNo
haltOnBlocklistHitSet to true to skip Azure's category analysis once a blocklist match is found. A blocklist match always blocks the content on its own, so the category severities Azure computes at false (the default) go unused in that caseNofalse
clientConfigHTTP client settings: timeoutSeconds (default 5), maxRetries (default 3), tls, headersNo
clientRequestFormatClient payload format: custom, ccr, responsesAPI, messagesAPI, or mcpNocustom
requestEnables scanning on the request. At least one of request/response is required. Without blockConditions, only a blocklistNames match can block itNo
responseEnables scanning on the response. At least one of request/response is required. Without blockConditions, only a blocklistNames match can block itNo

request / response​

ParameterDescriptionRequiredDefault
request.jsonQueriesJSON paths to scan in the request body. Only allowed for clientRequestFormat: customNo
request.blockConditions[].categoriesCategories this condition applies to. Falls back to the top-level categories, then every reported category, when emptyNo
request.blockConditions[].thresholdMinimum severity that blocks the content, greater than 0Yes (per condition)
request.onDenyResponseCustom deny response when a block condition matchesNo
request.onDenyResponse.statusCodeHTTP status code (100-599)No403
request.onDenyResponse.messageResponse body, sent as-is. Left empty, the body is the plain-text status text, for example ForbiddenNo
request.onDenyResponse.contentTypeResponse Content-TypeNomatches clientRequestFormat, or text/plain when message is empty (except on mcp)
response.jsonQueriesJSON paths to scan in the response body. Only allowed for clientRequestFormat: customNo
response.blockConditions[].categoriesSame as request.blockConditions[].categories, applied to the response directionNo
response.blockConditions[].thresholdSame as request.blockConditions[].threshold, applied to the response directionYes (per condition)
response.onDenyResponseCustom deny response when a block condition matchesNo
response.onDenyResponse.statusCodeHTTP status code (100-599)No403
response.onDenyResponse.messageResponse body, sent as-is. Left empty, the body is the plain-text status text, for example ForbiddenNo
response.onDenyResponse.contentTypeResponse Content-TypeNomatches clientRequestFormat, or text/plain when message is empty (except on mcp)

A blockConditions entry naming a category outside the categories filter is rejected: Azure only reports the categories it was asked to analyze, so that condition could never match.

onDenyResponse.statusCode doesn't apply when an MCP response is streamed as SSE. The status is already sent, so the block arrives as a JSON-RPC error event on the original 200.

jsonQueries is rejected when clientRequestFormat has a predefined extraction path (ccr, responsesAPI, messagesAPI, mcp). It only applies to custom.

clientRequestFormat and deny response formatting​

You can set clientRequestFormat explicitly, or leave it unset and let the middleware detect it automatically from the request context. Either way, once a format is active, the deny response body is automatically formatted to match the client's expected format:

clientRequestFormatNon-streaming responseStreaming response
customRaw message textRaw message text
ccrChat Completion JSON with message as assistant contentSSE chunk (data: {...})
responsesAPIResponses API JSON with status: "failed" and message in a structured error objectSSE events (response.failed) with the same status/error shape
messagesAPIMessages API JSON with message as assistant text contentSSE events (message_delta with stop_reason: refusal)
mcpJSON-RPC error echoing the request idJSON-RPC error as an SSE data: event

For a streaming client, the deny response's HTTP status code matters as much as its body shape. Most streaming clients treat any non-2xx status as a transport error. They retry the connection instead of reading the body. At a non-2xx status, this means the structured deny message above doesn't reach the user. The client only sees repeated connection failures. Set onDenyResponse.statusCode to 200 so the client reads the deny body as a normal, refused turn.