Azure Content Safety
The Azure Content Safety middleware scans requests and responses with Azure AI Content Safety's text moderation API, natively from the gateway.
Azure analyzes text against four moderation categories: Hate, SelfHarm, Sexual, and Violence. It reports a severity for each.
You define blockConditions that compare that severity against a threshold you choose, so the moderation policy lives in Hub
configuration rather than in a hand-built request template.
For prompt injection and jailbreak detection instead of content moderation, use the separate Azure Prompt Shield middleware. It calls a different Azure Content Safety API and is configured independently.
Key Features
- Native Azure moderation: calls Azure's
analyzeTextAPI directly, with no manual JSON template. - Per-severity block conditions: block only when a category's severity is at or above a threshold you set, not on any detection at all.
- Request and response scanning: enable either direction independently, or both.
- Blocklist support: match a custom blocklist alongside the moderation categories.
- Every client request format: works with
custom,ccr,responsesAPI,messagesAPI, andmcptraffic.
Requirements
-
AI Gateway must be enabled:
helm upgrade traefik traefik/traefik -n traefik --wait \--reset-then-reuse-values \--set hub.aigateway.enabled=true -
An Azure AI Content Safety resource, its endpoint, and an API key.
How It Works
- Reads the client request according to
clientRequestFormat, extracting the fields to scan. Either the format's predefined fields, or the paths you list injsonQueriesfor acustomformat.
When a Chat Completion, Responses API, Messages API, Bedrock Mantle, Google Agent Platform, or MCP middleware precedes this guard in the chain, it marks the client's request format in the request context.
The guard then uses that format automatically, and no clientRequestFormat value is required. Left unset with no such middleware in the chain, it falls back to custom.
- Calls Azure's
analyzeTextAPI with the extracted text, the configuredcategoriesfilter, andoutputType(severity scale). - Blocks immediately on any blocklist match, if
blocklistNamesis set. This check runs before block conditions are evaluated. It happens no matter howhaltOnBlocklistHitis set. That setting only controls whether Azure also computes category severities when a blocklist match already occurred. It does not control whether Hub blocks. - Otherwise, evaluates
blockConditionsin the order they are written, against the categories Azure reports. The content is blocked on the first matching condition, a severity at or abovethresholdfor one ofcategories. A condition with nocategoriesof its own falls back to the top-levelcategoriesfilter, or every reported category when that is also empty. - Returns a deny response on a match, formatted for the active
clientRequestFormat; otherwise the request or response continues unchanged.
List narrower conditions first. The deny response is the same whichever condition matches, but the logged
reason names the first one: list a categories: [Violence] condition before a broader condition with no
categories of its own, so a Violence match logs that specific reason instead of the broader condition's,
if both would otherwise match.
Configuration Examples
- Block on high-severity hate or violence
- Per-category thresholds and a blocklist
apiVersion: traefik.io/v1alpha1
kind: Middleware
metadata:
name: azure-content-safety
spec:
plugin:
azure-content-safety:
baseURL: https://my-resource.cognitiveservices.azure.com
apiKey: urn:k8s:secret:azure-content-safety-credentials:api-key
categories:
- Hate
- Violence
request:
blockConditions:
- threshold: 4
response:
blockConditions:
- threshold: 4
apiVersion: traefik.io/v1alpha1
kind: Middleware
metadata:
name: azure-content-safety-strict
spec:
plugin:
azure-content-safety:
baseURL: https://my-resource.cognitiveservices.azure.com
apiKey: urn:k8s:secret:azure-content-safety-credentials:api-key
blocklistNames:
- banned-terms
haltOnBlocklistHit: true
request:
blockConditions:
- categories: [Sexual]
threshold: 2
- categories: [Hate, Violence]
threshold: 4
onDenyResponse:
statusCode: 200
message: "Your request has been blocked."
Reference
General
| Parameter | Description | Required | Default |
|---|---|---|---|
baseURL | Azure Content Safety resource endpoint | Yes | |
apiKey | API key for the resource | Yes | |
apiVersion | Azure Content Safety API version | No | 2024-09-01 |
categories | Category filter sent to Azure, and the default for any blockConditions entry with no categories of its own. One or more of Hate, SelfHarm, Sexual, Violence. Empty matches every category | No | (every category) |
outputType | Severity scale returned by Azure: FourSeverityLevels or EightSeverityLevels | No | FourSeverityLevels |
blocklistNames | Names of Azure blocklists to match against, in addition to the moderation categories | No | |
haltOnBlocklistHit | Set to true to skip Azure's category analysis once a blocklist match is found. A blocklist match always blocks the content on its own, so the category severities Azure computes at false (the default) go unused in that case | No | false |
clientConfig | HTTP client settings: timeoutSeconds (default 5), maxRetries (default 3), tls, headers | No | |
clientRequestFormat | Client payload format: custom, ccr, responsesAPI, messagesAPI, or mcp | No | custom |
request | Enables scanning on the request. At least one of request/response is required. Without blockConditions, only a blocklistNames match can block it | No | |
response | Enables scanning on the response. At least one of request/response is required. Without blockConditions, only a blocklistNames match can block it | No |
request / response
| Parameter | Description | Required | Default |
|---|---|---|---|
request.jsonQueries | JSON paths to scan in the request body. Only allowed for clientRequestFormat: custom | No | |
request.blockConditions[].categories | Categories this condition applies to. Falls back to the top-level categories, then every reported category, when empty | No | |
request.blockConditions[].threshold | Minimum severity that blocks the content, greater than 0 | Yes (per condition) | |
request.onDenyResponse | Custom deny response when a block condition matches | No | |
request.onDenyResponse.statusCode | HTTP status code (100-599) | No | 403 |
request.onDenyResponse.message | Response body, sent as-is. Left empty, the body is the plain-text status text, for example Forbidden | No | |
request.onDenyResponse.contentType | Response Content-Type | No | matches clientRequestFormat, or text/plain when message is empty (except on mcp) |
response.jsonQueries | JSON paths to scan in the response body. Only allowed for clientRequestFormat: custom | No | |
response.blockConditions[].categories | Same as request.blockConditions[].categories, applied to the response direction | No | |
response.blockConditions[].threshold | Same as request.blockConditions[].threshold, applied to the response direction | Yes (per condition) | |
response.onDenyResponse | Custom deny response when a block condition matches | No | |
response.onDenyResponse.statusCode | HTTP status code (100-599) | No | 403 |
response.onDenyResponse.message | Response body, sent as-is. Left empty, the body is the plain-text status text, for example Forbidden | No | |
response.onDenyResponse.contentType | Response Content-Type | No | matches clientRequestFormat, or text/plain when message is empty (except on mcp) |
A blockConditions entry naming a category outside the categories filter is rejected: Azure only reports the categories it was asked to analyze, so that condition could never match.
onDenyResponse.statusCode doesn't apply when an MCP response is streamed as SSE. The status is already sent, so the block arrives as a JSON-RPC error event on the original 200.
jsonQueries is rejected when clientRequestFormat has a predefined extraction path (ccr, responsesAPI, messagesAPI, mcp). It only applies to custom.
clientRequestFormat and deny response formatting
You can set clientRequestFormat explicitly, or leave it unset and let the middleware detect it automatically from the request context. Either way, once a format is active, the deny response body is automatically formatted to match the client's expected format:
clientRequestFormat | Non-streaming response | Streaming response |
|---|---|---|
custom | Raw message text | Raw message text |
ccr | Chat Completion JSON with message as assistant content | SSE chunk (data: {...}) |
responsesAPI | Responses API JSON with status: "failed" and message in a structured error object | SSE events (response.failed) with the same status/error shape |
messagesAPI | Messages API JSON with message as assistant text content | SSE events (message_delta with stop_reason: refusal) |
mcp | JSON-RPC error echoing the request id | JSON-RPC error as an SSE data: event |
For a streaming client, the deny response's HTTP status code matters as much as its body shape. Most
streaming clients treat any non-2xx status as a transport error. They retry the connection instead of
reading the body. At a non-2xx status, this means the structured deny message above doesn't reach the
user. The client only sees repeated connection failures. Set onDenyResponse.statusCode to 200 so the
client reads the deny body as a normal, refused turn.
