Azure Prompt Shield
The Azure Prompt Shield middleware scans requests for direct prompt injection (jailbreak) attacks using Azure AI Content Safety's Prompt Shield API, natively from the gateway.
Prompt Shield only ever scans the request. A jailbreak or injection attempt is a property of the prompt, not the model's reply, so
unlike the other AI Gateway guards, this middleware has no separate request/response configuration.
For content moderation (hate speech, violence, and similar categories) instead of injection detection, use the separate Azure Content Safety middleware. It calls a different Azure Content Safety API and is configured independently.
Key Features
- Native Azure jailbreak and injection detection: calls Azure's Prompt Shield API directly, with no manual JSON template.
- Prompt attack detection: every extracted text, including tool output in the conversation, is checked as a user prompt. Azure's separate document-attack analysis isn't used.
- Every client request format: works with
custom,ccr,responsesAPI,messagesAPI, andmcptraffic.
Requirements
-
AI Gateway must be enabled:
helm upgrade traefik traefik/traefik -n traefik --wait \--reset-then-reuse-values \--set hub.aigateway.enabled=true -
An Azure AI Content Safety resource, its endpoint, and an API key.
How It Works
- Reads the client request according to
clientRequestFormat, extracting the fields to scan. Either the format's predefined fields, the whole body, or the paths you list injsonQueriesfor acustomformat.
When a Chat Completion, Responses API, Messages API, Bedrock Mantle, Google Agent Platform, or MCP middleware precedes this middleware in the chain, it marks the client's request format in the request context.
Prompt Shield then uses that format automatically, and no clientRequestFormat value is required. Left unset with no such middleware in the chain, it falls back to custom.
- Calls Azure's Prompt Shield API with the extracted text.
- Returns a deny response if Azure reports an attack, formatted for the active
clientRequestFormat; otherwise the request continues unchanged.
Configuration Examples
- Default format
- Custom format, scoped to one field
apiVersion: traefik.io/v1alpha1
kind: Middleware
metadata:
name: azure-prompt-shield
spec:
plugin:
azure-prompt-shield:
baseURL: https://my-resource.cognitiveservices.azure.com
apiKey: urn:k8s:secret:azure-prompt-shield-credentials:api-key
apiVersion: traefik.io/v1alpha1
kind: Middleware
metadata:
name: azure-prompt-shield-custom
spec:
plugin:
azure-prompt-shield:
baseURL: https://my-resource.cognitiveservices.azure.com
apiKey: urn:k8s:secret:azure-prompt-shield-credentials:api-key
clientRequestFormat: custom
jsonQueries:
- .prompt
onDenyResponse:
statusCode: 200
message: "Your request looks like a prompt injection attempt and was blocked."
Reference
| Parameter | Description | Required | Default |
|---|---|---|---|
baseURL | Azure Content Safety resource endpoint | Yes | |
apiKey | API key for the resource | Yes | |
apiVersion | Azure Content Safety API version | No | 2024-09-01 |
clientConfig | HTTP client settings: timeoutSeconds (default 5), maxRetries (default 3), tls, headers | No | |
clientRequestFormat | Client payload format: custom, ccr, responsesAPI, messagesAPI, or mcp | No | custom |
jsonQueries | JSON paths to scan within the request body. Only allowed for clientRequestFormat: custom. Left empty, the whole body is scanned | No | (whole body) |
onDenyResponse | Custom deny response when Prompt Shield detects an attack | No | |
onDenyResponse.statusCode | HTTP status code (100-599) | No | 403 |
onDenyResponse.message | Response body, sent as-is. Left empty, the body is the plain-text status text, for example Forbidden | No | |
onDenyResponse.contentType | Response Content-Type | No | matches clientRequestFormat, or text/plain when message is empty (except on mcp) |
jsonQueries is rejected when clientRequestFormat has a predefined extraction path (ccr, responsesAPI, messagesAPI, mcp). It only applies to custom.
clientRequestFormat and deny response formatting
You can set clientRequestFormat explicitly, or leave it unset and let the middleware detect it automatically from the request context. Either way, once a format is active, the deny response body is automatically formatted to match the client's expected format:
clientRequestFormat | Non-streaming response | Streaming response |
|---|---|---|
custom | Raw message text | Raw message text |
ccr | Chat Completion JSON with message as assistant content | SSE chunk (data: {...}) |
responsesAPI | Responses API JSON with status: "failed" and message in a structured error object | SSE events (response.failed) with the same status/error shape |
messagesAPI | Messages API JSON with message as assistant text content | SSE events (message_delta with stop_reason: refusal) |
mcp | JSON-RPC error echoing the request id | JSON-RPC error as an SSE data: event |
For a streaming client, the deny response's HTTP status code matters as much as its body shape. Most
streaming clients treat any non-2xx status as a transport error. They retry the connection instead of
reading the body. At a non-2xx status, this means the structured deny message above doesn't reach the
user. The client only sees repeated connection failures. Set onDenyResponse.statusCode to 200 so the
client reads the deny body as a normal, refused turn.
