Skip to main content

Azure Prompt Shield

The Azure Prompt Shield middleware scans requests for direct prompt injection (jailbreak) attacks using Azure AI Content Safety's Prompt Shield API, natively from the gateway.

Prompt Shield only ever scans the request. A jailbreak or injection attempt is a property of the prompt, not the model's reply, so unlike the other AI Gateway guards, this middleware has no separate request/response configuration.

For content moderation (hate speech, violence, and similar categories) instead of injection detection, use the separate Azure Content Safety middleware. It calls a different Azure Content Safety API and is configured independently.

Key Features​

  • Native Azure jailbreak and injection detection: calls Azure's Prompt Shield API directly, with no manual JSON template.
  • Prompt attack detection: every extracted text, including tool output in the conversation, is checked as a user prompt. Azure's separate document-attack analysis isn't used.
  • Every client request format: works with custom, ccr, responsesAPI, messagesAPI, and mcp traffic.

Requirements​

  • AI Gateway must be enabled:

    helm upgrade traefik traefik/traefik -n traefik --wait \
    --reset-then-reuse-values \
    --set hub.aigateway.enabled=true
  • An Azure AI Content Safety resource, its endpoint, and an API key.

How It Works​

  1. Reads the client request according to clientRequestFormat, extracting the fields to scan. Either the format's predefined fields, the whole body, or the paths you list in jsonQueries for a custom format.
Automatic format detection

When a Chat Completion, Responses API, Messages API, Bedrock Mantle, Google Agent Platform, or MCP middleware precedes this middleware in the chain, it marks the client's request format in the request context. Prompt Shield then uses that format automatically, and no clientRequestFormat value is required. Left unset with no such middleware in the chain, it falls back to custom.

  1. Calls Azure's Prompt Shield API with the extracted text.
  2. Returns a deny response if Azure reports an attack, formatted for the active clientRequestFormat; otherwise the request continues unchanged.

Configuration Examples​

apiVersion: traefik.io/v1alpha1
kind: Middleware
metadata:
name: azure-prompt-shield
spec:
plugin:
azure-prompt-shield:
baseURL: https://my-resource.cognitiveservices.azure.com
apiKey: urn:k8s:secret:azure-prompt-shield-credentials:api-key

Reference​

ParameterDescriptionRequiredDefault
baseURLAzure Content Safety resource endpointYes
apiKeyAPI key for the resourceYes
apiVersionAzure Content Safety API versionNo2024-09-01
clientConfigHTTP client settings: timeoutSeconds (default 5), maxRetries (default 3), tls, headersNo
clientRequestFormatClient payload format: custom, ccr, responsesAPI, messagesAPI, or mcpNocustom
jsonQueriesJSON paths to scan within the request body. Only allowed for clientRequestFormat: custom. Left empty, the whole body is scannedNo(whole body)
onDenyResponseCustom deny response when Prompt Shield detects an attackNo
onDenyResponse.statusCodeHTTP status code (100-599)No403
onDenyResponse.messageResponse body, sent as-is. Left empty, the body is the plain-text status text, for example ForbiddenNo
onDenyResponse.contentTypeResponse Content-TypeNomatches clientRequestFormat, or text/plain when message is empty (except on mcp)

jsonQueries is rejected when clientRequestFormat has a predefined extraction path (ccr, responsesAPI, messagesAPI, mcp). It only applies to custom.

clientRequestFormat and deny response formatting​

You can set clientRequestFormat explicitly, or leave it unset and let the middleware detect it automatically from the request context. Either way, once a format is active, the deny response body is automatically formatted to match the client's expected format:

clientRequestFormatNon-streaming responseStreaming response
customRaw message textRaw message text
ccrChat Completion JSON with message as assistant contentSSE chunk (data: {...})
responsesAPIResponses API JSON with status: "failed" and message in a structured error objectSSE events (response.failed) with the same status/error shape
messagesAPIMessages API JSON with message as assistant text contentSSE events (message_delta with stop_reason: refusal)
mcpJSON-RPC error echoing the request idJSON-RPC error as an SSE data: event

For a streaming client, the deny response's HTTP status code matters as much as its body shape. Most streaming clients treat any non-2xx status as a transport error. They retry the connection instead of reading the body. At a non-2xx status, this means the structured deny message above doesn't reach the user. The client only sees repeated connection failures. Set onDenyResponse.statusCode to 200 so the client reads the deny body as a normal, refused turn.