GCP Model Armor
The GCP Model Armor middleware scans requests and responses against a template you configure natively in GCP Model Armor, without hand-writing a raw JSON template.
Model Armor has no block-or-mask setting of its own on the Hub side. The template's own filter configuration decides that per
direction, and the middleware relays the result: a block becomes onDenyResponse, and a de-identify match is rewritten in place.
Key Features
- Native GCP template evaluation: reuses a Model Armor template you already configured instead of reimplementing its filters in Hub.
- Request and response scanning: enable either direction independently, or both, across every client request format, buffered and streamed.
- Response scanning with request context:
useRequestPromptsends the request text alongside the response call, so Model Armor judges the reply together with the prompt that produced it. - Flexible GCP authentication: a service account key, or Application Default Credentials.
Requirements
-
AI Gateway must be enabled:
helm upgrade traefik traefik/traefik -n traefik --wait \--reset-then-reuse-values \--set hub.aigateway.enabled=true -
A Model Armor template already created in your GCP project and region.
-
GCP credentials with the Model Armor User role (
roles/modelarmor.user) on the project: a service account key, or Application Default Credentials.
How It Works
- Reads the client request according to
clientRequestFormat, extracting the fields to scan. Either the format's predefined fields, or the paths you list injsonQueriesfor acustomformat.
When a Chat Completion, Responses API, Messages API, Bedrock Mantle, Google Agent Platform, or MCP middleware precedes this middleware in the chain, it marks the client's request format in the request context.
Model Armor then uses that format automatically, and no clientRequestFormat value is required. Left unset with no such middleware in the chain, it falls back to custom.
- Calls the Model Armor template in
projectId/locationId, identified bytemplateId:SanitizeUserPromptfor therequestdirection,SanitizeModelResponsefor theresponsedirection. - Relays the template's verdict:
- A blocking filter match returns a deny response (
onDenyResponse). - A de-identify match rewrites the matched text in place, using the replacement Model Armor returns.
- A blocking filter match returns a deny response (
- Sends request context for response scanning, when
useRequestPromptistrue: the request text is captured after the request direction's own masking, so a de-identified value is never re-sent to Model Armor. - Authenticates to GCP using the service account key in
credentials, or Application Default Credentials whencredentialsis empty.
Configuration Examples
- Response scanning with request context
- Request and response, custom format
apiVersion: traefik.io/v1alpha1
kind: Middleware
metadata:
name: gcp-model-armor
spec:
plugin:
gcp-model-armor:
projectId: my-gcp-project
locationId: europe-west1
templateId: my-model-armor-template
# Resolved from the "credentials.json" key of the gcp-model-armor-credentials secret, in the middleware's namespace.
credentials: urn:k8s:secret:gcp-model-armor-credentials:credentials.json
useRequestPrompt: true
response:
onDenyResponse:
statusCode: 200
message: "Your response has been blocked."
apiVersion: traefik.io/v1alpha1
kind: Middleware
metadata:
name: gcp-model-armor-custom
spec:
plugin:
gcp-model-armor:
projectId: my-gcp-project
locationId: europe-west1
templateId: my-model-armor-template
credentials: urn:k8s:secret:gcp-model-armor-credentials:credentials.json
clientRequestFormat: custom
request:
jsonQueries:
- .prompt
reason: prompt-scan
response:
jsonQueries:
- .completion
reason: response-scan
Reference
General
| Parameter | Description | Required | Default |
|---|---|---|---|
projectId | GCP project the Model Armor template belongs to | Yes | |
locationId | Region the template is deployed in | Yes | |
templateId | Identifier of the Model Armor template to call | Yes | |
credentials | Service account key (JSON), resolved from a Kubernetes secret (urn:k8s:secret:<name>:<key>) or inline. Falls back to Application Default Credentials when empty | No | (Application Default Credentials) |
allowedCredentialTypes | Credential types accepted from credentials or Application Default Credentials, for example service_account, authorized_user, external_account. Add authorized_user to use gcloud user credentials | No | [service_account] |
endpoint | Overrides the regional endpoint derived from locationId | No | derived from locationId |
sourceLanguage | Language code hinting the language of the scanned content, for example fr. Multi-language detection is always on, so leaving this empty means auto-detect rather than English-only | No | (auto-detect) |
clientConfig | HTTP client settings: timeoutSeconds (default 5, the time allowed for the whole scan including retries), maxRetries (default 3), tls, headers | No | |
useRequestPrompt | Sends the request text as context alongside a response scan, so Model Armor judges the reply together with its prompt | No | false |
clientRequestFormat | Client payload format: custom, ccr, responsesAPI, messagesAPI, or mcp | No | custom |
request | Enables scanning on the request. At least one of request/response is required | No | |
response | Enables scanning on the response. At least one of request/response is required | No |
request / response
| Parameter | Description | Required | Default |
|---|---|---|---|
request.jsonQueries | JSON paths to scan in the request body. Only allowed for clientRequestFormat: custom | No | |
request.reason | Labels this direction in traces, metrics, and logs. Never sent to the client | No | |
request.onDenyResponse | Custom deny response when the template blocks the request | No | |
request.onDenyResponse.statusCode | HTTP status code (100-599) | No | 403 |
request.onDenyResponse.message | Response body, sent as-is. Left empty, the body is the plain-text status text, for example Forbidden | No | |
request.onDenyResponse.contentType | Response Content-Type | No | matches clientRequestFormat, or text/plain when message is empty (except on mcp) |
response.jsonQueries | JSON paths to scan in the response body. Only allowed for clientRequestFormat: custom | No | |
response.reason | Labels this direction in traces, metrics, and logs. Never sent to the client | No | |
response.onDenyResponse | Custom deny response when the template blocks the response | No | |
response.onDenyResponse.statusCode | HTTP status code (100-599) | No | 403 |
response.onDenyResponse.message | Response body, sent as-is. Left empty, the body is the plain-text status text, for example Forbidden | No | |
response.onDenyResponse.contentType | Response Content-Type | No | matches clientRequestFormat, or text/plain when message is empty (except on mcp) |
A request or response block left present but empty (for example, a label written as request= with no value) decodes to nil
and silently disables that direction. Set at least one of request/response, and check the applied configuration if scanning
appears to have no effect.
onDenyResponse.statusCode doesn't apply when an MCP response is streamed as SSE. The status is already sent, so the block arrives as a JSON-RPC error event on the original 200.
jsonQueries is rejected when clientRequestFormat has a predefined extraction path (ccr, responsesAPI, messagesAPI, mcp). It only applies to custom.
clientRequestFormat and deny response formatting
You can set clientRequestFormat explicitly, or leave it unset and let the middleware detect it automatically from the request context. Either way, once a format is active, the deny response body is automatically formatted to match the client's expected format:
clientRequestFormat | Non-streaming response | Streaming response |
|---|---|---|
custom | Raw message text | Raw message text |
ccr | Chat Completion JSON with message as assistant content | SSE chunk (data: {...}) |
responsesAPI | Responses API JSON with status: "failed" and message in a structured error object | SSE events (response.failed) with the same status/error shape |
messagesAPI | Messages API JSON with message as assistant text content | SSE events (message_delta with stop_reason: refusal) |
mcp | JSON-RPC error echoing the request id | JSON-RPC error as an SSE data: event |
For a streaming client, the deny response's HTTP status code matters as much as its body shape. Most
streaming clients treat any non-2xx status as a transport error. They retry the connection instead of
reading the body. At a non-2xx status, this means the structured deny message above doesn't reach the
user. The client only sees repeated connection failures. Set onDenyResponse.statusCode to 200 so the
client reads the deny body as a normal, refused turn.
