heroui logo

GCP Vertex AI Response Blocked by Safety Filters

Elastic Detection Rules

View Source
Summary
Detects Vertex AI GenerateContent responses where the candidate's finish_reason is SAFETY, indicating a provider-side safety filter blocked or truncated the completion. The rule uses an ES|QL query against logs-gcp_vertexai.prompt_response_logs to identify records with finish_reason containing SAFETY and surfaces fields such as model, API method, the original prompt text, safety_settings.category/threshold, and safety_ratings alongside finish_reason. It supports triage by correlating with audit logs for source IP and user, checking for bursts of SAFETY finishes, and assessing whether a project is authorized to test safety filters (or if thresholds were weakened). The rule includes a False Positive analysis focusing on aggressive safety settings used in red-team or content-moderation workloads. Remediation emphasizes confirming authorization, revoking credentials if probing is confirmed, and monitoring for related STOP completions after threshold changes. The rule is mapped to MITRE ATLAS AML.TA0007 Defense Evasion to reflect potential attacker attempts to bypass safeguards. Setup requires Vertex AI prompt_response_logs with full_response.candidates.finish_reason. Risk score 47, severity medium, ES|QL rule type, with references to Vertex AI docs.
Categories
  • Cloud
  • GCP
Data Sources
  • Cloud Service
Created: 2026-10-02