
Summary
This ES|QL rule detects Elevated Safety Ratings in Google Cloud Vertex AI prompt-response logs. It flags GenerateContent responses where safety_ratings.probability is HIGH or MEDIUM, which signals a stronger risk of disallowed content even if generation is blocked. Requires the client to include safetySettings so those ratings are logged. The detection query operates on logs-gcp_vertexai.prompt_response_logs-* within a 60-minute window (now-60m, interval 10m) and filters for records where data_stream.dataset equals gcp_vertexai.prompt_response_logs and the candidate safety_ratings.probability is HIGH or MEDIUM. It returns key fields such as the full request text, safety_settings.threshold, finish_reason, safety_ratings.category/probability/severity, and model/api_method for triage. The rule is mapped to MITRE ATLAS Defense Evasion (AML.TA0007) and carries a medium risk_score (47). Investigation fields emphasize correlating the prompt, response safety ratings, and thresholds, plus pivoting to audit logs (logs-gcp_vertexai.auditlogs-*) to identify the caller (source IP, user email). Triage steps include reviewing the prompt, noting category/probability/severity, verifying threshold settings, and examining related alerts from the same project/model. Setup notes state that prompt_response_logs must have candidates.safety_ratings.* populated, and that safety_settings must be present in the GenerateContent request. False positives can arise from authorized red-team or testing workloads, as well as borderline prompts that naturally score MEDIUM. Remediation guidance covers quarantining malicious callers, revoking Vertex credentials, and reviewing IAM policies. If BLOCK_NONE was used, enforce stricter thresholds (e.g., BLOCK_MEDIUM_AND_ABOVE) via org policy or defaults. This rule supports proactive threat detection and policy enforcement in GenAI workflows by highlighting scenarios where unsafe content risks are signaled by elevated safety ratings even when not fully generated.
Categories
- GCP
- Cloud
Data Sources
- Cloud Service
Created: 2026-10-02