
Summary
Detects Google Cloud Vertex AI GenerateContent requests where safety_settings.threshold is explicitly set to BLOCK_NONE, effectively disabling category blocking for listed harm categories when the client sends safetySettings. This weakens built-in Gemini filters and can enable disallowed content. The rule queries the GCP Vertex AI prompt_response_logs dataset, matching on full_request.safety_settings.threshold containing BLOCK_NONE and surfacing fields such as cloud.project.id, model, api_method, full_request.contents.parts.text, full_request.safety_settings.category, and full_response.candidates.safety_ratings. It is tied to MITRE ATT&CK as Defense Evasion (T1562) and to MITRE ATLAS AML.T0015 Evade AI Model, indicating an intent to bypass model safety controls. Triaging involves validating which categories were weakened, inspecting the prompt for disallowed content, checking for subsequent high/medium safety ratings or successful unsafe completions, and identifying the caller via audit logs (source.ip, client.user.email). Consider organizational policy defaults or approved wrappers that may supersede client-sent safety_settings. False positives include documented research projects that deliberately set BLOCK_NONE; use labeled service accounts and exclude those principals. Remediation focuses on enforcing stricter thresholds through org policy, admission control, or application defaults; revoke non-compliant callers; and, if abuse is confirmed, rotate credentials and review related configuration changes (e.g., SetPublisherModelConfig).
Categories
- Cloud
Data Sources
- Cloud Service
- Application Log
ATT&CK Techniques
- T0015
- T1562
Created: 2026-10-02