heroui logo

GCP Vertex AI Jailbreak Prompt Heuristics

Elastic Detection Rules

View Source
Summary
Detects jailbreak/instruction-override attempts in Google Cloud Vertex AI GenerateContent prompts by analyzing Vertex AI prompt_response_logs for common jailbreak signals. The rule normalizes prompt text to lowercase and searches full_request.contents.parts.text for a broad set of patterns indicating instruction bypass (e.g., ignore previous instructions, DAN/Do Anything Now, unrestricted AI, no safety filter, jailbreak, developer mode, override your safety, hidden system prompt). It surfaces relevant fields such as model, API method, and both the prompt and model response, including finish_reason, to assess whether a jailbreak attempt was refused or complied. The rule maps to MITRE ATLAS AML.T0051 (LLM Prompt Injection) and AML.T0054 (LLM Jailbreak) with ATT&CK equivalents TA0005 (Execution) and TA0007 (Defense Evasion). It is designed for a higher-fidelity signal when Model Armor / AI Protection is enabled. The setup notes require Vertex AI prompt_response_logs and suggest enabling additional protections for stronger signals. False positives include approved red-team/evaluation prompts or non-adversarial mentions of jailbreak language. Triage involves verifying intent of the prompt, correlating with logs, and applying remediation such as credential review and input filtering, plus enabling first-party jailbreak signals where available.
Categories
  • Cloud
  • GCP
Data Sources
  • Cloud Service
ATT&CK Techniques
  • T0051
  • T0054
  • T1562
Created: 2026-10-02