heroui logo

GCP Vertex AI High Request and Token Volume

Elastic Detection Rules

View Source
Summary
Detects bursts in Vertex AI prompt-response activity by analyzing lookback window usage in gcp_vertexai.prompt_response_logs. The rule aggregates per-model and per-project usage, counting events (Esql.event_count) and summing token usage (Esql.total_tokens_sum, Esql.prompt_tokens_sum, Esql.candidates_tokens_sum) from usage_metadata.total_token_count (which includes thinking tokens). It triggers when event_count >= 100 and total tokens >= 10,000 within the last 60 minutes (lookback), evaluated in 10-minute intervals. The query groups results by model and project, surfacing elevated activity that could indicate model extraction, quota abuse, or a compromised caller. The rule explicitly notes that individual prompt text is not required for the alert. It maps to MITRE ATT&CK (Resource Hijacking) and MITRE ATLAS (Denial of AI Service, Cost Harvesting). False positives include approved batch processing or unusually high-traffic applications; remediation focuses on credential rotation, quotas, and spend review. Triage guidance recommends inspecting related logs (audit logs for source IP and user emails), sampling prompt responses, and reviewing recent SetPublisherModelConfig changes that could affect logging visibility. The alert supports suppression by model and project for 1 hour to reduce noise, and investigation fields highlight the key dimensions to inspect (model, project, event counts, and token sums). This rule is intended for GCP Vertex AI environments and helps detect overuse, misusage, or exploitation of AI services that could incur cost impact or data exposure.
Categories
  • Cloud
Data Sources
  • Cloud Service
  • Application Log
ATT&CK Techniques
  • T0029
  • T0034
  • T1496
Created: 2026-10-02