A newly released AI security checklist presents 222 specific tests designed to evaluate autonomous AI systems across 20 distinct attack categories. This comprehensive guide addresses a wide array of vulnerabilities, extending beyond typical prompt injection concerns to encompass infrastructure, cloud access, tools, memory, and agent communications.
Addressing Common Security Oversights
The checklist aims to fill a prevalent gap in security testing. Often, teams invest significant time in attempting to manipulate a model, inadvertently neglecting vulnerabilities such as exposed MLflow servers or accessible cloud metadata endpoints. These oversights can lead to compromised credentials and sensitive data breaches without requiring sophisticated model attacks.
Developed by security researcher Ravi Rajput, the checklist is grounded in the principles of the OWASP Web Security Testing Guide, offering a structured framework for systematic security assessments. Rajput’s independent spreadsheet provides testing goals, methodologies, recommended tools, expected outcomes, and severity ratings, although it is not an official OWASP resource.
Four Phases and 20 Attack Categories
The checklist organizes the security tests into four phases. Initially, teams map the attack surface, focusing on discovery, orchestration, cloud identity, and model supply chains. Subsequently, they explore inputs, prompt injection, system prompt leaks, and unsafe output handling. The third phase involves testing tools, excessive agency, memory, agent networks, and Model Context Protocol servers.
The final phase addresses deployment pipelines, privilege escalation, lateral movement, persistence, data theft, resource exhaustion, integrity failures, and multimodal input vulnerabilities. This structured approach assists testers in understanding an agent’s reach before assessing the impact of potentially harmful instructions.
Test Severity and Practical Examples
The checklist includes 75 tests rated as Critical, 108 as High, 30 as Medium, and nine as Low, highlighting the potential severity of vulnerabilities rather than confirmed flaws in specific products. Notable examples include cloud credential theft through server-side request forgery and unsafe Python pickle loading leading to remote code execution.
Other tests focus on risks like cross-customer document access and unauthorized data transfers through tool combinations. For memory and retrieval systems, tests examine whether removing tenant identifier filters could expose another customer’s documents, emphasizing inter-customer boundary security.
Rajput advises teams to clearly define their testing scope, document excluded checks, and maintain logs or screenshots to support each result. Destructive tests should be conducted only with explicit authorization, preferably in staging environments.
The downloadable spreadsheet aligns tests with OWASP and MITRE ATLAS frameworks, offering extensive coverage and a detailed record of conducted tests, thus enhancing AI system protection.
