SolidLab

Comprehensive security assessment (penetration testing)
of applications using AI

The adoption of generative AI and LLM agents introduces fundamental changes to application architecture that require a new approach to security testing. Traditional methods and tools do not provide the expected effectiveness for AI systems. We identify and demonstrate AI-specific risks that others would miss.

Why AI systems need specialized assessment

AI components differ from classical applications in three key characteristics, each of which creates new security challenges.

Probabilistic behavior

LLMs generate different responses to identical prompts — the output is nondeterministic and can change for the same input data. This makes typical deterministic tests insufficient and requires adaptive assessment methods that account for the probabilistic nature of AI.

AI-specific threat classes

Prompt Injection, Jailbreak, data disclosure, integration errors and excessive agent privileges are threats specific to AI systems and require specialized methods that go beyond the classic OWASP Top 10 for web applications.

Expanded attack surface

An AI agent becomes an independent execution unit with its own privileges. Multi-agent systems create additional exposure: every new tool can become an attack point, and every agent connected to internal services can become a privilege escalation point. The attack surface expands significantly and cannot be fully covered by classic vulnerability scanners or protected by a WAF.

Real threats for AI applications

Each of these threats can lead to data compromise, business disruption or reputational losses.

Prompt Injection

Attackers can take control of an AI agent through specially crafted prompts and gain access to data and actions on behalf of the system.

Confidential data leakage

An LLM may disclose sensitive information in responses — user data, internal documents, API keys or infrastructure details, including database content used for RAG.

Bypassing business restrictions

The agent context may not account for all aspects of the embedded business logic, which can be abused to bypass restrictions even without specialized payloads (for example, for an attack via prompt injection).

Excessive agent privileges

AI agents with excessive access rights to tools and services create opportunities for large-scale attacks, including use of integrated tools to access internal infrastructure, conduct spam attacks and phishing.

Excessive resource consumption

Targeted overload of system resources through AI components.

Who needs AI system assessment

Fintech and banks

AI chatbots, scoring and antifraud — protection against manipulation and customer data leaks.

SaaS and product companies

AI assistants, content generation and automation — protection against abuse of functionality.

E-commerce

AI recommendations and support chatbots — prevention of customer data leaks.

Healthcare

AI diagnostics and medical data processing — strict confidentiality requirements.

Government organizations

AI systems for request processing and analytics — protection of highly valuable information.

Telecom and industry

AI for network and infrastructure management — security of critical systems.

What we assess

Our methodology covers all layers of applications using AI — from the user interface to internal agent integrations. The checklists below are indicative, and the exact scope is determined depending on the target system. We also assess classic ML systems based on OWASP Top 10 for Machine Learning, covering the entire lifecycle from data preparation to production operation.

Architecture

Architecture and attack surface analysis

A comprehensive review from the user interface and API to specific agents in multi-agent systems, including LLMs, integrations and security mechanisms.

  • Identifying agent components and entry points for direct and indirect prompt injection
  • Understanding the agent’s business logic: goals, restrictions and tools used
  • Analyzing data flows between users, LLMs and tools
  • Identifying implemented security mechanisms
  • Identifying critical paths in multi-agent systems
OWASP LLM

AI-specific vulnerability analysis

Testing according to OWASP recommendations and methodologies for AI systems, including OWASP Top 10 for LLM Applications, OWASP Top 10 for Agentic Applications, OWASP MCP Top 10, MITRE ATLAS and Google SAIF.

  • Direct and indirect Prompt Injection
  • Insecure output handling
  • Confidential data leakage
  • Excessive resource consumption and other AI-specific risks
Integrations

Connected integrations and access rights analysis

Review of connected integrations, resources, tools, RAG systems and other agents, as well as the privileges available to agents.

  • Access rights audit for tools and APIs, including OWASP API Top 10
  • Searching for known API vulnerability classes and opportunities to generate payloads using LLMs
  • RAG system analysis for AI-specific vulnerabilities
  • Authorization checks for calls to external services
  • Horizontal and vertical privilege escalation analysis
  • Isolation assessment between user sessions
  • Checking for extraction of system prompts and configuration
  • Analysis of architecture and agent hierarchy in multi-agent systems
  • Attempting to exploit trust chains between agents in multi-agent systems
Security controls

Implemented security mechanisms analysis

Analysis of AI system components for hidden capabilities, potentially dangerous actions or other excessive privileges.

  • Jailbreak and content filter bypass
  • Agent restriction analysis
  • Creating and testing agent abuse scenarios
  • Testing business-logic restriction bypass scenarios, including request chains
Business logic

GenAI functionality abuse

Searching for opportunities to abuse implemented AI functionality in order to violate application business logic.

  • Manipulating agent behavior using system specifics and weaknesses discovered during previous assessment stages
  • Exploiting logical errors in agent chains

Assessment stages

A structured process ensuring completeness and transparency at every step.

01

Scope definition

Agreeing the scope, identifying AI components, reviewing documentation and system architecture.

02

Architecture analysis

Defining the attack surface: LLM models, agents, integrations, APIs, user interfaces and data flows.

03

Implementation analysis

Active vulnerability search according to OWASP AI methodologies, MITRE ATLAS and Google SAIF, integration testing and security control verification.

04

Vulnerability demonstration and risk analysis

Demonstrating real impact of identified vulnerabilities, assessing severity and business impact.

05

Report and recommendations

Preparing a detailed report with technical details, severity assessment and specific remediation recommendations.

What you get

A detailed report on identified risks, findings and their severity, along with recommendations for remediation and improving the overall security level.

Executive summary

A concise overview of results, key risks and the overall security level of the AI system — for high-level decision-making.

Technical report

Detailed description of each finding: root cause, demonstration steps, proof of concept, severity assessment using CVSS and AI-specific context.

Remediation recommendations

Concrete and practical instructions for developers and DevOps teams to fix each vulnerability and improve the overall protection level.

✓Retest after remediation
✓Consultations on remediation of identified issues

Why us

Our group of companies has been working in information security for more than 15 years. The team conducts research at the Faculty of Computational Mathematics and Cybernetics of Moscow State University, and the results have been presented at leading conferences — OWASP AppSec, DEF CON, Black Hat, Hack in the Box, Positive Hack Days, OffZone and ZeroNights.

Research expertise

Key employees conduct security research, publish and present at leading conferences.

Proven results

Numerous published CVEs in Apple products, Google Chrome, VMware vCenter, and MySQL2. Over 100 accepted reports in bug bounty programs from leading global companies.

Proprietary technologies

Our own tools and expertise allow us to analyze complex applications and APIs with intelligent result validation.

Confidentiality

An NDA is signed before work begins. All confidentiality requirements are strictly followed. All information is securely deleted after project completion.

Industry recognition

Acknowledgements from Alibaba, Amazon, Apple, Google, IBM, PlayStation, Mail.ru and other companies for responsible vulnerability disclosure.

Critical vulnerabilities discovered by our team

Apple

CVE-2025-24192

Google Chrome

CVE-2023-5480 · CVE-2024-10229 · CVE-2025-4664

MySQL2 for Node.js

CVE-2024-21507 · CVE-2024-21508 · CVE-2024-21509

Questions and answers

How does AI system assessment differ from standard web application assessment or penetration testing?

AI system assessment requires accounting for nondeterministic component behavior: the same input data can produce different LLM responses. The assessment may include searching for known web vulnerability classes, but primarily focuses on AI-specific vulnerabilities, LLM restriction mechanisms, AI-agent privileges and business-logic bypasses through AI components.

What types of AI systems do you test?

We analyze a wide range of systems: web applications integrated with LLMs, multi-agent systems, RAG systems, AI chatbots and assistants, automation systems based on AI agents, and custom solutions based on open models.

Do you need access to source code?

Assessment is possible in Black Box, Grey Box and White Box modes. For maximum coverage we recommend Grey/White Box. However, even in Black Box mode we can identify critical vulnerabilities at the user interface and external API levels.

How do you ensure the security of our data?

An NDA is signed before work begins. We protect the provided data, strictly follow confidentiality requirements and securely delete all information after project completion. Testing can be performed in an isolated environment.

What happens after the report is delivered?

We provide consultations on remediation of identified vulnerabilities and perform retesting after fixes to confirm their effectiveness. Consulting support for the remediation process is available if needed.

Find vulnerabilities in your AI application before attackers do

Contact us to discuss the security assessment of your AI system. We will estimate the scope and propose an optimal testing plan.