← All services
Penetration testing service
AI & LLM Application Security Assessment
Security testing for applications built on large language models — the chat interface, the RAG pipeline, and the tools the model is allowed to call. Mapped to the OWASP Top 10 for LLM Applications.
What the test covers
- Prompt injection, direct and indirect — including payloads delivered through retrieved documents, uploaded files and third-party content the model ingests
- System prompt extraction, and what an attacker learns from it: internal rules, filtering logic, permission structure
- Sensitive information disclosure — training data, other users’ context, secrets held in the prompt chain
- Excessive agency: what the model is permitted to do through its tools, and what happens when an attacker steers those tools
- Agent and tool-call abuse, including chained calls and multi-step attack paths across an agent workflow
- RAG pipeline security — corpus poisoning, embedding and vector store weaknesses, retrieval-boundary bypass between tenants
- Improper output handling, where model output reaches a browser, a shell, a database or another system unvalidated
- Unbounded consumption — cost amplification, resource exhaustion and model extraction through repeated querying
- Guardrail and content-filter bypass, tested against your actual policy rather than a generic jailbreak list
- The conventional application underneath: authentication, authorisation, tenancy isolation and the API surface, tested to the same standard as any other engagement
Deliverables
- Executive summary written for a non-technical reader
- Findings mapped to the OWASP Top 10 for LLM Applications (2025) and to MITRE ATLAS techniques
- Working reproduction steps for every finding, including the exact prompts and payloads used
- Prioritised remediation roadmap covering guardrails, tool permissions and architecture
- One retest round with a signed closure letter
Out of scope
- Model fairness, bias and alignment evaluation — a different discipline from security testing
- Formal EU AI Act conformity assessment, which requires a notified body
- Training infrastructure and MLOps pipeline review, unless separately scoped
- Accuracy benchmarking of model output