Skip to content
← All services

Penetration testing service

AI & LLM Application Security Assessment

Security testing for applications built on large language models — the chat interface, the RAG pipeline, and the tools the model is allowed to call. Mapped to the OWASP Top 10 for LLM Applications.

What the test covers

  • Prompt injection, direct and indirect — including payloads delivered through retrieved documents, uploaded files and third-party content the model ingests
  • System prompt extraction, and what an attacker learns from it: internal rules, filtering logic, permission structure
  • Sensitive information disclosure — training data, other users’ context, secrets held in the prompt chain
  • Excessive agency: what the model is permitted to do through its tools, and what happens when an attacker steers those tools
  • Agent and tool-call abuse, including chained calls and multi-step attack paths across an agent workflow
  • RAG pipeline security — corpus poisoning, embedding and vector store weaknesses, retrieval-boundary bypass between tenants

Deliverables

  • Executive summary written for a non-technical reader
  • Findings mapped to the OWASP Top 10 for LLM Applications (2025) and to MITRE ATLAS techniques
  • Working reproduction steps for every finding, including the exact prompts and payloads used
  • Prioritised remediation roadmap covering guardrails, tool permissions and architecture
  • One retest round with a signed closure letter

Out of scope

  • Model fairness, bias and alignment evaluation — a different discipline from security testing
  • Formal EU AI Act conformity assessment, which requires a notified body
  • Training infrastructure and MLOps pipeline review, unless separately scoped
  • Accuracy benchmarking of model output