The 2025 edition of the OWASP Top 10 for LLM Applications is not a tidy-up of the previous list. It reflects a change in what these systems actually are.
What moved
- System prompt leakage entered the list outright, after repeated real-world cases where extracted prompts exposed internal rules, filtering logic and permission structure.
- Vector and embedding weaknesses were added, reflecting how many teams reach for retrieval rather than fine-tuning — which makes the corpus and the vector store part of the attack surface.
- Excessive agency was substantially expanded as applications began granting models the ability to act rather than only to answer.
- Overreliance was renamed Misinformation, sharpening it onto the model generating and propagating false claims.
Why it matters for testing
A great deal of what is currently sold as AI security testing is a jailbreak corpus run against a chat endpoint. That measures what a model can be made to say. It says nothing about what it can be made to do.
Where a model can call tools, the interesting question is not whether it can be talked into producing something disallowed. It is what happens when an attacker steers a legitimate tool call — whether the retrieval boundary between tenants holds, whether output crossing into a shell or a database is validated, and whether a chained sequence of individually reasonable calls adds up to something that is not.
Those findings are specific to your architecture, which means they are found by hand. Automated corpora are a starting point for breadth, not the engagement.