Instead of just determining if the LLM can be jailbroken, security testers advance the attack by one step and check whether it is possible to actually break the trust boundary. The research on AI agent hijacking conducted by NIST has shown that the baseline attack success rate was at least 11%, while for red team attacks the rate was 81%.
This guide is going to cover the practical side of LLM red teaming. We are going to talk about prompt injection, RAG attacks, data exposure, misuse of agents and tools, attack chaining, and impact validation and remediation.
Source: https://qualysec.com/llm-re ...