AI Attacks attack research · no vendor filter updated 2026-08-21
// archive
Every model has an attack surface.
AI attack techniques against deployed LLMs and agents: jailbreaks, prompt injection, model extraction, evasion, and data poisoning. Every technique is documented from published research, disclosed incidents, and vendor documentation, with the source named so you can check it.
Enter the archive →Latest entries
// index10 of 21 entries
How Indirect Prompt Injection Works
Attack TechniquesLLM Jailbreak Defenses: Why Static Filters Fail
DefenseLLM Jailbreak Examples: 10 Documented Patterns
Attack TechniquesPrompt Injection vs Jailbreak: How They Differ and Why It Matters
ExplainerHow to Detect Prompt Injection: Four Approaches Ranked
DefenseAdversarial Examples Explained Simply: How Pixels Fool a Model
ExplainerOWASP Top 10 LLM Explained: Every Entry and What to Fix
LLM SecurityEvasion Attacks on Production Classifiers: Malware, Spam, Fraud
Adversarial MLPoisoning Web-Scale Training Sets: Split-View and Frontrunning
Adversarial MLAdversarial Examples Against Vision Models in 2026
Adversarial MLStart here
The reference pages on this site, independent of publication date. Each one is sourced to published research and updated as the technique or the defense against it moves.
- Ten documented jailbreak patterns Each pattern with the mechanism, the reported success rate from its source, and the signal that detects it.
- LLM jailbreak techniques explained The technique families and the defenses used against each, from persona hijacking to adversarial suffixes.
- Many-shot jailbreaking How hundreds of in-context demonstrations erode safety training, and why longer context windows make it worse.
- Model extraction via black-box queries Query budgets, extraction fidelity, and how each deployed countermeasure is defeated.
- Why static jailbreak filters fail What happens to twelve published defenses when the attacker is allowed to adapt to them.
- Red Team Plan Builder Pick the components and trust boundaries a target has; get an ordered checklist of the techniques that reach it.
Independent, specialist, and free to read
AI Attacks publishes focused, sourced guides on a single topic. No paywall, no account, no ad tracking.
Subscribe
AI Attacks — in your inbox
Practitioner-grade AI red team techniques and tooling — delivered when there's something worth your inbox.
No spam. Unsubscribe anytime.