Overview
Speaker: Ibrahim Al Azher
Date: Friday, September 18, 2026, 2:30 PM
Room: PM103
This session covers adversarial attacks on language models and the defenses against them, organized in three parts: attacks, defenses, and tool-based attack and defense. The reading list for each part is below.
Relevant Materials
Attack
- Universal and Transferable Adversarial Attacks on Aligned Language Models
- Improved Techniques for Optimization-Based Jailbreaking on Large Language Models
- Tree of Attacks: Jailbreaking Black-Box LLMs Automatically
- PANDORA: Jailbreak GPTs by Retrieval Augmented Generation Poisoning
- A Gradient Control Method for Backdoor Attacks on Parameter-Efficient Tuning
- FLASH: Focused Layer Attention Sink Hijacking
Defense
- Safety Alignment Should Be Made More Than Just a Few Tokens Deep
- AGD: Adversarial Game Defense Against Jailbreak Attacks in Large Language Models
- Mitigating Fine-tuning based Jailbreak Attack with Backdoor Enhanced Safety Alignment
- AutoDefense: Multi-Agent LLM Defense against Jailbreak Attacks
- Defending Large Language Models Against Jailbreak Attacks via Layer-specific Editing
Tool-Based Attack and Defense
- Same Payload, Different Channel: Measuring Trust Asymmetry in Tool-Using Language Models
- Mind the GAP: Text Safety Does Not Transfer to Tool-Call Safety in LLM Agents
- Prompt Injection as Role Confusion
- LLMs Encode Harmfulness and Refusal Separately
- When Safety Speaks a Language: A Mechanistic Analysis of Safety–Language Identity Entanglement in LLMs
- Refusal in Language Models Is Mediated by a Single Direction
- Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs
- ASIDE: Architectural Separation of Instructions and Data in Language Models (ICLR 2026)