Language Model Evasion Attacks

Recent momentum

-32%

13 papers in the last 28 days · 0.2% of indexed attention

Twelve weeks of publication activity for this topic as it is defined today.

Weekly history

Recent digests

What was published in this topic, kept on the site without email delivery.

Period ending 2026-09-21

10 new papers

A weekly snapshot of new work published in Language Model Evasion Attacks.

Period ending 2026-09-14

1 new paper

A weekly snapshot of new work published in Language Model Evasion Attacks.

Period ending 2026-09-07

1 new paper

A weekly snapshot of new work published in Language Model Evasion Attacks.

204 papers

Latest in Language Model Evasion Attacks

  1. Do LLMs Make Neural Distinguishers Wise?

    Jun 9, 2026Tatsuya Sakagami, Masashi Hisai, Naoto YanaiLanguage Model Evasion AttacksText2Cypher

  2. Steering Vectors are an Adversarial Attack Surface

    Jun 4, 2026Abzal Aidakhmetov, Donato Crisostomi, Tommaso Mencattini +3Language Model Evasion AttacksActivation Steering

  3. Automatically Attacking Software Reverse Engineering AI Agents

    May 28, 2026Brian Crawford, Justin Phillips, Patrick McClureReverse EngineeringDisassembly