Sabotage

Recent momentum

emerging

0 papers in the last 28 days · 0.0% of indexed attention

Twelve weeks of publication activity for this topic as it is defined today.

17 papers

Latest in Sabotage

Open your feed →
CardsList
  1. Diffuse AI Control on Fuzzy Tasks

    Jun 8, 2026Mikhail Terekhov, Caglar Gulcehre, Vivek Hebbar +1Artificial Intelligence SafetySabotage

  2. Training on Documents About Monitoring Leads to CoT Obfuscation

    May 14, 2026Reilly Haskins, Bilal Chughtai, Joshua EngelsObfuscationSabotage

  3. Removing Sandbagging in LLMs by Training with Weak Supervision

    Apr 23, 2026Emil Ryd, Henning Bartsch, Julian Stastny +2SabotageElicitation