LLM Refusal

LLM: Large Language Model

Momentum

0 papers in the last four weeks, with none the four weeks before. 0.0% of all new papers.

Jul 13Week of Sep 28

Latest papers 8

All topics
CardsList
  1. Mark, Don't Erase: Token Inoculation for Dual-Use Knowledge in LLMs

    Jul 21, 2026Seunghyun Lee, Dongyoon Han, Sangdoo YunLLM AlignmentLLM Refusal

  2. Faithfulness to Refusal: A Causal Audit of Neuron Selectors

    Jul 6, 2026Ananth Eswar, Pratinav Seth, Utsav Avaiya +1Transformer InterpretabilityFeature Attribution

  3. PsychoSafe: Eliciting Psychologically-Informed Refusals in Large Language Models

    Jun 8, 2026Gianluca Barmina, Federico Torrielli, Sven Harms +7Emotional Support ConversationMental Health

  4. Distinguishable Deletion: Unifying Knowledge Erasure and Refusal for Large Language Model Unlearning

    May 16, 2026Puning Yang, Junchi Yu, Qizhou Wang +3LLM RefusalLLM Safety

  5. Targeted Neuron Modulation via Contrastive Pair Search

    May 12, 2026Sam Herring, Jake Naviasky, Karan MalhotraLanguage Model SteeringLLM Interpretability

  6. FinRAG-12B: A Production-Validated Recipe for Grounded Question Answering in Banking

    May 6, 2026Denys Katerenchuk, Pablo Duboue, Keelan Evanini +6Financial ServicesLLM Grounding