cs.CLMay 28, 2026

A Study on Question-Answer Dataset for LLM Safety Evaluation with a Focus on Illegal Activities

Authors: Kenji ImamuraMasao IdeuchiAtsushi Fujita

Organizations: National Institute of Information and Communications Technology 3-5 Hikaridai, Seika-cho, Soraku-gun, Kyoto, 619-0289, Japan

Abstract

In this paper, we discuss question-answer dataset for LLM safety evaluation, with a focus on illegal activities. Specifically, on the basis of manual analysis of AnswerCarefully, we introduce several additional information, methods for creating question-answer examples, and a rubric for evaluating LLM-generated responses. The outcomes of this study are intended to be shared with the "JAI-Trust" project.

Explore similar work

CardsList
  1. Schützen: Evaluating LLM Safety in Bulgarian and German Contexts

    Jun 9, 2026Kiril Georgiev, Yuxia Wang, Dimitar Iliyanov Dimitrov +2Large Language Model SafetyGerman