Organizations: Japan National Institute of Occupational Safety and Health, Kanagawa, Japan · Kaze To Taiyo, Tokyo, Japan · Saga Occupational Health Association, Saga, Japan · Department of Pharmacy, Zikei Hospital/Zikei Institute of Psychiatry, Okayama, Japan · Department of Medical Welfare, Suzuka University of Medical Science, Mie, Japan · Graduate School of Human Sciences, Ritsumeikan University, Osaka, Japan · Faculty of Nursing, National Defense Medical College, Saitama, Japan · Support Center for Students with Disabilities, Aoyama Gakuin University, Tokyo, Japan
Large language models (LLMs) may support counseling training, yet evidence from Japanese-language interactions and automated quality ratings remains limited. We examined 18 fixed Japanese-language counseling transcripts generated through artificial intelligence (AI)-to-AI interactions under three counselor conditions: GPT-minimal (GPT-4-turbo with a minimal role instruction), GPT-SMDP (GPT-4-turbo with the Structured Multi-step Dialogue Prompt [SMDP]), and Claude-SMDP (Claude-3-Opus with SMDP). Fifteen counseling experts rated transcripts on four adapted global scales from the Motivational Interviewing Treatment Integrity coding manual and an overall-quality item; three newer LLMs independently rated the same transcripts in three iterations. In this fixed stimulus set, SMDP-condition dialogues received higher expert ratings for cultivating change talk, partnership, empathy, and overall quality than GPT-minimal dialogues; the two SMDP counselor models did not differ. LLM ratings were reproducible but generally more lenient than expert-reference ratings, particularly for softening sustain talk and overall quality. Simulated-client naturalness was below the scale midpoint. These findings provide an expert-referenced benchmark for Japanese-language AI counseling simulations and show that reproducible LLM ratings should not be treated as calibrated counseling-quality evidence without expert validation. This study does not test clinical effectiveness or human-client outcomes.