cs.CLApr 22, 2026

Exploiting LLM-as-a-Judge Disposition on Free Text Legal QA via Prompt Optimization

Authors: Mohamed Hesham ElganayniRunsheng ChenSebastian NaglMatthias Grabmair

Organizations: Technical University of Munich Germany

Abstract

This work explores the role of prompt design and judge selection in LLM-as-a-Judge evaluations of free text legal question answering. We examine whether automatic task prompt optimization improves over human-centered design, whether optimization effectiveness varies by judge feedback style, and whether optimized prompts transfer across judges. We systematically address these questions on the LEXam benchmark by optimizing task prompts using the ProTeGi method with feedback from two judges (Qwen3-32B, DeepSeek-V3) across four task models, and then testing cross-judge transfer. Automatic optimization consistently outperforms the baseline, with lenient judge feedback yielding higher and more consistent gains than strict judge feedback. Prompts optimized with lenient feedback transfer better to strict judges than the reverse direction. Analysis reveals that lenient judges provide permissive feedback, yielding prompts with broader applicability, whereas strict judges produce restrictive feedback, leading to judge-specific overfitting. Our findings demonstrate algorithmically optimizing prompts on training data can outperform human-centered prompt design and that judges' dispositions during optimization shape prompt generalizability.

Explore similar work

CardsList
  1. JudgeSense: A Benchmark for Prompt Sensitivity in LLM-as-a-Judge Systems

    Apr 26, 2026Rohith Reddy Bellibatlu, Edward Raff, Wenbin ZhangLlm-As-A-JudgeJudges

  2. Magis-Bench: Evaluating LLMs on Magistrate-Level Legal Tasks

    May 8, 2026Ramon Pires, Thales Sales Almeida, Celio Larcher Junior +6Legal DomainLlm-As-A-Judge