cs.CLSep 28, 2026

CRISP: Cultural Reward Modeling for Implicit Situated Propriety

Authors: Zekun Yuan, Yangfan Ye, Baohang Li, Shuaibo Zhao, Zekun Zhou, Ziming Li, Qichen Hong, Kun Chen, +1 more

Organizations: Harbin Institute of Technology · Huawei Technologies Co., Ltd · Peng Cheng Laboratory

Abstract

As large language models (LLMs) are increasingly deployed across countries and regions, the ability to recognize and respond appropriately to diverse cultural contexts becomes increasingly important. However, existing research has largely focused on cultural knowledge or tasks with predefined response spaces, while open-ended culturally situated behavior remains comparatively underexplored. In this work, we introduce CRISP-RM, a culturally situated reward model that assigns rewards according to cultural appropriateness in open-ended social scenarios. During policy optimization, we further introduce Norm Grounding Supervision (NGS), providing guidance that enhances the policy's sensitivity to relevant cultural norms. To construct culturally situated data, we employ a collaborative multi-agent framework that instantiates implicit cultural norms into diverse social scenarios and further curate NormCompass as a dedicated testbed. We conduct comprehensive experiments to evaluate the effectiveness of CRISP-RM in both reward modeling and policy optimization. Best-of-NN experiments show that CRISP-RM consistently outperforms strong general reward models. During GRPO policy optimization, CRISP-RM generally improves culturally situated behavior, while incorporating NGS yields further gains. Further analyses demonstrate the advantages of CRISP-RM in distinguishing culturally appropriate behavior beyond superficial fluency and politeness, while NGS provides complementary gains during policy optimization by improving norm grounding.

Figures & tables

Appendix figures & tables23 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. CultureForest: Understanding and Evaluating Cultural Norm Grounded Reasoning in LLMs

    Jun 1, 2026Yangfan Ye, Xiaocheng Feng, Jialong Tang +5Cultural AwarenessTest Time

  2. Steerable Cultural Preference Optimization of Reward Models

    Jun 17, 2026Minsik Oh, Advit Deepak, Sophie Wu +2Large Language Model AlignmentSequential Preference Optimization

  3. Whose Norms? Disentangling Cultural and Personal Alignment in Large Language Models

    Jun 5, 2026Angana Borah, Isabelle Augenstein, Rada MihalceaLarge Language Model AlignmentLarge Language Model Bias