Human communities are governed by normative systems: shared standards that produce \textit{norms} dictating acceptable behavior, enforced through community sanctioning. Aligning increasingly autonomous AI systems with these norms is a central alignment challenge, complicated by the fact that norms are vast in number, change quickly, and are often arbitrary (e.g., dress or language conventions). Thus, alignment requires \textit{normative competence}: the ability to discern from interaction alone what norms a community enforces without relying on static pretrained knowledge. We introduce a multi-agent community debate setting, where access to debate is governed by synthetic norms, to study normative competence in isolation from pretraining exposure. We show that baseline LLM agents fail to learn norms even when doing so would improve their accuracy. We then experiment with various \textit{normative modules} -- architectural components for norm inference -- finding that norm-following is highly sensitive to both the style of norm and the model powering the normative module, suggesting a lack of generalizability. Furthermore, when idiosyncratic, non-normative behaviors accompany the true norm, LLM agents exhibit an unselective attribution failure: they indiscriminately copy idiosyncratic noise alongside enforced rules, a pattern that persists even when imitating unnecessary behaviors is explicitly penalized. To the best of our knowledge, our work is the first to operationalize and evaluate normative competence in LLMs, demonstrating that current AI systems excel at behavioral mimicry but lack the capacity to discern socially enforced order.
Figures & tables
Figure 1: (a) An illustration of our community debate setting for evaluating normative competence. A newcomer agent to the debate community must learn to successfully use the hidden community norm to gain access to other agents’ responses and reasoning, thus benefiting from debate. (b) We find in our environment that the newcomers which are able to follow norms better see a direct increase in accuracy, as a result of gaining access to community engagement. Depicted are all the baselines and approaches that we evaluate within our experiments. This provides strong incentive for newcomer agents to learn the community norm.
Figure 2: Norm-following in community debate. Per-round rate at which the newcomer’s messages satisfy the community norm, by Normative Module x LLM. The Without Normative Module group allows the newcomer to generate a response given a history of community interactions, but without a normative module to explicitly steer the model to learn the hidden norm. This baseline, even when powered by a frontier model like Claude Sonnet, is largely incapable of following norms; adding a normative module significantly boosts norm-following performance. All error bars are 1 standard error.
Figure 3: Even with a variety of normative module architectures and frontier-level LLMs, we find poor performance overall, as well as significant variability and weak points in following different types of norms. We consider the following types of norms: Local edit (insert a symbol, emoji, or fixed keyword); Boundary (required start/end/sentence placement); Substitution (replace a word with a pseudoword); Repeated-word (a token must appear exactly or at least N times); Global structure (length, sentence & paragraph count requirements); Register / Discourse (rhetorical or self-referential markers). Results shown are aggregated across models (left) and across normative modules (right). See § A.3 for further analysis and Table 1 for performance on individual norms.
Figure 4: Single-community norm with idiosyncratic behaviors (Direct Edit / Infer + Edit / Contrastive). We find that the better a model follows the community norm, the more it also over-imitates. This occurs even on the normative modules (ToM, Contrastive) which are explicitly designed to infer norms by attending primarily to sanctioning behavior , though we do see evidence that attending to sanctioning behavior reduces overimitation. Many of these models also copy these idiosyncratic behaviors more frequently than they follow the true norm. An example learned norm description, showing that models do not distinguish between the norm and irrelevant behaviors, is shown as well.
Figure 5: We see a steep drop in both accuracy and norm-following ability as a result of overimitation across nearly all normative modules and models when we introduce “expensive” idiosyncratic behaviors in the environment. This effect is particularly dramatic for the normative modules and models that exhibit higher overimitation rates in Fig. 4 , indicating that they in fact do not reduce their overimitation rates under this new penalty.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Norm
Following rate
nc_at_least_3_paragraphs
0.658±0.054
sy_plus_minus
0.654±0.087
pseudo_frim_min_3
0.550±0.086
sy_checkmark
0.545±0.088
sy_emoji_target
0.533±0.100
st_at_least_7_sentences
0.521±0.180
Appendix
Table 1: Norm-following rate by norm, aggregated over normative modules and models.
Figure 6: We investigate whether normative modules that can learn a more accurate description reflecting the true community norm also exhibit reduced overimitation rates due to the description “separating out” the norm from the noise.
Figure 7: Expected newcomer participation probability 1−ϵt on non-final rounds under exponential vs. linear ϵ -decay over T=50 questions ( ϵ0=1 , ϵT−1=0 ). Lines show the schedule; markers are empirical per-question means. The final round of each debate is always participate and is excluded.
Figure 8: We find that the accuracies of our chosen background and newcomer agents do not change significantly with larger numbers of debate rounds. The newcomer achieves most of its total performance gains by T=3 rounds so we set that as our cutoff.
Norm configuration
Rule type
ifeval_smiley
Symbol inclusion
tk_true_exactly_2
Exact token count
nc_under_30_words
Length constraint
st_exactly_3_paragraphs
Paragraph structure
pp_snondle_bookend
Pseudoword placement
ps_drendal_for_and
Pseudoword substitution
Appendix
Table 2: The six norms included in the ToM ablation analysis. All belong to the norm catalog.
Condition
Introspection
First order
Second order
Confidence
Calls/update
ToM reference
Yes
Yes
No
No
5
+ Second order
Yes
Yes
Yes
No
7
+ Confidence
Yes
Yes
No
Yes
5
+ Both
Yes
Yes
Yes
Yes
7
Infer + Edit
Separate Infer + Edit baseline
1
Appendix
Table 3: Conditions for two background agents. Call counts include norm inference after a debate and exclude response editing, retries, and debate-agent calls. Confidence denotes elicitation and downstream use of norm-confidence scores.
Condition
Norm following (%)
Difference from reference (pp)
ToM reference
20.0±14.8
—
Second-order inference
21.1±14.7
+1.1
Confidence elicitation
20.9±15.0
+0.9
Both additions
20.5±14.7
+0.5
Infer + Edit
2.1±1.4
−17.9
Appendix
Table 4: Mean norm following and one standard error across six norms.
Figure 9: ToM component comparison with Qwen3 8B as the normative backend, two GPT-5 Nano background agents, and one Gemma 3 4B IT newcomer. Bars show mean norm following over six norms; error bars show one standard error across norm-level runs. The audited runs use seed 0 and 50 MMLU-Pro questions.
Figure 10: Per-norm ToM ablations with Qwen3 8B. Bars use the audited seed-0 runs of 50 debates. Background and newcomer models are held fixed as in Figure 9 .
The conformity bias exhibited by large language models (LLMs) can pose a significant challenge to decision-making in LLM-based multi-agent systems (LLM-MAS). While many prior studies have treated "conformity" simply as a matter of opinion change, this study introduces the social psychological distinction between informational conformity and normative conformity in order to understand LLM conformity at the mechanism level. Specifically, we design new tasks to distinguish between informational conformity, in which participants in a discussion are motivated to make accurate judgments, and normative conformity, in which participants are motivated to avoid conflict or gain acceptance within a group. We then conduct experiments based on these task settings. The experimental results show that, among the six LLMs evaluated, up to five exhibited tendencies toward not only informational conformity but also normative conformity. Furthermore, intriguingly, we demonstrate that by manipulating subtle aspects of the social context, it may be possible to control the target toward which a particular LLM directs its normative conformity. These findings suggest that decision-making in LLM-MAS may be vulnerable to manipulation by a small number of malicious users. In addition, through analysis of internal vectors associated with informational and normative conformity, we suggest that although both behaviors appear externally as the same form of "conformity," they may in fact be driven by distinct internal mechanisms. Taken together, these results may serve as an initial milestone toward understanding how "norms" are implemented in LLMs and how they influence group dynamics.
Online group chats are social spaces with local conversational norms that are rarely stated explicitly. The ability and willingness of LLM-based agents to recognize and adapt to these norms remains mostly unexplored. We introduce LoSoNA, a benchmark for local social norm adaptation in multi-party chat. Each scenario gives a subject model a curated group-chat transcript in which non-subject participants demonstrate a hidden local norm, followed by a final elicitor turn that forces a response revealing whether the subject has inferred that norm. We evaluate eight frontier and open-weight models under four prompting conditions that vary how explicitly the model is told to treat the prior conversation as evidence for how it should answer. Naive prompting remains limited for most models; explicit norm-aware prompting helps unevenly, with Gemini 3.1 Pro reaching 84.2% and Claude Fable 5 reaching 81.6%, while several other models show small gains or regressions. LoSoNA contributes to recent calls for evaluating LLM social capabilities by testing whether models can infer local conversational norms from precedent and use them in a one-turn group-chat response.
AI agents are increasingly deployed in shared environments where they pursue diverse goals and compete for rewards. This multi-agent competition can lead to behaviors that serve individual gains at collective cost -- for instance, marketing agents may post misleading content as a result of competing for engagement on social media. Human societies address such problems through norms that constrain acceptable behavior, supported by enforcement mechanisms that detect and penalize violations. Motivated by this, we study norm enforcement mechanisms for language model agents. We find that simple enforcement mechanisms are exploited by misaligned agents for competitive advantage, even when they are not explicitly trained or prompted to do so. We thus turn our attention to designing more robust mechanisms, and identify two key ingredients: estimating each agent's reliability over time, and updating this estimate with escalating penalties for repeated misbehavior. Across three simulated environments and a variety of agent populations, mechanisms built on these principles resist exploitation, while still penalizing norm violations at comparable or lower cost than baselines. Our results position norm enforcement mechanisms as scalable levers for shaping agents' behavior, but only when designed to anticipate becoming part of the system they govern. Our code and data are available at https://yaowenye.com/norm-enforcement.