Test-time scaling and post-training have improved LLM performance in coding and mathematical reasoning, but their effectiveness for individual stance prediction remains unclear. We study this question by predicting a person's stance in a new discussion from their history. We evaluate widely used test-time scaling strategies and post-training methods, such as supervised fine-tuning and reinforcement learning, and identify four failure modes across generation, selection, and learning: (1) incorrect consensus, where repeated samples agree on the wrong stance; (2) selection failure, where generation covers the observed stance but selection misses it; (3) response overfitting, where supervised fine-tuning improves imitation but harms prediction; and (4) early plateau, where reinforcement learning shows modest initial gains followed by limited further improvement. We expose these failures using STANCE-BENCH, which contains 2499 prediction tasks from 500 Hacker News users. Guided by this analysis, we explore a simple approach that combines direct scores for all candidate stances with explicit assessments of support from the individual's history. On the 781-task test set, this approach achieves 21.83 discussion-specific Macro F1 with Qwen3-8B, compared with 19.27 for direct scoring. Our results motivate evaluating candidate generation, final selection, and person-specific evidence use separately. Our data is available at https://github.com/stance-bench/Stance-Bench.
Figures & tables
Figure 1: Overview. Stance-Bench is built by collecting Hacker News threads, inducing discussion-specific stances, and annotating user contributions. Given a user’s history, the discussion, and its stance options, a model predicts the user’s stance. We identify four failure modes: incorrect consensus (samples agree on a wrong stance), selection failure (the observed stance is generated but not selected), response overfitting (SFT fits responses better but predicts stances worse), and early plateau (RL gains little after early training, and reward contrast stays sparse).
Figure 2: Topic composition of all 2499 Stance-Bench tasks.
Figure 3: Candidate coverage and final prediction accuracy for (a) Gemini 3.7 Flash, (b) Gemini 3.1 Pro, and (c) Qwen3-8B. Coverage measures whether the target user’s observed stance appears among the generated candidates and accuracy measures whether the final prediction matches that stance. Their difference identifies cases in which the observed stance is available but not selected.
Figure 4: Training diagnostics for Qwen3-8B. (a) Training and validation average token loss during comment-and-label SFT with 3 random seeds. (b) Sampled training stance-match rate and first-response validation accuracy during GRPO. (c) Fractions of eight-response rollout groups with all responses incorrect, all correct, or zero reward advantages. Curves show means across three training seeds; shading and error bars indicate ±1 sample standard deviation. The SFT epoch-0 point is a shared base-model evaluation without an error bar. GRPO rollout statistics use non-overlapping 50-step windows.
Gemini 3.7 Flash
Gemini 3.1 Pro
Qwen3-8B
Method
k
Acc. ↑
F1 ↑
Acc. ↑
F1 ↑
Acc. ↑
F1 ↑
Direct generation
1
27.53
22.43
24.46
22.05
26.89 ± 1.30
20.53 ± 1.37
Verbalized Sampling
4
25.48
21.46
21.77
19.97
20.57 ± 0.07
16.67 ± 0.47
Universal Self-Consistency
4
27.27
22.41
24.58
21.69
27.83 ± 0.27
20.66 ± 0.24
AlphaCode-style consensus
4
27.53
22.30
24.58
21.84
27.49 ± 0.66
20.64 ± 0.42
PlanSearch
4
21.25
18.14
23.30
20.65
22.36 ± 1.42
18.42 ± 1.49
Table 1: Test-time scaling across Gemini and Qwen. Accuracy (Acc.) and discussion-specific Macro F1 (F1) are percentages on test set. Qwen entries report mean ± sample s.d. over three generation seeds. k counts prediction candidates and AlphaCodium uses iterative refinement. Candidate counts do not equalize total inference compute. Bold marks column maxima.
Method
Accuracy ↑
Macro F1 ↑
Direct generation
26.89 ± 1.30
20.53 ± 1.37
Direct scoring
25.86
19.27
LEVER (verifier, k=16 )
27.53
20.13
LEVER (product, k=16 )
25.48
19.78
SFT
19.50 ± 1.37
16.35 ± 1.19
GRPO
27.78 ± 0.90
21.35 ± 0.95
Table 2: Learning and adaptation on Qwen3-8B. Test-set accuracy and Macro F1 (%). Subscripted entries report means with sample s.d. over three generation seeds (direct generation and steering) or training seeds (SFT, GRPO and profiles). Other entries are deterministic or single fitted systems. Bold marks column maxima, not significance. Settings and matched controls are in Appendix C.3 .
Variant
Accuracy ↑
Macro F1 ↑
Full method ( aB+bS )
28.04
21.83
Direct scores only ( b=0 )
25.48
19.00
Evidence scores only ( a=0 )
23.94
20.06
Different-user history
25.65
19.93
Recent history
27.91
22.30
Retrieved comments in direct prompt
27.53
19.94
Table 3: Ablations of direct and evidence scores on Qwen3-8B. We report accuracy and Macro F1 (%) on the test set. Branch ablations set one coefficient to zero. History substitutions keep the direct scores and full-method coefficients fixed. Different-user results average three donor assignments. The final row appends the same retrieved comments to the direct prompt. Shading identifies the full method. Details are in Appendix C.4 .
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Table 4: Diagnostic coverage of existing evaluations and Stance-Bench . R1: users compared within a shared context; R2: personal evidence removed or replaced with the prediction target fixed; R3: candidate and final predictions recorded separately in a shared stance space. ULSP (CB): user-level stance prediction on Connected Behaviour. ✓ : supported; : partial; ✗ : unsupported; ?: unclear from the reported evaluation.
Method
T
p
ktop
Token cap
Prediction readout
Direct generation
0.2
Default
Default
1,024
One response
LEVER (both variants)
0.95
Default
Default
512
Rank 16 responses
SFT / SFT + user profile
0.95
0.95
20
2,048
One response
GRPO
1.0
0.95
20
384
First of eight responses
Latent profile
0.95
1.0
Default
2,048
One response
Steering vector
0.95
1.0
Default
2,048
One response
Appendix
Table 5: Test decoding for Qwen3-8B baselines. T denotes temperature; p and ktop denote nucleus and top- k sampling. Token caps apply per response. “Default” denotes the generation setting inherited from the loaded model or server configuration.
Stance detection identifies the attitude of a text author toward a given target. Recent studies have explored various LLM-based strategies for this task, from zero-shot prompting to multi-agent debate. However, existing works differ in data splits, base models, and evaluation protocols, making fair comparison difficult. We conduct a systematic comparison that evaluates five methods across two categories -- prompt-based inference (Direct Prompting, Auto-CoT, StSQA) and agent-based debate (COLA, MPRF) -- on four datasets with 14 subtasks, using 15 LLMs from six model families with parameter sizes from 7B to 72B+. Our experiments yield several findings. First, on all models with complete results, the best prompt-based method outperforms the best agent-based method, while agent methods require 7 to 12 times more API calls per sample. Second, model scale has a larger impact on performance than method choice, with gains plateauing around 32B. Third, reasoning-enhanced models (DeepSeek-R1) do not consistently outperform general models of the same size on this task.
Genan Dai, Zini Chen, Yi Yang +1
School of Artificial Intelligence, Shenzhen Technology University, Shenzhen, China
Existing approaches to LLM personalization focus on constructing better personalized models or inputs, while treating inference as a single-shot process. In this work, we study Test-Time Personalization (TTP) along an unexplored axis: scaling inference-time computation by sampling N candidates from a personalized policy model and selecting the best with a personalized reward model. We prove that oracle selection yields expected utility growing logarithmically with the number of sampled candidates, establishing a theoretical ceiling for test-time scaling. However, standard reward models fail to realize this potential. To diagnose why, we derive a unified scaling law that decomposes any reward model's Best-of-N curve into four measurable quantities and reveals two failure modes, user-level collapse (near-constant prediction for some users) and query-level reward hacking (negative correlation with true quality for some queries). Guided by this law, we propose a probabilistic personalized reward model whose learned variance effectively mitigates both failure modes. Experiments confirm both elements of our framework: TTP delivers consistent scaling across multiple policy models and personalized text generation tasks, and our scaling law closely matches observed scaling curves across reward-model variants.
Large language models (LLMs) are increasingly used as surrogates for human participants, but it remains unclear which models best capture human behavior and why. To address this, we introduce Psych-201, a novel dataset that enables us to measure behavioral alignment at scale. We find that post-training -- the stage that turns base models into useful assistants -- consistently reduces alignment with human behavior across model families, sizes, and objectives. Moreover, this misalignment widens in newer model generations even as base models continue to improve. Finally, we find that persona-induction -- a popular technique for eliciting human-like behavior by conditioning models on participant-specific information -- does not improve predictions at the level of individuals. Taken together, our results suggest that the very processes that are currently employed to turn LLMs into useful assistants also make them less accurate models of human behavior.