Where Do Test-Time Scaling and Training Fall Short in Individual Stance Prediction?
Organizations: University of California, San Diego
Abstract
Test-time scaling and post-training have improved LLM performance in coding and mathematical reasoning, but their effectiveness for individual stance prediction remains unclear. We study this question by predicting a person's stance in a new discussion from their history. We evaluate widely used test-time scaling strategies and post-training methods, such as supervised fine-tuning and reinforcement learning, and identify four failure modes across generation, selection, and learning: (1) incorrect consensus, where repeated samples agree on the wrong stance; (2) selection failure, where generation covers the observed stance but selection misses it; (3) response overfitting, where supervised fine-tuning improves imitation but harms prediction; and (4) early plateau, where reinforcement learning shows modest initial gains followed by limited further improvement. We expose these failures using STANCE-BENCH, which contains 2499 prediction tasks from 500 Hacker News users. Guided by this analysis, we explore a simple approach that combines direct scores for all candidate stances with explicit assessments of support from the individual's history. On the 781-task test set, this approach achieves 21.83 discussion-specific Macro F1 with Qwen3-8B, compared with 19.27 for direct scoring. Our results motivate evaluating candidate generation, final selection, and person-specific evidence use separately. Our data is available at https://github.com/stance-bench/Stance-Bench.
Figures & tables
| Gemini 3.7 Flash | Gemini 3.1 Pro | Qwen3-8B | |||||
| Method | Acc. | F1 | Acc. | F1 | Acc. | F1 | |
| Direct generation | 1 | 27.53 | 22.43 | 24.46 | 22.05 | 26.89 1.30 | 20.53 1.37 |
| Verbalized Sampling | 4 | 25.48 | 21.46 | 21.77 | 19.97 | 20.57 0.07 | 16.67 0.47 |
| Universal Self-Consistency | 4 | 27.27 | 22.41 | 24.58 | 21.69 | 27.83 0.27 | 20.66 0.24 |
| AlphaCode-style consensus | 4 | 27.53 | 22.30 | 24.58 | 21.84 | 27.49 0.66 | 20.64 0.42 |
| PlanSearch | 4 | 21.25 | 18.14 | 23.30 | 20.65 | 22.36 1.42 | 18.42 1.49 |
| Method | Accuracy | Macro F1 |
|---|---|---|
| Direct generation | 26.89 1.30 | 20.53 1.37 |
| Direct scoring | 25.86 | 19.27 |
| LEVER (verifier, ) | 27.53 | 20.13 |
| LEVER (product, ) | 25.48 | 19.78 |
| SFT | 19.50 1.37 | 16.35 1.19 |
| GRPO | 27.78 0.90 | 21.35 0.95 |
| Variant | Accuracy | Macro F1 |
|---|---|---|
| Full method ( ) | 28.04 | 21.83 |
| Direct scores only ( ) | 25.48 | 19.00 |
| Evidence scores only ( ) | 23.94 | 20.06 |
| Different-user history | 25.65 | 19.93 |
| Recent history | 27.91 | 22.30 |
| Retrieved comments in direct prompt | 27.53 | 19.94 |
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
| Method | Token cap | Prediction readout | |||
|---|---|---|---|---|---|
| Direct generation | 0.2 | Default | Default | 1,024 | One response |
| LEVER (both variants) | 0.95 | Default | Default | 512 | Rank 16 responses |
| SFT / SFT + user profile | 0.95 | 0.95 | 20 | 2,048 | One response |
| GRPO | 1.0 | 0.95 | 20 | 384 | First of eight responses |
| Latent profile | 0.95 | 1.0 | Default | 2,048 | One response |
| Steering vector | 0.95 | 1.0 | Default | 2,048 | One response |