Human Judgment

Momentum

14 papers in the last four weeks, up 250% on the four weeks before. 0.1% of all new papers.

Jul 13Week of Sep 28

Latest papers 57

All topics
CardsList
  1. Collective intelligence through aggregation

    Oct 5, 2026Franz Dietrich, Christian ListCollective BehaviorsHuman Judgment

  2. Who Owns That? Evaluating Ownership Intuitions in Large Language Models

    Sep 30, 2026Xizhi Xiao, Yue Wu, Shan Xu +1Large Language Model DecisionsHuman Judgment

  3. Multidimensional Observer Model and Perceptual Dimensions of Human Image Quality Assessment

    Sep 29, 2026Sheng Zhao, Weikai Lin, Yuhao ZhuImage Quality AssessmentMultimodal Evaluation

  4. JudgeCast: Time Series Forecasting with Experience-Informed Covariate Judgements

    Sep 29, 2026Donguk Kwon, Wooseok Jeong, Dongha LeeData-Driven ForecastingTime Series Forecasting

  5. Multi-Dimensional Comparative Scale Construction for Efficient Personalized Subjective Judgment in High-Traffic Applications

    Sep 27, 2026Xianglong Shi, Shifeng Liu, Sirui Zhao +2Model JudgmentsHuman Judgment

  6. Knowing Is Not Choosing: What Explicit Verification Adds Beyond Generative Preference

    Sep 27, 2026Yilong Li, Chengpo Yan, Aayan Arish +1Large Language Model ResponsesFactual Recall

  7. Can Jev Judge Radiology Reports? Evaluating a System One Model for Clinical Factuality

    Sep 23, 2026Jiaju Huang, Hao Yang, Xinyu Ma +6Radiology Report GenerationClinician Trust

  8. The Ethics of Artificial Intelligence in Military Operations

    Sep 22, 2026Nicolas Drapier, Florian Mauberger, Aladine Chetouani +1Artificial Intelligence GovernanceAccountability

  9. Calibrated Decisions at Scale: Converting Police Crash Narratives into Probabilistic Crash Variables with a System One Model (Jev)

    Sep 21, 2026Amir Rafe, Subasish DasAccidentsCrashes

  10. Trustworthy FinAInce: Unpacking How AI-Mediated Financial Advice is Judged

    Sep 17, 2026Aryan Ramchandra Kapadia, Eshwar Chandrasekharan, Koustuv SahaAdviceFinancial Strategy Research

  11. Measuring AI Leadership: Development and Validation of a Multidimensional Measure for AI-Native Organizations

    Sep 16, 2026Mustafa Akben, Leslie CoyneHuman JudgmentExploratory Factor Analysis

  12. NovGauge: A Fine-Grained Benchmark for Diagnosing LLMs' Capability in Paper Novelty Assessment

    Sep 12, 2026Guoqiang Zhang, Kexin Tan, Ming Zhang +12Human-Annotated BenchmarkHuman Judgment

  13. Post-hoc Alignment of LLM-judges to Human Judgment Distribution

    Sep 1, 2026Sebastian Steindl, Nikos Voskarides, Alberto Gasparin +1Llm-As-A-JudgeHuman Judgment

  14. Scaling Large Reasoning Models beyond Human Supervision: A Path toward Superintelligence

    Aug 31, 2026Zhiqin Yang, Jingwen Fu, Yuhan Liu +16Process-Level SupervisionLarge Reasoning Models

  15. One AI Signal, Many Human Judgments: A Bayesian Cascade Analysis of AI-based Credibility Indicators in Online Information Spread

    Aug 31, 2026Zhuoran Lu, Weilong Wang, Yangyang Yu +4CredibilityMisinformation

  16. How Do Large Language Models Judge Social Attraction? Evidence from Theory-Grounded Persona Ratings Across Multiple LLMs and Humans

    Aug 10, 2026Hasan Mahmud, Khawaja Abaid Ullah, Mohammad Javad Khojasteh +2RaterPersonality

  17. The Judgment-Consequence Gap: LLM Moral Reasoning in Healthcare Decisions

    Aug 6, 2026Hadi Hosseini, Samarth Khanna, Leona PierceMoral ReasoningClinical Reasoning Training

  18. The ethics of artificial intelligence in the life sciences: Universality, cultural diversity and an architecture of care

    Aug 5, 2026Jean-Pierre Changeux, Gustavo Deco, Morten L. KringelbachEthicsHuman Judgment

  19. Responsibility in Multi-Agent Sequential Decision-Making: Comparing Human Judgments to Formal Models of Causal Attribution

    Aug 5, 2026Nripsuta Ani Saxena, Stelios Triantafyllou, Goran RadanovićHuman JudgmentCausal Reasoning

  20. Query Timing Produces Opposite Positional Biases Between LLMs and Humans

    Aug 1, 2026Jasin Cekinmez, Addison J. Wu, Thomas L. GriffithsPrimingBiases

  21. Why Large Language Models and Humans Converge and Diverge in Evaluating Creativity

    Jul 24, 2026Pengzhao Lyu, Yeun Joon Kim, Hanlin Xiao +1CreativityLarge Language Model Evaluation

  22. Beyond Sycophancy: Structured Resistance and Compliance in LLM Moral Reasoning

    Jul 23, 2026Baihui Wang, Bernard KochMoral ReasoningSycophancy

  23. pAI-Econ-claude: A Gated Human-in-the-Loop Multi-Agent Architecture for AI-Assisted Economic Theory Development

    Jul 23, 2026Chen Zhu, Xiaolu Wang, Weilong ZhangEconomiesHuman-In-The-Loop

  24. AI advice suppresses people's willingness to say "I don't know", even when the advice is wrong and accuracy is incentivized

    Jul 15, 2026Chiara Marcoccia, Walter Quattrociocchi, Valerio CapraroAdviceStudent Misconceptions

  25. LLM Judges Can Be Too Generous When There Is No Reference Answer

    Jul 14, 2026Chalamalasetti Kranti, Sowmya VajjalaLarge Language Model JudgesHuman Judgment

  26. Consensus vs. Dissent: Dynamic LLM Modeling of Subjective Preferences in Group Recommenders

    Jul 11, 2026Cedric Waterschoot, Nava Tintarev, Francesco BarilePreference Alignment LearningLarge Language Model Decisions

  27. Beyond Accuracy: How Humans Evaluate Legally Correct but Socially Controversial Legal Advice from Machines

    Jul 6, 2026Benjamin Minhao Chen, Zhiyu LiLegal Reasoning TasksHuman Judgment

  28. Benchmarking the Alignment of Data-Quality Metrics, Human Judgment and Land-Cover Segmentation Performance for Earth Observation

    Jun 23, 2026Ümit Mert Çağlar, Alptekin TemizelSynthetic DataData Quality

  29. AI Alignment From Social Choice Perspectives

    Jun 19, 2026Daniel Halpern, Evi Micha, Ariel D. Procaccia +3Artificial Intelligence AlignmentReinforcement Learning From Human Feedback

  30. MemSlides: A Hierarchical Memory Driven Agent Framework for Personalized Slide Generation with Multi-turn Local Revision

    Jun 15, 2026Ye Jin, Yangyang Xu, Jun Zhu +1Working MemoryPersonalization

  31. LLMs Can Better Capture Human Judgments--With the Right Prompts

    Jun 10, 2026Danica Dillion, Chen Cecilia Liu, Baihui Wang +5Human JudgmentMoral Reasoning

  32. Hidden Consensus:Preference-Validity Compression in Human Feedback

    Jun 9, 2026Dorcas Chia Ern Chua, Karen Myn Hui Lee, Jia Yue Tan +9Preference AlignmentReinforcement Learning From Human Feedback

  33. Personalized to Persuade: The Effects of Contextualization and Warmth on Trust and Reliance in Conversational AI

    May 29, 2026Mert Yazan, Suzan Verberne, Frederik Bungaran Ishak SitumeangPersuasionArtificial Intelligence Literacy

  34. Label Over Logic? How Source Cues Bias Human Fallacy Judgments More Than LLMs

    May 28, 2026Mahjabin Nahar, Nafis Irtiza Tripto, Aiping Xiong +2Logical FallaciesHuman Judgment

  35. Argument Quality Assessment with Large Language Models: A Pairwise Bradley-Terry Approach

    May 27, 2026Nicolás Benjamín Ocampo, Agnes Paullate Nyiranziza, Davide CeolinHuman Judgment

  36. Presupposition and Reasoning in Conditionals: A Theory-Based Study of Humans and LLMs

    May 18, 2026Tara Azin, Yongan Yu, Raj Singh +1Human JudgmentLLM Reasoning Strategies

  37. MHGraphBench: Knowledge Graph-Grounded Benchmarking of Mental Health Knowledge in Large Language Models

    May 15, 2026Weixin Liu, Congning Ni, Shelagh A. Mulvaney +4Mental HealthMedical Ontology

  38. Editor's Choice: Evaluating Abstract Intent in Image Editing through Atomic Entity Analysis

    May 14, 2026Mor Ventura, Roy Hirsch, Yonatan Bitton +2Image EditingMulti-Modal Data

  39. Why Expert Alignment Is Hard: Evidence from Subjective Evaluation

    May 6, 2026Tzu-Mi Lin, Wataru Hirota, Tatsuya Ishigaki +2Expert DemonstrationsLarge Language Model Alignment

  40. Brief chatbot interactions produce lasting changes in human moral values

    Apr 23, 2026Yue Teng, Qianer Zhong, Kim Mai Tich Nguyen Thordsen +2Moral ReasoningChatbots

  41. Not All Subjectivity Is the Same! Defining Desiderata for the Evaluation of Subjectivity in NLP

    Mar 30, 2026Urja Khurana, Michiel van der Meer, Enrico Liscio +2SubjectivityDesiderata

  42. On the Context Sensitivity of LLM Moral Judgment

    Mar 24, 2026Adrian Sauter, Mona SchirmerMoral ReasoningLarge Language Model Decisions

  43. RegCheck: A tool for structured comparisons between study registrations and papers

    Jan 19, 2026Jamie Cummins, Beth Clarke, Ian Hussey +1Research AutomationReproducibility

  44. P-Check: Advancing Personalized Reward Model via Learning to Generate Dynamic Checklist

    Jan 6, 2026Kwangwook Seo, Dongha LeeProgress Reward ModelingReward Functions

  45. Learning to Think Like a Cartoon Captionist: Incongruity-Resolution Supervision for Multimodal Humor Understanding

    Date pendingHatice Merve Vural, Doga Kukul, Ege Erdem Ozlu +4HumorHuman Judgment