Human-Annotated Benchmark

Momentum

18 papers in the last four weeks, up 125% on the four weeks before. 0.2% of all new papers.

Jul 13Week of Sep 28

Latest papers 87

All topics
CardsList
  1. PlotGround: Grounding Plot Digitization in Real Scientific Figures and Their Source Data

    Oct 5, 2026Yaohui Zhang, Binxu Li, Haoyi Duan +7Human-Annotated BenchmarkQuantifying

  2. GNN-CB: A Graph Neural Network Competition Benchmark for Human and LLM Evaluation

    Oct 4, 2026Murad Hossen, Tasneem Selim, Gurur Gamgam +22Human-Annotated Benchmark

  3. A Citation-Grounded Benchmark for Trustworthy Earnings Call Transcript Analysis with Large Language Models

    Oct 1, 2026Yingzhu Zhao, Vlad Pandelea, Han Yuan +4Financial Question AnsweringHuman-Annotated Benchmark

  4. R-GroundBench: A Diagnostic Benchmark for R-Group Groundingin Markush Molecular Editing

    Sep 30, 2026Xin Wang, Zichuan Ying, Xinna Lin +8Chemical StructuresMoleculenet

  5. EngramBench: A Capability-Grounded Benchmark for Skill-Evolution Harnesses

    Sep 30, 2026Zhixuan Tan, Pengjie Gu, Zhao Li +4Coding AgentsAgentic Benchmarks

  6. OPIS: An Input-Grounded Benchmark for Multi-Object Memory in Video World Models

    Sep 28, 2026Hao Wang, Tao Yu, Liuzhou Zhang +13Video World ModelsMultimodal Memory

  7. MolDesignBench: Evaluating LLM-based Agent for Scenario-grounded Molecular Design

    Sep 23, 2026Yongjun Jeong, Hanbum Ko, Ye Rin Kim +6Molecular DesignHuman-Annotated Benchmark

  8. GroundedGEO: Auditing the Evidence Gap in Generative Search Rankings

    Sep 21, 2026Yihan Xia, Huiling Fan, Kangrong Zhong +1Evidence-Grounded QuestionsEvidence Conflict

  9. ReSTI: A Source-Grounded Audit and Repair of STI-Bench

    Sep 21, 2026Pengzhan Sun, Ramanathan Rajaraman, Shiu-Hong Kao +2Stable Spatial UnderstandingHuman-Annotated Benchmark

  10. Mind the Gaps: A Curated Benchmark for Form Field Detection

    Sep 20, 2026Iheb Brini, Omar Moured, Hamza Gbada +1Field EvaluationHuman-Annotated Benchmark

  11. Omni2Web: Benchmarking Audiovisual Website Development

    Sep 20, 2026Minghao Han, Zhenghao Xing, Xize Cheng +7OmniHuman-Annotated Benchmark

  12. DocAttriBench: Benchmarking Answer Grounding in Document Visual Question Answering

    Sep 17, 2026Luca De Grandis, Silvia Cappelletti, William Raccagni +3Human-Annotated BenchmarkDocument

  13. Map2Route: Benchmarking Compositional Language-Grounded Route Planning over Semantic Maps

    Sep 15, 2026Muyi Bao, Hang Xu, Jingfan Tang +5Semantic MappingHuman-Annotated Benchmark

  14. MedSNIP: Building and Benchmarking Snippet-Level Granularity for Medical Fact Verification

    Sep 14, 2026Hasan Iqbal, Sarfraz Ahmad, Hyunjae Kim +5FactsHuman-Annotated Benchmark

  15. NovGauge: A Fine-Grained Benchmark for Diagnosing LLMs' Capability in Paper Novelty Assessment

    Sep 12, 2026Guoqiang Zhang, Kexin Tan, Ming Zhang +12Human-Annotated BenchmarkHuman Judgment

  16. ReactHuman: A Physics-Grounded Benchmark for Human-Like Reactive Decision-Making in Embodied Multimodal LLMs

    Sep 12, 2026Yizhan Li, Jianxin You, Mengyang Xiong +5Embodied Artificial IntelligenceHuman-Annotated Benchmark

  17. GPS-Bench: A Governance Policy Benchmark for Automating Policy Analysis

    Sep 3, 2026Linh Le, Melanie Bui, My Chiffon Nguyen +2Policy EvaluationMulti-Agent Simulations

  18. Can LLMs Discover Scientific Laws in Real and Parallel Worlds?

    Sep 1, 2026Yiming Huang, Ziche Liu, Junxia Cui +11Scientific DiscoveryHuman-Annotated Benchmark

  19. SocialReasonBench: A Video-QA Benchmark for Social Reasoning with Counterfactual Narrative Videos

    Aug 31, 2026Zheyu Huang, Zijing Shi, Haozhe Luo +4Social ReasoningReasoning Benchmark

  20. HumanTracker: Towards Comprehensive and Human-Aligned Motion Tracking Benchmark

    Aug 13, 2026Dairu Liu, Zekun Qi, Jiayu Zeng +11Human MotionMotion Imitation

  21. BavGround: A Benchmark for Regional Cultural Grounding and Dialect Competence in Bavarian

    Aug 13, 2026Jophin John, Michael Hoffmann, Jan Fillies +2Multilingual BenchmarkDialects

  22. SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information

    Aug 11, 2026Junjie Ye, Zhuohui Sheng, Shaofan Liu +12Human-Annotated BenchmarkDevice

  23. MRBench: A Comprehensive Benchmark for Human Motion-Text Retrieval

    Aug 8, 2026Fulong Liu, Liang Xu, Chengqun Yang +3Motion-Language AlignmentHuman-Annotated Benchmark

  24. FinRank: An Evidence-Grounded Benchmark for Financial Question Answering and Retrieval over SEC Filings

    Aug 7, 2026Sasan Mansouri, Daniel Saad, Mark Wahrenburg +2Financial Question AnsweringHuman-Annotated Benchmark

  25. Recipes for Creativity: Iterative Generation and Evaluation in Large Language Models

    Aug 7, 2026Rens Anderson, Tessa Verhoef, Amirhossein ZohrehvandCreativityLarge Language Model Evaluation

  26. TRACE: A Multi-Layer Benchmark for Human AI Controller Coordination Under Drift and Failure

    Aug 7, 2026Joshua Zuniga, Srinivasan Subramanian, Ramya Madhuri Narapureddy +1TracesArtificial Intelligence Control

  27. GAUGE: A Measurement-Grounded Benchmark for Physical Fidelity in Simulation Engines and Video World Models

    Aug 6, 2026Shuai Wang, Yaxin Feng, Xuekun Jiang +12Physics SimulationVideo World Models

  28. M3^3R-Bench: A Unified Benchmark for Evidence-Grounded Multimodal Metaphor Understanding

    Aug 6, 2026Hong Jiang, Junnan Zhu, Jingwang Huang +9MetaphorsHuman-Annotated Benchmark

  29. FinReportBench: Measuring and Improving Institution-Grade Financial Report Generation

    Aug 5, 2026Yinghao Tang, Tan Zhenwei, Yiyao Wang +4ReportModel Auditing

  30. COMEX: A Composition-Grounded Benchmark and Learning Framework for Explainable Aesthetic Image Cropping

    Aug 4, 2026Rui Yang, Wei Zhou, Dingyong Gou +5CroppingMulti-Image Editing Benchmarks

  31. CrossLex: A Source-Grounded Benchmark for Cross-Jurisdictional Legal Reasoning in Large Language Models

    Aug 2, 2026Xiaocui Yang, Xican Tan, Shoujie Chen +3Legal Reasoning TasksHuman-Annotated Benchmark

  32. MetaSpace: Metamorphic Testing for Spatial Cognition in Embodied Agents

    Jul 26, 2026Gengyang Xu, Dongwei Xiao, Yiteng Peng +1Embodied AgentsHuman-Annotated Benchmark

  33. GPE: Evaluating Robust Evidence Aggregation for Fact Verification under Controllable GEO-Style Poisoning

    Jul 22, 2026Zhaoqi Wang, Zijian Zhang, Xiaomei Yuan +4Fact-CheckingEvidence-Grounded Questions

  34. Benchmarking Human and Automatic Speech Recognition of Diverse Speech: Initial Results

    Jul 21, 2026Ilse Huisman, Rares Popa, Yuanyuan Zhang +1Automatic Speech Recognition EvaluationAutomatic Speech Recognition

  35. Frontier AI performance across the business disciplines: a case-grounded benchmark of knowledge work and analytical reasoning

    Jul 17, 2026Ajay Patel, Kartik Hosanagar, Ramayya Krishnan +2Artificial Intelligence BenchmarksHuman-Annotated Benchmark

  36. Large Language Models for Code Generation from Multilingual Prompts: A Curated Benchmark and a Study on Code Quality

    Jul 16, 2026Saima Afrin, Alessandro Midolo, Camilo Escobar-Velásquez +5Code GenerationMultilingual Benchmark

  37. Enjoy Your Talk: A Human-Centered Benchmark for Multi-Turn Dialogue with Decoupled User Simulation, Target Modeling, and Judging

    Jul 11, 2026Jinglan Gong, Jiefan Lu, Hewei Guo +5Dialogue BenchmarksUser Simulation

  38. Evaluating Neural Decompilation of Dart AOT Binaries: Fine-Tuning, Metric Validity, Specification Leakage, and Reliability

    Jul 7, 2026Raafat Abualazm, Ayman AboElhassan, Amr G. WassalModel Fine-TuningHuman-Annotated Benchmark

  39. Overview of the TalentCLEF 2026: Skill and Job Title Intelligence for Human Capital Management

    Jun 30, 2026Luis Gasco, Hermenegildo Fabregat, Laura García-Sardiña +8JobsSkills

  40. A Physics-Grounded Benchmark for Multi-Agent Dynamics in World Models

    Jun 27, 2026Nuo Chen, Lulin Liu, Zihao Li +12Physics SimulationCollision Prediction

  41. AGORA: An Archive-Grounded Benchmark for Agentic Workplace Document Reasoning

    Jun 23, 2026Honglin Guo, Qi Zhang, Yu Zhang +7Human-Annotated BenchmarkTool-Grounded Reasoning

  42. Weaving Multi-Source Evidence for Biomedical Reasoning: The BioMedHop Benchmark and BioWeave Framework

    Jun 15, 2026Xingyu Tan, Shiyuan Liu, Xiaoyang Wang +5Multi-Hop ReasoningHuman-Annotated Benchmark

  43. Human-Centered Benchmarking of Driver Monitoring Models

    Jun 6, 2026Ruben Dario Florez-ZelaMobilenetv2Human-Annotated Benchmark

  44. Impostor: An Agent-Curated Benchmark for Realistic AIGC Manipulation Localization

    Jun 3, 2026Zhenliang Li, Yutao Hu, Qixiong Wang +5Human-Annotated BenchmarkLocalization

  45. OVO-S-Bench: A Hierarchical Benchmark for Streaming Spatial Intelligence in Multimodal LLMs

    Jun 2, 2026Yifei Li, Pengyiang Liu, Yuhang Zang +4Stable Spatial UnderstandingMultimodal Large Language Models

  46. InsightVQA: High-Dimensional Emotion-Cognitive Visual Question Answering Benchmark

    Jun 1, 2026Shiyu Wang, Ziyu Liu, Chaoyi Yu +6Emotion RecognitionHuman-Annotated Benchmark

  47. "I've Seen How This Goes": Characterizing Diversity via Progressive Conditional Surprise

    Jun 1, 2026Matthew Khoriaty, David Williams-King, Shi FengDiversityCreativity

  48. Position: Every Ground Truth is a Human Construction, not an Objective Truth

    May 28, 2026Charlotte Högberg, Ericka Johnson, Kiri L. WagstaffGround TruthHuman-Annotated Benchmark

  49. AstroMind: A High-Fidelity Benchmark for Spacecraft Behavior Reasoning Based on Large Language Models

    May 23, 2026Hao Liu, Siyuan Yang, Qinglei Hu +1Human-Annotated BenchmarkHigh-Fidelity