Synthetic Benchmark Generation

Latest papers 213

All topics
CardsList
  1. WorldCoder-Bench: Benchmarking Physically Grounded 3D World Synthesis

    Jun 1, 2026Shuo Lu, Yinuo Xu, Kecheng Yu +8LLM-Based Program SynthesisSynthetic Benchmark Generation

  2. ContinuousBench: Can Differentially Private Synthetic Text Improve Capabilities?

    Jun 1, 2026Peihan Liu, Lucas Rosenblatt, Weiwei Kong +7Privacy-Preserving Language ModelsSynthetic Benchmark Generation

  3. BenchEvolver: Frontier Task Synthesis via Solution-Centric Evolution

    May 31, 2026Yangzhen Wu, Aaron J. Li, Wenjie Ma +10Synthetic Benchmark GenerationCode Generation Evaluation

  4. GraphARC: A Comprehensive Benchmark for Graph-Based Abstract Reasoning

    May 29, 2026Saku Peltonen, August Bøgh Rønberg, Andreas Plesner +1LLM Reasoning with GraphsSynthetic Benchmark Generation

  5. XLGoBench: Detecting cross-lingual skill gaps with algorithmic tasks

    May 29, 2026Purvam Jain, Preethi Jyothi, Vihari Piratla +1Multilingual Language Model EvaluationCross-Lingual Reasoning in Language Models

  6. PhoneWorld: From Real-App Trajectories to Dynamic and Verifiable Environments for Phone-Use Agents

    May 28, 2026Yuxuan Liu, Xin Lai, Junyi Li +21Mobile GUI AutomationSynthetic Benchmark Generation

  7. A Matter of TASTE: Improving Coverage and Difficulty of Agent Benchmarks

    May 27, 2026Tomer Keren, Nitay Calderon, Asaf Yehudai +3Computer-Use Agent BenchmarksTool-Use Evaluation

  8. SNARE: Adaptive Scenario Synthesis for Eliciting Overeager Behavior in Coding Agents

    May 27, 2026Yubin Qu, Yi Liu, Gelei Deng +4Coding AgentsAI Agent Security

  9. ReverseMath: Answer Inversion for Scalable and Verifiable Mathematical Problem Generation

    May 26, 2026Raoyuan Zhao, Yihong Liu, Yupei Du +2Mathematical Reasoning BenchmarksSynthetic Benchmark Generation

  10. OSMa-Bench++: Toward Open-Ended Benchmarking of Semantic Mapping for Manipulation with Prompt-Generated Synthetic Scenes

    May 26, 2026Regina Kurkova, Maxim Popov, Sergey KolyubinRobotics SimulationSynthetic Benchmark Generation

  11. A Controlled Synthetic Benchmark for Educational Aspect-Based Sentiment Analysis

    May 25, 2026Yehudit Aperstein, Alexander ApartsinEducational TechnologySynthetic Benchmark Generation

  12. PDEInvBench: A Comprehensive Dataset and Design Space Exploration of Neural Networks for PDE Inverse Problems

    May 25, 2026Divyam Goel, Nithin Chalapathi, Sanjeev Raja +1Neural Network OptimizationSynthetic Benchmark Generation

  13. ScaleWoB: Guiding GUI Agents with Coding Agents via Large-Scale Environmental Synthesis

    May 24, 2026Guohong Liu, Jialei Ye, Pengzhi Gao +4Mobile GUI AutomationGUI Agents

  14. GENSTRAT: Toward a Science of Strategic Reasoning in Large Language Models

    May 22, 2026Vartan Shadarevian, Kia Ghods, Alex Kenich +1LLM EvaluationImperfect-Information Games

  15. VGenST-Bench: A Benchmark for Spatio-Temporal Reasoning via Active Video Synthesis

    May 21, 2026Jinho Park, Youbin Kim, Hogun Park +1VLM EvaluationSpatiotemporal Reasoning

  16. SynAE: A Framework for Measuring the Quality of Synthetic Data for Tool-Calling Agent Evaluations

    May 21, 2026Shuaiqi Wang, Aadyaa Maddi, Zinan Lin +1Tool-Use EvaluationLLM Agent Evaluation

  17. VeriScale: Adversarial Test-Suite Scaling for Verifiable Code Generation

    May 21, 2026Yifan Bai, Xiaoyang Liu, Zihao Mou +7Code GenerationSynthetic Benchmark Generation

  18. SWE-Mutation: Can LLMs Generate Reliable Test Suites in Software Engineering?

    May 21, 2026Yuxuan Sun, Yuze Zhao, Yufeng Wang +6LLM EvaluationAutomated Software Testing

  19. RankJudge: A Multi-Turn LLM-as-a-Judge Synthetic Benchmark Generator

    May 20, 2026Zhenwei Tang, Zhaoyan Liu, Rasa Hosseinzadeh +3LLM EvaluationPairwise Comparison

  20. PlanningBench: Generating Scalable and Verifiable Planning Data for Evaluating and Training Large Language Models

    May 20, 2026Ziliang Zhao, Zenan Xu, Shuting Wang +7LLM EvaluationLLM Planning

  21. MTR-Suite: A Framework for Evaluating and Synthesizing Conversational Retrieval Benchmarks

    May 20, 2026Junhao Ruan, Abudukeyumu Abudula, Bei Li +8Conversational SearchSynthetic Benchmark Generation

  22. ShadeBench: A Benchmark Dataset for Building Shade Simulation in Sustainable Society

    May 19, 2026Longchao Da, Mithun Shivakoti, Xiangrui Liu +3Benchmark ConstructionSynthetic Benchmark Generation