Synthetic Benchmark Generation

Latest papers 213

All topics
CardsList
  1. DrugTargetWorld: A Synthetic Biobank for Training and Benchmarking AI Scientists

    Oct 7, 2026Samuel Margolis, Paul Schmiedmayer, Alan Huang +12Drug DiscoveryAI Agent Benchmarks

  2. RSI-Forge: From Research Papers to Environments for Recursive Self-Improvement

    Oct 7, 2026Renxiong Wang, Darvin Yi, Abril Herrlein +16Self-Improving AgentsRecursive Self-Improvement

  3. MS-Exam-Gen: Source-Grounded Benchmark Construction for Evaluating LLMs on Textual Multiple Sclerosis MRI Knowledge

    Oct 5, 2026Abdul Basit, Muhammad Abdullah Hanif, Muhammad ShafiqueLLM EvaluationSynthetic Benchmark Generation

  4. Ontology-Grounded, Reasoner-Verified Benchmarks for Evaluating LLM Reasoning in Scientific AI

    Sep 30, 2026Nishtha N. Vaidya, Stephan Grimm, Thomas Hubauer +1LLM EvaluationSynthetic Benchmark Generation

  5. Teaching LLMs to Generate Challenging MILP Instances via Solver Feedback

    Sep 29, 2026Jitin Singla, Parikshit Pareek, Pratik Jawanpuria +1Mixed-Integer Linear ProgrammingLarge Language Model-Guided Optimization

  6. MetaBench-Harness: Unlocking End-to-End Optimization of Benchmark Harnesses

    Sep 27, 2026Xuanjun Chen, Hua-Hsuan Chen, Wei-Chung Lu +3LLM EvaluationBenchmark Design

  7. A Living Benchmark for Information Retrieval from Electronic Health Records

    Sep 24, 2026Jordan L. Cahoon, Chloe O. Stanwyck, Sulaiman Somani +23Synthetic Benchmark GenerationClinical QA

  8. SciWalker: Synthesizing Scientific Coding Problems with Operator Graphs and Execution Feedback

    Sep 24, 2026Chenxi Li, Wenxuan Zeng, Yun Luo +4Execution-Guided Code GenerationScientific Code Generation

  9. SSP-Bench: A Hybrid Data Generation Framework for Safety, Security, and Privacy Evaluation

    Sep 21, 2026Fatih Deniz, Yazan Boshmaf, Issa KhalilLLM Safety BenchmarksLLM Evaluation

  10. A paired synthetic construction-site image dataset for robust computer vision under adverse conditions

    Sep 21, 2026Viet Huy Duong, Ruoxin Xiong, Md Abdullah Al Forhad +1Synthetic Data GenerationSynthetic Benchmark Generation

  11. AURORA: A Natural Language-Driven Agentic Framework for Understanding, Reasoning, and Orchestrating Reliable Air-Ground Co-Simulation

    Sep 17, 2026Keshu Wu, Hao Zhang, Rui Gan +4LLM Agent OrchestrationSynthetic Benchmark Generation

  12. WinSyn: An Automated Pipeline for Realistic Enterprise Question-Answering Evaluation

    Sep 14, 2026Amey Varhade, Ananya Sutradhar, Ravishankar Krishnaswamy +1Synthetic Data GenerationAI Agent Evaluation

  13. Embodied-BenchForge: A Closed-Loop Agentic Workflow for Embodied Benchmark Construction

    Sep 14, 2026Baoyang Jiang, Fengchun Zhang, Leyuan Wang +9Benchmark ConstructionEmbodied QA

  14. SynCo: Synthetic Community-Aware Attributed Graph Generator for Graph Neural Network Benchmarking

    Sep 9, 2026Guilherme Henrique Messias, Mariana Caravanti de Souza, Sylvia Iasulaitis +1Synthetic Data AugmentationGraph Generation

  15. FinCUABuild: Can Agents Build Reliable Benchmarks for Dynamic Financial Computer Use?

    Sep 7, 2026Jingpu Yang, Fengxian Ji, Jinri Guo +7Computer-Use Agent BenchmarksFinancial Services

  16. SHELF: A Synthetic Harness for Multi-Task Bibliographic Benchmarking

    Sep 2, 2026Michael J. BommaritoLLM EvaluationLibrary and Information Science

  17. Peg-in-Bench: A Modular Benchmark for High-Precision Robotic Insertion

    Sep 1, 2026Yosel Delgado, José G. Buenaventura-Carreón, Floris Erich +4Contact-Rich Robotic ManipulationGeneralization in Robotic Manipulation

  18. Synthetic Worlds for Temporal Evaluation and Knowledge Updating in LLMs

    Aug 31, 2026Jonathan Zheng, Zirui Shao, Alan Ritter +1Synthetic Data PretrainingKnowledge Editing