Synthetic Benchmark Generation

Latest papers 213

All topics
CardsList
  1. HawkesNest: A Multi-Axis Synthetic Benchmark for Spatiotemporal Pattern Complexity

    Jun 15, 2026Yahya Aalaila, Sumantrak Mukherjee, Gerrit Großmann +1Temporal Point ProcessesSynthetic Benchmark Generation

  2. A Validated LBM Dataset and Pipeline for Surrogate Modeling of Turbulent 3D Obstructed Channel Flows

    Jun 15, 2026Lukas Schröder, Shubham Kavane, Harald KöstlerNeural Surrogate ModelingSynthetic Benchmark Generation

  3. PAL-Bench: Evidence-Grounded Profile Reconstruction from Longitudinal Personal Albums

    Jun 15, 2026Qiwei Yan, Zhiqiang Yuan, Zexi Jia +4Data ProvenanceSynthetic Benchmark Generation

  4. The Data Manifold under the Microscope

    Jun 14, 2026Marios Koulakis, Constantin SeiboldIntrinsic DimensionalitySynthetic Benchmark Generation

  5. Benchmarking Instance-Dependent Label Noise with Controlled Corruptions

    Jun 12, 2026Shadman Islam, Agustinus Kristiadi, Mostafa MilaniSynthetic Benchmark GenerationNoisy-Label Learning

  6. VHDLSuite: Unified Pipeline for LLM VHDL Generation with Data Synthesis and Evaluation

    Jun 11, 2026Yijun Shen, Minghao Shao, Yichen Zhao +4LLM EvaluationCode Generation

  7. EvoBrowseComp: Benchmarking Search Agents on Evolving Knowledge

    Jun 11, 2026Yunhan Wang, Jiaan Wang, Lianzhe Huang +2Benchmark ContaminationWeb Search Agents

  8. SciR: A Controllable Benchmark for Scientific Reasoning in LLMs

    Jun 11, 2026Pierre Beckmann, Marco Valentino, Andre FreitasLLM EvaluationScientific Reasoning

  9. How Fine-Grained Should a RAG Benchmark Be? A Hierarchical Framework for Synthetic Question Generation

    Jun 11, 2026Chase M. Fensore, Kaustubh Dhole, Jason Fan +2Benchmark DesignSynthetic Data Generation

  10. Net-Ev2^2: A Generative Simulator for Network Event Evolution

    Jun 10, 2026Guangyu Wang, Zhaonan WangIntelligent Transportation SystemsSynthetic Benchmark Generation

  11. Embodied-BenchClaw: An Autonomous Multi-Agent System for Embodied Spatial Intelligence Benchmark Construction

    Jun 10, 2026Baoyang Jiang, Fengchun Zhang, Leyuan Wang +7Spatial Reasoning BenchmarksSynthetic Benchmark Generation

  12. ClusBench: The Clustering Benchmark Data Resource You've All Been Waiting For (?)

    Jun 9, 2026David P. HofmeyrBenchmark ConstructionClustering

  13. STAGE-Claw: Automated State-based Agent Benchmarking for Realistic Scenarios

    Jun 9, 2026Sirui Liang, Bohan Yu, Peiyu Wang +8Computer-Use Agent BenchmarksLLM Agent Evaluation

  14. BSTabDiff: Block-Subunit Diffusion Priors for High-Dimensional Tabular Data Generation

    Jun 8, 2026Al Zadid Sultan Bin Habib, Md Younus Ahamed, Prashnna Gyawali +2Copula ModelsSynthetic Tabular Data Generation

  15. PIPE-Cypher: Automatic Enterprise Benchmark Generation for Text-to-Cypher Systems

    Jun 7, 2026Suraj Ranganath, Anish RaghavendraBenchmark DesignText-to-Cypher

  16. SurgiQ: A Large-Scale Multi-Domain Benchmark for Evaluating Surgical Understanding in Large Language Models

    Jun 6, 2026Ayah Al-Naji, Edoardo Fazzari, Saif Alkindi +3LLM EvaluationSynthetic Benchmark Generation

  17. The CIFAR Synthetic Evidence Corpus for Detecting AI-Generated Evidence

    Jun 6, 2026Kelly McConvey, Jalehsadat Mahdavimoghaddam, Nima Jamali +8Digital ForensicsSynthetic Benchmark Generation

  18. A Framework for Evaluating and Benchmarking Concept Drift Detection Methods

    Jun 5, 2026Vitor Cerqueira, Heitor Murilo Gomes, Marco Heyden +2Concept DriftSynthetic Benchmark Generation

  19. UnpredictaBench: A Benchmark for Evaluating Distributional Randomness in LLMs

    Jun 4, 2026Amirhossein Abaskohi, Amirhossein Dabiriaghdam, Liang Luo +4LLM EvaluationLanguage Model Generation Evaluation

  20. Benchmark Everything Everywhere All at Once

    Jun 4, 2026Shiyun Xiong, Dongming Wu, Peiwen Sun +5LLM EvaluationBenchmark Design

  21. PhyRoGen: Synthetic Generation of Physical Robot Manipulation Puzzles Using Procedural Content Generation

    Jun 4, 2026Lennart Julian Droß, Andreas Orthey, Marc ToussaintSynthetic Benchmark GenerationRobotic Manipulation

  22. GenTI: Benchmarking LLMs for Autonomous IDPS Rule Generation for Unseen Attacks

    Jun 4, 2026Hassan Jalil Hadi, Rehana Yasmin, Ali ShokerNetwork Intrusion DetectionSynthetic Benchmark Generation

  23. SANE Schema-aware Natural-language Evaluation of Biological Data

    Jun 3, 2026Rolf Gattung, Martin Krueger, Markus ReischlLLM EvaluationLLM Grounding

  24. Measuring What Matters: Synthetic Benchmarks for Concept Bottleneck Models

    Jun 3, 2026Julian Skirzynski, Harry Cheon, Shreyas Kadekodi +2Concept Bottleneck ModelsBenchmark Design

  25. CoEval: Ranking Language Models for Custom Tasks Without Labeled Data or Trustworthy Benchmarks

    Jun 2, 2026Alexander Apartsin, Yehudit ApersteinModel SelectionLLM Evaluation

  26. FinStressTS: A Parametric Synthetic Benchmark for Time-Series Forecasting in Finance

    Jun 2, 2026Jiaze Sun, Kelvin J. L. Koa, Ruiyang Ni +3Financial ForecastingTime Series Forecasting