Benchmark Datasets

Momentum

11 papers in the last four weeks, up 83% on the four weeks before. 0.1% of all new papers.

Jul 13Week of Sep 28

Latest papers 120

All topics
CardsList
  1. TabJoinBench: A Benchmark for Joinable Table Discovery

    Sep 30, 2026Sandipan De, Jin Wang, Vivek GuptaData LakesBenchmark Datasets

  2. Scores That Hold, Benchmarks That Leak: Measuring Dataset Contamination in Public Brain-Tumor MRI Classification

    Sep 30, 2026Bhanu Prakash Vangala, Sowmya Guda, Latha Peddi +1Brain Tumor ClassificationMedical Imaging Datasets

  3. A 3GPP-Compliant Benchmark Dataset for RIS-Aided Beyond 5G Networks

    Sep 30, 2026Pujitha Mamillapalli, Pankaj Singh Rathour, Abhinav KumarReconfigurable Intelligent SurfacesMillimeter Wave

  4. Large Language Models for Automated Cross-Domain Machine Learning Task Type Identification: A Benchmark Dataset and Evaluation

    Sep 28, 2026Petros Tsialis, Steffen Limmer, Tobias Rodemann +1Large Language Model PipelinesBenchmark Datasets

  5. Can Vision-Language Models Judge Olympic Diving? From Reasoning to Scores in Zero-Shot Action Quality Assessment

    Sep 16, 2026Henry O. Velesaca, David Freire-Obregon, Luigi Miranda +1AthleticismOpen-Vocabulary Action Recognition

  6. ParsHate: A Benchmark Dataset for Hate and Target Detection in Persian

    Sep 14, 2026Zahra Bokaei, Walid Magdy, Bonnie WebberHate Speech DetectionHate Speech

  7. Skeletal Prototypes on Iterative Nerve Expansions

    Sep 14, 2026Jordan Eckert, Henry SchenckLearnable PrototypesSpherical Latent Space

  8. AdamX: Cosine similarity meets gradient descent

    Sep 10, 2026Francisco Caldas, Ruben Belo, Cláudia SoaresAdamGradient Descent

  9. BarkNet-Lite: A Lightweight Texture and Colour Network with the BarkBD Benchmark for Bark-Based Tree Species Recognition in Bangladesh

    Sep 7, 2026Aroshi Ali, Saad Ahmed, Md. Khalid SyfullahSpecies IdentificationBenchmark Datasets

  10. FrameBench:A Language Understanding Benchmark Based on Frame Semantics

    Sep 3, 2026Chihiro Yano, Ryohei SasanoLegal Natural Language Processing BenchmarksBenchmark Datasets

  11. UFPR-PEs: A Brazilian Face Recognition Benchmark with Self-Declared Race/Color Labels

    Aug 31, 2026Alexandre Diano, Bernardo Biesseck, Gabriel Polo +4Face Recognition ModelBenchmark Datasets

  12. VICBench: A Multi-Language Benchmark for Code Vulnerability Detection

    Aug 12, 2026Jin Lu, Xuening Han, Yang Zhong +4Common Vulnerabilities And ExposureBenchmark Datasets

  13. Large-scale AI-Ready Data for Anti-Cancer Drug Response Modeling

    Aug 11, 2026Vincent Lavelle, Yitan Zhu, Kaitlyn Marlor +2PharmacokineticsMulti-Omic Integration

  14. Towards Unified Dynamic Face Landmark Detection

    Aug 11, 2026Sebastian Regalado, Varshanth R. Rao, Ruowei Jiang +2One-Shot Medical Landmark DetectionModality-Specific Encoders

  15. Uncertainty-Aware Ensemble Deep Randomized Neural Networks for Classification

    Aug 7, 2026M. Sajid, A. Quadir, A. Rahaman +2Random ForestUncertainty-Aware Classification

  16. JUMP-lite: Compact, reproducible benchmarking of cell representations

    Aug 7, 2026Alán F. Muñoz, Johan Fredin Haslum, Runxi Shen +2CellsPhenotypes

  17. Towards a satellite image manipulation and deepfake localization benchmark dataset

    Aug 5, 2026Jacob Arndt, Debvrat Varshney, Philipe Dias +1Deepfake DetectionVision Datasets

  18. Why Ranking Anomaly Detection Algorithms Isn't as Reliable as You May Think

    Aug 5, 2026Simon Klüttermann, Jérôme Rutinowski, Frederik Polachowski +1Anomaly ScoreBenchmark Datasets

  19. A Benchmark Dataset for MLLM-Generated Image Detection: GPT Image2 & Nano Banana2

    Aug 2, 2026Zirui Zhang, Yinbo Yu, Donghai Guan +3Ai-Generated Image DetectionImage Generation

  20. ARB: A Matched Authorship-Rewriting Benchmark Dataset for AI-Text Detector Evaluation

    Jul 31, 2026Gaetano Perrone, Simon Pietro RomanoMachine-Generated Text DetectionLarge Language Model Benchmarks

  21. Benchmarks Are Not Monolithic: Sample-Level Auditing and Orchestration for LLM Evaluation

    Jul 30, 2026Philipp D. Siedler, Jordan SassoonBenchmark DatasetsModel Auditing

  22. Enhancing Automated Machine Learning via Homogeneous Train-Test Splitting Methods

    Jul 29, 2026Yearn Tan Yin Tze, Charles GrelloisModel EvaluationBenchmark Datasets

  23. BanClickThumb: A Multimodal Dataset and Transformer Fusion Benchmarks for Clickbait Detection in Bengali YouTube Videos

    Jul 19, 2026Md. Ariful Islam, Md Tanvirul Islam, Md. Maruf Hossain Miru +1BanglaAi-Generated Video Detection

  24. Beyond Benchmarks: Exposing the Hidden Crisis in Bangla Hate Speech Detection

    Jul 13, 2026Faria Afrin Tisha, Fariya Tabassum, Hafsa Binte Kibria +2Hate Speech DetectionBangla

  25. TSAI-MetaFraud: A Benchmark Dataset for Financial Fraud Transaction and Behavioral Risk Detection in Metaverse Ecosystems

    Jul 10, 2026Refat Ishrak Hemel, Ehsan Hallaji, Roozbeh Razavi-FarFraud DetectionBenchmark Datasets

  26. Event Stream based Multi-Modal Video Anomaly Detection: A Benchmark Dataset and Algorithms

    Jul 10, 2026Peipei Zhu, Yueqing Niu, Lin Zhu +3Training-Free Video Anomaly DetectionFine-Grained Video Understanding

  27. Data Analysis in the Wild: Benchmarking Large Language Models Against Real-World Data Complexities

    Jul 7, 2026So Hasegawa, Shailaja Keyur Sampat, Lei Liu +1Large Language Model BenchmarksBenchmark Datasets

  28. BatteryLake: Agentic, Physics-Grounded Curation of Heterogeneous Battery Aging Data and Benchmarking

    Jul 6, 2026Tianwen Zhu, Hao Wang, Yonggang WenLithium-Ion BatteriesBenchmark Datasets

  29. When Do Foundation Models Pay Off? A Break-Even Analysis of Pretrained Time Series Forecasters

    Jul 6, 2026Nicholas Tan Jerome, Frank SimonTime Series Foundation ModelsFoundation Model

  30. A Fair Benchmarking of Deep Relational Database Learning Models

    Jul 4, 2026Kazi F. Akhter, Bharath Ajendla, Manar D. SamadRelational Deep LearningRelational Databases

  31. ClinOCR-Bench: A Comprehensive Clinical Scanned Document Dataset for Optical Character Recognition Model Evaluation

    Jul 4, 2026Enshuo Hsu, Jin Zhou, Kirk RobertsOptical Character RecognitionClinical

  32. Deriving Benchmarking Datasets from Long-Form Recordings: Challenges and Opportunities

    Jul 3, 2026Kaveri K. Sheth, Lawrence Borst, Tarek Kunze +6Language AcquisitionSeed-Tts-Eval Benchmark

  33. AGVBench: A Reliability-Oriented Benchmark of Data Augmentation for Vein Recognition

    Jul 2, 2026Haiyang Li, Yuming Fu, Qun Song +4BiometricsData Augmentation

  34. IonSense-QKG: A Quantum-Readiness Metadata Framework for Lithium-Ion Battery Dataset Discovery

    Jul 1, 2026Sakthi Prabhu Gunasekar, Prasanna Kumar RangarajanLithium-Ion BatteriesBenchmark Datasets

  35. Exploring Differences Between Tabular Enterprise Data and Public Benchmarks

    Jun 29, 2026Myung Jun Kim, Maximilian Schambach, Frank Essenberger +2Tabular DataBenchmark Datasets

  36. Beyond IID: How General Are Tabular Foundation Models, Really?

    Jun 29, 2026Lennart Purucker, Andrej Tschalzev, Nick Erickson +7Tabular Foundation ModelsBenchmark Datasets

  37. SFBench: The SciFy Scientific Feasibility Benchmark

    Jun 28, 2026Cash Costello, James Mayfield, Elsbeth Turcan +7Scientific DiscoveryBenchmark Datasets

  38. KidRisk: Benchmark Dataset for Children Dangerous Action Recognition

    Jun 24, 2026Minh-Kha Nguyen, Trung-Hieu Do, Kim Anh Phung +3Open-Vocabulary Action RecognitionBenchmark Datasets

  39. The Two-Hump Problem: Bridging the Difficulty Gap in Mathematical Reinforcement Learning

    Jun 19, 2026Lucas Fagan, Michele Tarquini, Ali Shehper +6Offline Reinforcement LearningSearch Algorithms

  40. Heterogeneous SAR-optical fusion for near-real-time land use and land cover mapping under cloud contamination: A novel framework and global benchmark dataset

    Jun 16, 2026Jiangong Xu, Weibao Xue, Xiaoyu Yu +3Sentinel-2Land Cover

  41. The BD-LSC Dataset: Facilitating the Benchmarking of Models for Lexical Semantic Change Detection in Slang and Standard Usage

    Jun 15, 2026Afnan Aloraini, Riza Batista-NavarroBenchmark DatasetsChange Detection

  42. Uncertainty Quality of VGGT: An Analysis on the DTU Benchmark Dataset

    Jun 15, 2026Markus Hillemann, Robert Langendörfer, Steven Landgraf +1Visual Geometry Grounded TransformerReconstruction Fidelity

  43. LessonBench-V1: A Benchmark Dataset for Evaluating AI Lesson Generation Agents

    Jun 12, 2026Ravidu Suien Rammuni Silva, Ahmad Lotfi, Isibor Kennedy Ihianle +2TutorsFaithful Question Generation

  44. HybridCodeAuthorship: A Benchmark Dataset for Line-Level Code Authorship Detection

    Jun 10, 2026Luke Patterson, Li Wang, Adam FaulknerAuthorshipCode Generation

  45. An Electric Potential-Augmented Benchmark Dataset for Physics-Guided Image Reconstruction of Electrical Capacitance Tomography

    Jun 10, 2026Xinqi Zhang, Qiming Ma, Lihui PengImage ReconstructionNeural Fields

  46. Benchmarking and Exploring the Capabilities of LLMs for Attack Investigations

    Jun 9, 2026Aniket Anand, Yiwei Hou, Daniel Fields +4Security EvaluationAttacker Large Language Model

  47. POPSICLE: Benchmark Datasets for Segmentation and Localization in CryoET

    Jun 8, 2026Jonathan Schwartz, Utz Heinrich Ermel, C. Braxton Owens +6Cryo-Electron MicroscopyBenchmark Datasets

  48. RIDE: An Open Dataset and Benchmark for Train Delay Prediction

    Jun 3, 2026Clément Elliker, Mathis Le Bail, Clément Mantoux +2RelightingAccurate Traffic Forecasting

  49. New Benchmarking Shows Limited Generalization Power of TCR Antigenic Epitope Prediction Models

    Jun 3, 2026Yiming Liao, Yiheng Li, Ning Jiang +2Epitope PredictionBenchmark Datasets

  50. AICompanionBench: Benchmarking LLMs-as-Judges for AI Companion Safety

    Jun 3, 2026Yanjing Ren, Reza Ebrahimi, TengTeng MaArtificial Intelligence CompanionsLlm-As-A-Judge

  51. The Reliability Gap in Benchmark Auditing: Distribution Shift and Scale as Failure Modes of Contamination Detection

    Jun 2, 2026Wojciech Zarzecki, Jan Dubiński, Sebastian CygertModel AuditingBenchmark Datasets

  52. Benchmark Dataset for Catalysis on 2D MXenes

    May 30, 2026Pavlo Melnyk, Anmar Karmush, Mårten Wadenbäck +4Heterogeneous CatalysisFirst-Principles Calculation