Data Set
Datasets are crucial for training and evaluating machine learning models, particularly in areas like natural language processing, computer vision, and audio analysis. Current research emphasizes creating diverse and high-quality datasets addressing specific challenges, such as data imbalance, cross-lingual inconsistencies, and the need for realistic representations of real-world scenarios. This involves developing novel annotation techniques, incorporating multiple data modalities (e.g., text, images, audio), and employing various model architectures (e.g., transformers, convolutional neural networks) for analysis and benchmark creation. The availability of well-designed datasets directly impacts the development of robust and reliable machine learning models, ultimately advancing scientific understanding and improving practical applications across numerous fields.
Papers
MedGPTEval: A Dataset and Benchmark to Evaluate Responses of Large Language Models in Medicine
Jie Xu, Lu Lu, Sen Yang, Bilin Liang, Xinwei Peng, Jiali Pang, Jinru Ding, Xiaoming Shi, Lingrui Yang, Huan Song, Kang Li, Xin Sun, Shaoting Zhang
Open-WikiTable: Dataset for Open Domain Question Answering with Complex Reasoning over Table
Sunjun Kweon, Yeonsu Kwon, Seonhee Cho, Yohan Jo, Edward Choi
Fashionpedia-Taste: A Dataset towards Explaining Human Fashion Taste
Mengyun Shi, Serge Belongie, Claire Cardie
A Survey on Dataset Distillation: Approaches, Applications and Future Directions
Jiahui Geng, Zongxiong Chen, Yuandou Wang, Herbert Woisetschlaeger, Sonja Schimmler, Ruben Mayer, Zhiming Zhao, Chunming Rong
NorQuAD: Norwegian Question Answering Dataset
Sardana Ivanova, Fredrik Aas Andreassen, Matias Jentoft, Sondre Wold, Lilja Øvrelid
Fine Tuning with Abnormal Examples
Will Rieger
HeySQuAD: A Spoken Question Answering Dataset
Yijing Wu, SaiKrishna Rallabandi, Ravisutha Srinivasamurthy, Parag Pravin Dakle, Alolika Gon, Preethi Raghavan
DiffuseExpand: Expanding dataset for 2D medical image segmentation using diffusion models
Shitong Shao, Xiaohan Yuan, Zhen Huang, Ziming Qiu, Shuai Wang, Kevin Zhou
LoRaWAN-enabled Smart Campus: The Dataset and a People Counter Use Case
Eslam Eldeeb, Hirley Alves
ZRG: A Dataset for Multimodal 3D Residential Rooftop Understanding
Isaac Corley, Jonathan Lwowski, Peyman Najafirad