Data Synthesis
Data synthesis focuses on generating artificial datasets that mimic the statistical properties and structure of real-world data, primarily to address data scarcity, privacy concerns, and the need for diverse training data in machine learning. Current research emphasizes the synthesis of complex data types, including relational databases and time series, often employing generative models like diffusion models and large language models (LLMs) to achieve high fidelity and utility. These techniques are proving valuable in various applications, from improving the performance of large language models and vision systems to enhancing medical image analysis and enabling privacy-preserving data sharing. The field is also actively developing robust evaluation metrics and methods to ensure the quality and reliability of synthetic data.
Papers
DS$^2$-ABSA: Dual-Stream Data Synthesis with Label Refinement for Few-Shot Aspect-Based Sentiment Analysis
Hongling Xu, Yice Zhang, Qianlong Wang, Ruifeng Xu
How to Synthesize Text Data without Model Collapse?
Xuekai Zhu, Daixuan Cheng, Hengli Li, Kaiyan Zhang, Ermo Hua, Xingtai Lv, Ning Ding, Zhouhan Lin, Zilong Zheng, Bowen Zhou
MegaPairs: Massive Data Synthesis For Universal Multimodal Retrieval
Junjie Zhou, Zheng Liu, Ze Liu, Shitao Xiao, Yueze Wang, Bo Zhao, Chen Jason Zhang, Defu Lian, Yongping Xiong
A Survey on Data Synthesis and Augmentation for Large Language Models
Ke Wang, Jiahui Zhu, Minjie Ren, Zeming Liu, Shiwei Li, Zongye Zhang, Chenkai Zhang, Xiaoyu Wu, Qiqi Zhan, Qingjie Liu, Yunhong Wang
Robust RL with LLM-Driven Data Synthesis and Policy Adaptation for Autonomous Driving
Sihao Wu, Jiaxu Liu, Xiangyu Yin, Guangliang Cheng, Meng Fang, Xingyu Zhao, Xinping Yi, Xiaowei Huang