cs.LGOct 5, 2026

CCQ: A Multi-State Child Care Quality Dataset to Support AI for Children's Health Research

Authors: Victor Li, Yuzhang Xie, Ziwei Dong, Qingyang Zhu, Wenjing Ma, Carl Yang, Jinbing Bai, Huiwen Xu, +1 more

Organizations: Emory University, Atlanta, GA, USA

Abstract

High-quality child care in early life is a critical determinant of children's growth and development. Research on child care quality has been constrained by fragmented, non-research-friendly, and privacy-bound datasets. We present CCQ (Child Care Quality), a large-scale, de-identified dataset for applied data science research at the intersection of AI and early childhood health. CCQ integrates 59,372 child care provider records across 12 U.S. states, covering diverse provider types as well as data schemas. To ensure research utility while protecting privacy, we implement an automated, LLM-based curation pipeline that anonymizes, cleans, and standardizes raw state records into two complementary releases: a cleaned textual release and a fully preprocessed tabular release. We also benchmark traditional machine learning models, tabular foundation models, and language models on quality rating prediction and important features analytics. Within a state, tabular classifiers on the preprocessed tables perform best. Across states, zero-shot transfer is near chance, but modest target-state supervision recovers most of the within-state performance, and pretraining on other states benefits finetuned language models. We release both datasets with all code to accelerate AI-driven research on child care quality and ultimately improve children's health and development.

Figures & tables

Explore similar work

CardsList
  1. CAREBench: A Child-Safety Risk Benchmark for Language Models

    Jun 29, 2026Kaavya Krishna-Kumar, Elaine Lau, Vaughn Robinson +6Artificial Intelligence SafetyFrontier Models

  2. Deriving Benchmarking Datasets from Long-Form Recordings: Challenges and Opportunities

    Jul 3, 2026Kaveri K. Sheth, Lawrence Borst, Tarek Kunze +6Language AcquisitionSeed-Tts-Eval Benchmark

  3. ChildEval: When large language models meet children's personalities

    May 27, 2026Yanyan Luo, Xue Han, Chunxu Zhao +5Large Language Model EvaluationPersonality