High-quality child care in early life is a critical determinant of children's growth and development. Research on child care quality has been constrained by fragmented, non-research-friendly, and privacy-bound datasets. We present CCQ (Child Care Quality), a large-scale, de-identified dataset for applied data science research at the intersection of AI and early childhood health. CCQ integrates 59,372 child care provider records across 12 U.S. states, covering diverse provider types as well as data schemas. To ensure research utility while protecting privacy, we implement an automated, LLM-based curation pipeline that anonymizes, cleans, and standardizes raw state records into two complementary releases: a cleaned textual release and a fully preprocessed tabular release. We also benchmark traditional machine learning models, tabular foundation models, and language models on quality rating prediction and important features analytics. Within a state, tabular classifiers on the preprocessed tables perform best. Across states, zero-shot transfer is near chance, but modest target-state supervision recovers most of the within-state performance, and pretraining on other states benefits finetuned language models. We release both datasets with all code to accelerate AI-driven research on child care quality and ultimately improve children's health and development.
Figures & tables
Figure 1 . Overview of the CCQ study. Stage 1 curates the QRIS records of twelve states in four steps: collection, anonymization, cleaning, and verification (Figure 2 ). The raw release retains textual feature while the preprocessed release is strictly numerical. Stage 2 benchmarks tabular classifiers and language models on two tasks. Within-state prediction uses five-fold cross-validation and cross-state prediction holds out one state (orange) and trains on the others (blue). Because the preprocessed data schemas differ across states, it is only used for within-state prediction. The trained models are analyzed for prediction accuracy, feature attribution, and cross-state transferability.
State
Data Source
#Provider
#Rated
QRIS
QR Scale
#Feature
CA
CA Child Care Resource & Referral Network ( MyChildCarePlan, 2026 )
10,715
847
Quality Counts California
1–5
142
CO
CO Dept of Early Childhood ( Colorado Dept of Early Childhood, 2026 )
4,508
3,423
Colorado Shines Rating
1–5
64
GA
GA Dept of Early Care and Learning ( Georgia Dept of Early Care and Learning, 2026 )
8,192
2,897
Quality Rated
1–3
203
KY
KY Cabinet for Health and Family Services ( Kentucky Cabinet for Health and Family Services, 2026 )
1,929
1,896
Kentucky All STARS
1–5
61
MD
MD State Dept of Education ( Maryland State Department of Education, 2026 )
7,106
5,005
Maryland EXCELS
1–5
32
MT
MT Dept of Public Health & Human Services ( Montana Department of Public Health and Human Services, 2026 )
213
181
Best Beginnings STARS to Quality
1–5
12
Table 1 . Data sources overview. #Feature includes the provider ID and QR score.
Figure 2 . The CCQ construction pipeline. (1) The Georgia reference is implemented by hand. Its human-curated data are the ground truth for testing two agents on twenty curation tasks: a per-sample agent (E2E) and a code-writing agent (Coder), which is adopted. (2) The other eleven states are collected, anonymized, and cleaned with the help of LLMs. Human verification at each step is shown with icons indicating who performed each step. The orange code icon indicate when Georgia’s code is used as and in-context example. Lastly, two annotators spot-checked the curated data against the live portals, finding 97.2% of the checked values correct.
Curation Task
Model
#Tasks
Acc%
In Tok.
Out Tok.
Anonymization
Column removal
E2E
1
97.12
71.0
1.0
Column removal
Coder
1
99.28
103.1
1.0
De-identification
E2E
3
35.2±42.8
107.2±22.6
4.4±2.5
De-identification
Coder
3
100.0±0.0
0.03±0.00
0.05±0.01
Cleaning
Table 2 . LLM-based curation results on the reference state Georgia. Token counts are per provider record, so the code-writing agent’s one-time cost is amortized over the 8,192 records in the Georgia reference implementation.
State
QR score
Provider type
Licensed capacity
Ages served
H1
H2
κ
H1
H2
κ
H1
H2
κ
H1
H2
κ
CA
100.0
100.0
100.0
NA
NA
NA
100.0
100.0
100.0
95.0
95.0
100.0
CO
100.0
100.0
100.0
80.0
80.0
91.3
90.0
95.0
94.6
100.0
100.0
100.0
KY
90.0
95.0
91.6
100.0
100.0
100.0
70.0
95.0
68.0
NA
NA
NA
MD
100.0
100.0
100.0
100.0
100.0
100.0
95.0
95.0
100.0
100.0
100.0
100.0
NC
100.0
95.0
92.4
100.0
100.0
100.0
100.0
100.0
100.0
100.0
100.0
100.0
Table 3. Manual data verification results per state (20 providers per state). H1 and H2 give the percentage of released values each annotator judged correct against the live portal. κ is Cohen’s kappa ( ×100 ) between their correct/incorrect judgments. The ALL row pools the states that publish each field. “NA” indicates that the state does not publish the field.
Figure 3 . Overview of the CCQ dataset. (a) Distribution of quality rating scores across rated providers in each state. (b) Feature missing rates for each state’s raw release, excluding the provider identifier and rating column. Boxes span the interquartile range and whiskers mark the 5th and 95th percentiles.
5-fold stratified, shuffled; folds cached per state and shared by all methods
Target test block
stratified 20%, carved once per target, scored at every supervision level
Validation split
stratified 15% of each training phase; best epoch by validation QWK, patience 2 (textual classifiers; TabNN below)
Tabular classifiers
Table 4 . Complete hyperparameter configuration. Values are held fixed across all 12 states, both tasks, and all six supervision levels. “Phase” is either (1) source pretraining or (2) target adaptation in sequential training. Each phase derives its own class weights and validation split.
Figure 4 . Text-based representation of a provider record, following PedCA-FT ( Lu et al., 2025 ) . (1) A simplified record in Georgia’s schema is written as text. False and blank values are left out, and the provider identifier and the rating never appear. (2) The textual classifiers read this text directly. For the tabular classifiers, all-MiniLM-L6-v2 embeds the text in 254-token chunks, and the normalized mean of the chunk embeddings is the 384-dimensional input.
Method
CA
CO
GA
KY
MD
MT
NC
NE
OK
SC
WA
WI
BA
QWK
BA
QWK
BA
QWK
BA
QWK
BA
QWK
BA
QWK
BA
QWK
BA
QWK
BA
QWK
BA
QWK
BA
QWK
BA
QWK
Dummy
20.00
0.00
20.00
0.00
33.33
0.00
20.00
0.00
20.00
0.00
20.00
0.00
20.00
0.00
20.00
0.00
20.00
0.00
20.00
0.00
25.00
0.00
23.00
0.00
Preprocessed feature representation
LR-num
39.45
29.33
41.37
51.69
45.23
21.04
41.51
32.43
44.75
64.31
37.86
27.23
50.29
44.90
42.44
29.03
49.48
64.64
41.68
36.25
43.51
27.09
44.71
50.59
RF-num
34.01
44.82
37.06
56.82
35.47
6.45
35.95
40.38
43.23
65.07
42.67
31.67
42.37
59.19
39.27
26.79
52.83
76.43
38.76
39.65
32.81
20.44
38.52
52.75
XGB-num
36.22
42.59
39.53
57.98
41.15
19.00
41.76
48.10
44.27
65.23
43.85
26.63
46.02
61.73
38.84
29.99
53.07
76.44
39.66
42.18
39.07
33.63
40.80
52.63
Table 5 . Within-state QR prediction results on each state’s native QR scale. Each cell is the mean over five-fold cross-validation and reports BA and QWK, both as percentages. The best value per state and metric is in bold , and the second best is underlined .
Figure 5 . Elo ratings from the within-state and cross-state experiments. The Elo scale uses a 400-point gap, which corresponds to a 91% expected win rate. Error bars show 95% confidence intervals from 2,000 bootstrap rounds that resample the twelve states. Cross-state scores are averaged over the six supervision levels. RF-txt is the 1,000-point anchor in both panels. The two panels are computed on different sets of comparisons, so scores are comparable within a panel but not across panels.
Figure 6 . Pareto frontier of mean Elo score versus number of tuned parameters for within-state QR prediction. The dashed line marks the Pareto-optimal set. Parameter counts are averaged over the 12 per-state models. For tree ensembles we count leaves. TabPFN updates no weights and is plotted at 100 .
State
QWK (%)
#Features
XGB-num
XGB-reg
Δ
Above floor
Total
CA
42.59
45.85
+ 3.25
4
82
CO
57.98
58.95
+ 0.97
7
117
GA
19.00
18.26
− 0.74
5
278
KY
48.10
49.65
+ 1.55
9
112
MD
65.23
67.95
+ 2.72
3
34
Table 6 . Fidelity of the XGBoost regression model (XGB-reg) used for SHAP analysis. Δ is its QWK difference from the XGBoost classifier (XGB-num). Above floor counts the features that exceed the shadow-feature noise floor, and Total counts the features of the preprocessed representation.
Figure 7 . Ten highest-attribution features per state, ranked by mean ∣SHAP∣ over that state’s providers. The dashed line is the shadow-feature noise floor, and bars to its left are indistinguishable from shuffled noise. Color gives the direction: green where higher feature values push the predicted rating up, red where they push it down. Feature names are state-specific and horizontal scales differ across panels, so bar lengths are comparable within a panel but not across panels.
Method
CA
CO
GA
KY
MD
MT
NC
NE
OK
SC
WA
WI
BA
QWK
BA
QWK
BA
QWK
BA
QWK
BA
QWK
BA
QWK
BA
QWK
BA
QWK
BA
QWK
BA
QWK
BA
QWK
BA
QWK
0% target supervision
Dummy
20.00
0.00
20.00
0.00
33.33
0.00
20.00
0.00
20.00
0.00
20.00
0.00
20.00
0.00
20.00
0.00
20.00
0.00
20.00
0.00
0.00
0.00
0.00
0.00
LR-txt
14.77
-3.88
22.38
17.31
32.00
2.06
20.48
5.80
23.04
35.10
24.71
5.19
17.78
-5.89
18.35
11.36
20.51
1.06
26.79
21.66
27.12
1.01
25.92
7.60
RF-txt
13.53
-9.36
17.58
2.62
32.92
-0.55
14.51
2.41
27.73
48.96
19.00
-8.22
23.26
-0.57
19.66
4.57
16.61
-9.08
24.99
20.88
22.41
3.87
23.34
12.52
XGB-txt
41.28
-1.22
18.67
-0.03
29.30
-6.32
17.74
2.75
23.35
40.56
29.36
-1.44
36.48
23.06
18.21
3.28
14.57
-7.02
20.12
-0.70
21.50
4.65
27.86
3.43
Table 7 . Leave-one-state-out cross-state QR prediction results. At every target supervision level p , each cell is scored on the same stratified 20% test set of the held-out target state. The best value per state and metric within each supervision level, excluding Dummy, is in bold , and the second best is underlined . Dummy predicts the majority class of its training set (the pooled source states plus the target draw).
Figure 8 . (a) Mean test QWK over the twelve target states as a function of the target supervision level p for all ten models. (b) Mean test QWK of the same ten models over the twelve states, with and without pretraining on the pooled source states. “w/ pretraining” is the p=80 cross-state setting. “w/o pretraining” is the within-state setting.
How can we evaluate whether frontier AI systems recognize child-safety risks before they escalate into explicit harm? Existing child safety evaluations focus on child sexual abuse material, yet many child-safety failures begin earlier: in model assistance that helps adults manipulate, impersonate, profile, or isolate minors, and in model responses that deepen children's emotional dependence on AI systems rather than redirecting them toward human support. We introduce CAREBench (Child AI Risk Evaluation), a benchmark to assess such upstream child-safety risks in language models. CAREBench contains 500 prompts spanning twelve risk categories, including grooming and relationship engineering, deception and impersonation, surveillance and privacy, sextortion and sexual abuse, AI anthropomorphization, emotional dependency, and mental illness sensitivity. Developed with response annotations from parents and clinicians, the benchmark excludes explicit abuse material and imagery; instead, it evaluates whether models recognize, refuse, de-escalate, or redirect risky interactions before harm becomes overt. Evaluating seven frontier models on our benchmark, we find failure rates ranging from 2% to 58%, with failure patterns that vary across risk categories. CAREBench provides a responsibly scoped evaluation for LLM developers to identify and close gaps in child safety policies.
Long-form recordings (LFRs) of child-centered audio are ecologically valid sources for studying early language development, but three problems limit their use. First, LFR corpora are collected across sites with heterogeneous formats and consent structures, making cross-corpus use non-trivial. Second, without standardized benchmarks, assessing whether tools generalize across languages and conditions is hard. Third, ML workflows rarely respect privacy constraints governing sensitive child speech. This paper presents a framework addressing all three: a standardized collection of 27 child-centered datasets built with open-source tools (S1); a replicable pipeline for four speech-processing benchmarks (S2); and ELSI, a role-based ecosystem embedding ethical governance into the ML workflow (S3). We demonstrate the framework via a voice type classification case study and show the three solutions are mutually dependent.
Kaveri K. Sheth, Lawrence Borst, Tarek Kunze +6
LAAC, LSCP, DEC, ENS, EHESS, CNRS, PSL University, Paris, France · Laboratoire d’Informatique et Syst`emes, Universit´e Aix-Marseille, CNRS, France · Signal Processing Research Centre, Tampere University, Finland
While LLMs enable personalized chatbots, their effectiveness in child-centered personalization remains unclear, as systematic evaluation of child-specific preferences is still lacking. To address this gap, we introduce ChildEval, a benchmark for evaluating LLMs' ability to infer and follow child-centered preferences in long-context conversations. ChildEval contains 29K synthesized persona profiles of children aged 3-6, providing relatively static background information. Each persona is associated with a child preference-which may align with, conflict with, or be independent of the persona-expressed either explicitly in a single sentence or implicitly through 6-10 turn dialogues. Explicit and implicit preferences are designed to reflect the same underlying preference but differ in expression, capturing dynamic aspects of preference expression rather than changes in the static persona. The benchmark spans five top-level and fourteen sub-level categories covering children's daily lives and development. We further propose fine-grained, child-centric evaluation protocols to systematically assess open-source LLMs. Experimental results demonstrate how different personalized representations affect LLM responses and suggest that finetuning on ChildEval can enhance child-centered performance. Our code and dataset are available at https://github.com/ziyanluo/ChildEval.