CLARA: Can AI Assess Developmental Appropriateness in Children's Stories?
Organizations: University of Auckland
Abstract
Assessing the developmental suitability of children's narratives is important for educational recommendation and developmental literacy research, yet such assessment typically relies on subjective and difficult-to-scale human judgment. This raises an important question: Can AI systems approximate human developmental judgments of children's stories? To study this problem, we introduce CLARA, a cognitively grounded framework for developmental narrative understanding through structured annotation across cognitive (COG), language (LAN), and social-emotional (SEL) dimensions, together with a bilingual benchmark resource containing 1107 Chinese--English children's stories with normalized silver developmental references and structured developmental annotations. We evaluate CLARA through benchmark comparison, component analysis, translated bilingual consistency analysis, and blinded human evaluation with educators. Experimental results show that structured developmental annotation achieves substantially stronger alignment with developmental references and human judgments than readability-based methods and direct prompting baselines. Overall, our findings suggest that AI systems can approximate certain aspects of human developmental judgment when guided by structured developmental annotation, while also highlighting the importance of interpretability and human oversight in educational NLP.
Figures & tables
| Age Group | #Mapped Stories | Share |
|---|---|---|
| 0–2 | 30 | 2.9% |
| 2–4 | 544 | 52.2% |
| 4–6 | 687 | 65.9% |
| 6–8 | 510 | 48.9% |
| 8–12 | 345 | 33.1% |
| Annotated stories | 1043 | 94.2% |
| Model | Acc. | Cov. | L-Stab. | A-Stab. |
|---|---|---|---|---|
| GPT-4o | 0.71 0.01 | 17.73 0.15 | 0.87 | 0.89 |
| Gemini-2.5-Pro | 0.87 0.01 | 29.56 0.65 | 0.66 | 0.79 |
| Qwen-Max | 0.73 0.02 | 26.86 0.71 | 0.77 | 0.79 |
| DeepSeek-Chat | 0.69 0.02 | 62.41 0.48 | 0.91 | 0.91 |
| Metric | Score |
|---|---|
| Overall Label Jaccard | 0.772 |
| COG Jaccard | 0.793 |
| LAN Jaccard | 0.774 |
| SEL Jaccard | 0.775 |
| Age Agreement | 0.860 |
| Strict Exact Full-Set Match | 0.180 |
| Method | Overlap | Mean Distance |
|---|---|---|
| Dale–Chall | 0.469 | 0.882 |
| FRE | 0.517 | 0.681 |
| FKGL | 0.603 | 0.502 |
| GPT-4o | 0.730 | 0.672 |
| Gemini-2.5-Pro | 0.768 | 0.497 |
| Qwen-Max | 0.775 | 0.467 |
| Variant | Overlap | Mean Distance |
|---|---|---|
| GPT-4o Direct | 0.730 | 0.672 |
| Free-form + Agg. | 0.791 | 0.411 |
| Taxonomy Labels | 0.852 | 0.231 |
| Taxonomy + Mapping | 0.889 | 0.148 |
| Full CLARA | 0.904 | 0.096 |
| Variant | Overlap | Distance |
|---|---|---|
| Random Mapping | 0.611 | 0.593 |
| Dimension-swapped | 0.742 | 0.366 |
| Coarse Taxonomy | 0.861 | 0.181 |
| Full Taxonomy | 0.904 | 0.096 |
| Criterion | Score |
|---|---|
| COG Label Quality | 4.24 |
| LAN Label Quality | 4.18 |
| SEL Label Quality | 4.31 |
| Educational Interpretability | 4.42 |
| Predicted Age Appropriateness | 4.17 |
| Average | 4.26 |
| Method | Human Agreement |
|---|---|
| FKGL | 0.58 |
| GPT-4o Direct | 0.67 |
| GPT-4o Rubric Direct | 0.73 |
| CLARA (ours) | 0.81 |
| Preference | Human Evaluators |
|---|---|
| Prefer CLARA | 61% |
| Prefer Publisher | 19% |
| Both Reasonable | 20% |
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
| Variant | Overlap | Mean Distance |
|---|---|---|
| Full CLARA | 0.904 | 0.096 |
| w/o COG | 0.861 | 0.164 |
| w/o LAN | 0.873 | 0.151 |
| w/o SEL | 0.842 | 0.183 |
| Metric | Score |
|---|---|
| COG F1 | 0.84 |
| LAN F1 | 0.81 |
| SEL F1 | 0.79 |
| Representative Macro F1 | 0.73 |
| Representative Micro F1 | 0.78 |
| Human–Human Cohen’s | 0.76 |
| Metric | Score |
|---|---|
| Exact Agreement | 0.68 |
| Adjacent Agreement | 0.89 |
| Cohen’s | 0.72 |
| Story Excerpt |
| A cloth bear receives a letter inviting him to visit the big bears when the moon is full. On the journey, the bear takes a bus, climbs a mountain path, and finally joins the celebration with the other bears. |
| Developmental Labels |
| COG: simple cause; action consequence; routine understanding LAN: basic vocabulary; story sequence SEL: emotion expression; preference expression; imitation |
| Publisher Reference |
| 2–5 (mapped to 2–4 and 4–6) |
| CLARA Prediction |
| Story Excerpt |
| A man is rescued by the Nine-Colored Deer after falling into a river. Although he promises to keep the deer secret, he later betrays the deer in exchange for a reward. The story follows the consequences of betrayal, trust, and moral responsibility. |
| Developmental Labels |
| COG: cause effect; prediction; sequence order LAN: narrative inference; discourse understanding SEL: emotion cause; perspective taking; rule awareness |
| Publisher Reference |
| 3–8 (mapped to 2–4, 4–6, 6–8) |
| CLARA Prediction |
| 0–2 years (Sensorimotor stage) |
|---|
| COG: visual tracking, sound response, face recognition, object recognition, object permanence, color recognition, shape recognition, size discrimination, pattern recognition, sensory integration |
| LAN: sound association, babbling pattern, word recognition, word-object mapping, phonetic awareness |
| SEL: emotion reaction, emotion recognition, attachment security, social smiling, caregiver response |
| 2–4 years (Early preoperational stage) |
| COG: simple cause, action consequence, routine understanding, object function, basic counting, category learning, similarity detection, difference detection |
| LAN: basic vocabulary, two-word sentence, sentence imitation, naming objects, story sequence, picture description, symbolic play language |