cs.CVOct 6, 2026

A Deterministic Evidence Layer for Vision-Language Autism Screening from Naturalistic Home Video

Authors: Wenqi Li, Mindi Ruan, Chuanbo Hu, Shuo Wang, Xin Li

Organizations: Department of Computer Science, University at Albany, SUNY, Albany, NY, USA · Lane Department of Computer Science and Electrical Engineering, West Virginia University, Morgantown, WV, USA · Department of Radiology, Washington University in St. Louis, St. Louis, MO, USA

Abstract

Autism spectrum disorder (ASD) is diagnosed through specialist observation of a child's social behavior, and access to that expertise is the bottleneck for early identification. Vision-language models (VLMs) describe a child's behavior from video well; the verdict drawn from the description is unstable: at temperature~0, across eight pipeline configurations on one backbone, 16--37% of clips change their predicted label between repeated runs, and the cause lies in the serving stack. We keep the VLM frozen and move the decision out of the model. Under our grounded perception constraints the VLM writes an event table of timestamped, glossary-labeled events that names the eliciting press and logs counter-evidence; a text-only stage age-calibrates the confidence of each row; a deterministic weight-of-evidence scorer sums it into an evidence total and stratifies it into a risk category, so every decision decomposes into named per-feature contributions and can be re-scored from the saved table. On 43 caregiver-recorded, protocol-free home free-play clips of preschool children, the pipeline reaches AUC 0.851±0.0120.851 \pm 0.012, 86.0% accuracy, and F1_1 71.8 over three runs. It labels 74.4% of clips correctly in every run (60.5% for the zero-shot baseline) and flags no typically developing clip in every run (9 of 31 at zero-shot). An ablation on the same backbone attributes the gain to the grounded perception constraints read through the deterministic scorer.

Figures & tables

Appendix figures & tables9 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Fine-tuning a multimodal large language model for clinician-grade autism behavioral scoring from short home videos

    Jun 25, 2026Mohammadmahdi Honarmand, Parnian Azizian, Aaron Kline +9Autism Spectrum DisorderLLM Fine-Tuning

  2. A multimodal large language model for evidence-based autism spectrum disorder screening

    Sep 15, 2026Jun Chen, Qi Zhao, Yunliang Jiang +8Autism Spectrum DisorderMultimodal Large Language Models

  3. UniAR: A Unified Framework for Autism Recognition Enhanced by Multi-View Prompt Learning

    Sep 25, 2026Lei Xin, Zeheng Wang, Jiayin Zhu +6Autism Spectrum DisorderMultimodal Learning