UniAR: A Unified Framework for Autism Recognition Enhanced by Multi-View Prompt Learning
Authors: Lei Xin, Zeheng Wang, Jiayin Zhu, Shihong Huang, Fanhu Zeng, Changjiang Jiang, Dengbo He, Yutao Yue, +1 more
Organizations: Wuhan University Wuhan, Hubei, China · Northeast Normal University Changchun, Jilin, China · The Hong Kong University of Science and Technology (Guangzhou) Guangzhou, Guangdong, China · Nanjing Agricultural University Nanjing, Jiangsu, China · Imperial College London London, United Kingdom · The Hong Kong University of Science and Technology Hong Kong, China · Harvard University Boston, USA
Autism Spectrum Disorder (ASD) is a complex neurodevelopmental disorder for which early and accurate diagnosis is critical to improving long-term developmental outcomes. However, existing ASD recognition methods are often constrained by the scarcity of diagnostic text data, forcing them to rely mainly on visual analysis and limiting their ability to model clinically meaningful semantic reasoning. To address this challenge, we propose UniAR, a unified framework enhanced by multi-granularity prompt learning for robust ASD recognition under heterogeneous data variations. Specifically, UniAR leverages a large multimodal model to generate hierarchical diagnostic descriptions at the word, phrase, and sentence levels, compensating for the lack of paired clinical reports. To align the generated semantics with visual evidence, we further design a Mixture-of-Experts-based Multi-Scale Alignment Module, which dynamically matches vector-quantized visual prototypes with semantic representations at corresponding granularities. Extensive experiments on four benchmarks covering brain MRI and facial expression scenarios show that UniAR consistently outperforms existing state-of-the-art methods, achieving average accuracies of 75.9% on MRI benchmarks and 91.6% on facial benchmarks, while improving average Accuracy on MRI benchmarks by 1.5 percentage points and average Accuracy on facial benchmarks by 1.2 percentage points over baselines. These results demonstrate that UniAR offers a robust and interpretable framework for ASD screening under semantic scarcity.
Figures & tables
Figure 1. The framework comparisons.Existing autism recognition methods (a-b) are mainly implemented by domain experts or neural networks. In contrast, UniAR (c) integrates expert knowledge with neural networks to provide a more interpretable recognition solution, which is capable of outputting both diagnostic reports and results simultaneously. Comparison of clinical diagnosis, conventional black-box AI diagnosis, and the proposed multi-view prompting framework, which produces both an ASD prediction and interpretable evidence.
Figure 2. Overview of UniAR. The visual stream extracts multi-scale visual prototypes via a learnable visual codebook-augmented multi-hierarchical encoder. The semantic stream uses hierarchical prompts with GPT-4o to generate word-, phrase- and sentence-level descriptions, encoded by a frozen text encoder. A Scale-Aware Alignment Module aligns and fuses these multi-modal features for final classification. Pipeline of UniAR, including hierarchical visual encoding, AI-generated word-, phrase-, and sentence-level diagnostic descriptions, multi-scale cross-modal alignment, and final ASD classification.
Figure 3. Detailed illustration of the Multi-Scale Alignment Module (MSAM) designed to tackle feature inconsistencies. Three-stage multi-scale alignment module showing semantic and visual features aligned through cross-modal attention and fused for prediction.
Method
ABIDE
ABIDE II
AVG
Acc
Recall
Pre
AUC
Acc
Recall
Pre
AUC
Acc
Recall
Pre
AUC
CPM
61.4
58.3
53.3
64.7
60.3
61.0
58.5
62.0
60.9
59.7
55.9
63.4
Braingnn
60.2
60.4
50.3
58.2
58.2
59.3
50.2
57.3
59.2
59.9
50.3
57.8
SpectBGNN
59.6
67.3
51.9
61.4
69.6
71.0
72.2
71.4
64.6
69.2
62.1
66.4
STAGIN
68.5
71.2
69.2
74.1
68.4
67.8
69.1
72.0
68.5
69.5
69.2
73.1
Bolt
67.2
68.4
66.3
71.3
70.2
70.5
71.4
75.0
68.7
69.5
68.9
73.2
Table 1. Comparative experiments on MRI scenarios.
Method
HRM
Kaggle
AVG
Acc
Recall
Pre
AUC
Acc
Recall
Pre
AUC
Acc
Recall
Pre
AUC
RAN
78.4
73.5
75.8
81.0
98.2
97.8
97.5
99.1
88.3
85.7
86.7
90.1
SCN
79.3
74.8
76.5
82.5
94.5
93.2
94.8
95.5
86.9
84.0
85.7
89.0
DMUE
80.5
77.1
77.5
83.9
97.1
96.5
97.8
98.2
88.8
86.8
87.7
91.1
RUL
80.8
76.9
78.4
84.2
97.6
96.8
98.1
98.5
89.2
86.9
88.3
91.4
EAC
81.2
77.5
78.9
84.8
96.3
95.9
96.7
97.5
88.8
86.7
87.8
91.2
Table 2. Comparative experiments on facial expression scenarios.
Method
Modules
ABIDE
ABIDE II
AVG
Word
Phrase
Sent.
Fusion
Acc
Recall
Pre
AUC
Acc
Recall
Pre
AUC
Acc
Recall
Pre
AUC
Baseline
×
×
×
×
76.2
77.1
76.5
80.1
69.1
65.8
70.2
72.3
72.7
71.5
73.4
76.2
w/ Word
✓
×
×
×
76.9
77.6
77.2
80.7
70.0
66.9
71.5
73.1
73.5
72.3
74.4
76.9
w/ Word, Phrase
✓
✓
×
×
77.6
78.3
78.0
81.4
71.0
68.1
72.6
74.2
74.3
73.2
75.3
77.8
w/ Word, Phrase, Sentence
✓
✓
✓
×
78.1
78.8
78.5
82.0
71.8
68.8
73.2
75.0
75.0
73.8
75.9
78.5
Ours
✓
✓
✓
✓
78.5
79.2
78.8
82.5
73.3
69.4
73.7
76.7
75.9
74.3
76.3
79.6
Table 3. Ablation study on MRI scenarios.
Codebook Size
HRM
Kaggle
ABIDE
ABIDE II
Accuracy
Recall
Precision
AUC
Accuracy
Recall
Precision
AUC
Accuracy
Recall
Precision
AUC
Accuracy
Recall
Precision
AUC
64
83.8
80.5
81.1
86.9
98.4
97.6
97.9
98.1
78.1
78.6
78.2
81.9
71.8
68.9
73.2
75.2
128
84.5
81.2
81.8
87.6
98.7
97.9
98.2
98.7
78.5
78.9
78.8
82.5
72.3
69.4
73.7
75.7
256
84.2
81.4
81.5
87.5
98.5
97.7
98.1
98.8
78.3
78.9
78.6
82.6
72.1
69.3
73.5
75.8
Table 4. Experimental results of different codebook sizes on multiple datasets.
Number of Experts
HRM
Kaggle
ABIDE
ABIDE II
Accuracy
Recall
Precision
AUC
Accuracy
Recall
Precision
AUC
Accuracy
Recall
Precision
AUC
Accuracy
Recall
Precision
AUC
2
84.2
80.9
81.5
87.2
98.4
97.7
97.9
98.4
78.2
78.8
78.6
82.2
72.2
69.2
73.5
75.3
3
84.4
81.2
81.7
87.4
98.7
97.7
98.1
98.6
78.4
78.7
78.9
82.3
72.1
69.2
73.6
75.5
4
84.5
81.2
81.8
87.6
98.7
97.9
98.2
98.7
78.5
79.2
78.8
82.5
72.3
69.4
73.7
75.7
Table 5. Comparative experiments with different numbers of experts on multiple datasets.
Figure 4. Qualitative visualization of attention mechanisms and hierarchical semantic alignment.
Figure 5. Data volume statistics of different categories across platforms. Pie charts showing the distribution of normal, mild, moderate, and severe samples across Bilibili, YouTube, and TikTok.
Figure 6. Cross-platform performance comparison on social media datasets. Bar charts comparing UniAR and the EAC baseline on accuracy, F1-score, and recall across three social-media platforms.
Platform
Rouge-1
Rouge-2
Rouge-L
Bilibili
72.3
65.7
70.1
YouTube
78.6
72.4
76.9
TikTok
77.8
71.3
76.2
Table 6. Evaluation metrics on cross-platform datasets.
The clinical management of autism spectrum disorder (ASD) faces a bottleneck in early screening, mainly because trained specialists are scarce and conventional assessment tools are subjective. Here, we introduce ASDchat, a multimodal large language model designed for evidence-based ASD screening, which takes video, audio, and dialogue as input. ASDchat adopts a dual-branch architecture, where the decision branch generates screening probabilities and the evidence branch generates traceable, timestamped behavioral evidence aligned with standardized clinical criteria (ADOS-2). The model was trained and evaluated on a dataset of 1,035 participants from 27 sites in China, which covered typically developing (TD) children, children with ASD, and children with other disorders. For ASD versus TD, ASDchat reached an area under the receiver operating characteristic curve (AUC) of 0.953 ± 0.021. On 9 held-out sites that were not used for training, the mean AUC was 0.932. Furthermore, unsupervised clustering of the behavioral dimensions split the ASD cases into six subtypes with different phenotypic profiles, and ASDchat suggests an intervention for each subtype. ASDchat provides a feasible path for large-scale, evidence-based early ASD screening in clinical practice.
Jun Chen, Qi Zhao, Yunliang Jiang +8
National Special Education Resource Center for Children with Autism, Zhejiang Normal University, China · School of Computer Science and Technology, Zhejiang Normal University, China · School of Mathematics and Statistics, Xi’an Jiaotong University, China +5
Accurate Autism Spectrum Disorder (ASD) screening for school-age children is crucial to identify cases that may have been missed earlier and to enable timely interventions supporting social, cognitive, and academic development. Current ASD screening relies on subjective assessments and 2D analysis methods that fail to capture spatial displacement patterns characteristic of ASD behaviors. In this study, a novel 3D temporal analysis framework is presented, built on top of DECA (Detailed Expression Capture and Animation), a 3D modeling framework, to extract comprehensive head pose parameters (including translational components Tx,Ty,Tz) and facial expressions independent of pose variations. LSTM and GRU-based temporal classifiers were trained on the extracted 3D features from video data collected from 39 participants (19 ASD, 20 TD) aged 7-12 years during Virtual Reality-Continuous Performance Test tasks. The GRU-based models demonstrated superior performance, with 3D head pose features achieving 83.9% accuracy and 3D facial features reaching 81.4% accuracy, outperforming 2D baseline approaches by 10.7% and 7.5%, respectively. Furthermore, multimodal fusion of 3D head pose and facial features with PCA-based dimensionality reduction achieved the highest accuracy of 84.6%, outperforming unimodal approaches. This work establishes a foundation for objective, automated screening tools addressing current diagnostic limitations in ASD identification for school-age populations.
Inam Qadir, Elizabeth B Varghese, Dena Al-Thani +1
College of Science and Engineering, Hamad Bin Khalifa University, Qatar Foundation, Doha, Qatar
Autism spectrum disorder (ASD) affects 1 in 31 US children, yet median age at diagnosis exceeds four years. Artificial intelligence pipelines that provide quantified diagnosis using easy to access observational data (e.g., home videos) could help with earlier diagnosis, and timely delivery of early treatments. We fine-tuned Gemini 2.5 Pro on 400 clinician-rated home videos with low-rank adaptation, training only on 30 behavioral features previously validated to produce reliable predictions when passed to various ML models. On 99 held-out children (49 ASD, 50 neurotypical), inter-rater reliability with clinicians (per-feature weighted Cohen's kappa) improved by 40% (p<0.001), with 27 of 28 evaluable features improving. As an emergent zero-shot capability, direct ASD diagnosis F1 improved by 53% (p<0.001), matching or exceeding clinician outcomes. Classifier-assisted pipelines using fine-tuned LLM-derived behavioral features matched clinician-scored inputs across all tested pathways and achieved 77% accuracy (95% CI: 68-85%) and an AUC of 86% (95% CI: 78-92%). Fine-tuned multimodal LLMs can serve as scalable behavioral feature extractors for use in autism assessment and diagnosis.
Department of Mechanical Engineering, Stanford University, Stanford, CA 94305, USA · Department of Pediatrics, Stanford University, Stanford, CA 94305, USA · Department of Biomedical Data Science, Stanford University, Stanford, CA 94305, USA +2