UniAR: A Unified Framework for Autism Recognition Enhanced by Multi-View Prompt Learning
Authors: Lei Xin, Zeheng Wang, Jiayin Zhu, Shihong Huang, Fanhu Zeng, Changjiang Jiang, Dengbo He, Yutao Yue, +1 more
Organizations: Wuhan University Wuhan, Hubei, China · Northeast Normal University Changchun, Jilin, China · The Hong Kong University of Science and Technology (Guangzhou) Guangzhou, Guangdong, China · Nanjing Agricultural University Nanjing, Jiangsu, China · Imperial College London London, United Kingdom · The Hong Kong University of Science and Technology Hong Kong, China · Harvard University Boston, USA
Autism Spectrum Disorder (ASD) is a complex neurodevelopmental disorder for which early and accurate diagnosis is critical to improving long-term developmental outcomes. However, existing ASD recognition methods are often constrained by the scarcity of diagnostic text data, forcing them to rely mainly on visual analysis and limiting their ability to model clinically meaningful semantic reasoning. To address this challenge, we propose UniAR, a unified framework enhanced by multi-granularity prompt learning for robust ASD recognition under heterogeneous data variations. Specifically, UniAR leverages a large multimodal model to generate hierarchical diagnostic descriptions at the word, phrase, and sentence levels, compensating for the lack of paired clinical reports. To align the generated semantics with visual evidence, we further design a Mixture-of-Experts-based Multi-Scale Alignment Module, which dynamically matches vector-quantized visual prototypes with semantic representations at corresponding granularities. Extensive experiments on four benchmarks covering brain MRI and facial expression scenarios show that UniAR consistently outperforms existing state-of-the-art methods, achieving average accuracies of 75.9% on MRI benchmarks and 91.6% on facial benchmarks, while improving average Accuracy on MRI benchmarks by 1.5 percentage points and average Accuracy on facial benchmarks by 1.2 percentage points over baselines. These results demonstrate that UniAR offers a robust and interpretable framework for ASD screening under semantic scarcity.
Figures & tables
Figure 1. The framework comparisons.Existing autism recognition methods (a-b) are mainly implemented by domain experts or neural networks. In contrast, UniAR (c) integrates expert knowledge with neural networks to provide a more interpretable recognition solution, which is capable of outputting both diagnostic reports and results simultaneously. Comparison of clinical diagnosis, conventional black-box AI diagnosis, and the proposed multi-view prompting framework, which produces both an ASD prediction and interpretable evidence.
Figure 2. Overview of UniAR. The visual stream extracts multi-scale visual prototypes via a learnable visual codebook-augmented multi-hierarchical encoder. The semantic stream uses hierarchical prompts with GPT-4o to generate word-, phrase- and sentence-level descriptions, encoded by a frozen text encoder. A Scale-Aware Alignment Module aligns and fuses these multi-modal features for final classification. Pipeline of UniAR, including hierarchical visual encoding, AI-generated word-, phrase-, and sentence-level diagnostic descriptions, multi-scale cross-modal alignment, and final ASD classification.
Figure 3. Detailed illustration of the Multi-Scale Alignment Module (MSAM) designed to tackle feature inconsistencies. Three-stage multi-scale alignment module showing semantic and visual features aligned through cross-modal attention and fused for prediction.
Method
ABIDE
ABIDE II
AVG
Acc
Recall
Pre
AUC
Acc
Recall
Pre
AUC
Acc
Recall
Pre
AUC
CPM
61.4
58.3
53.3
64.7
60.3
61.0
58.5
62.0
60.9
59.7
55.9
63.4
Braingnn
60.2
60.4
50.3
58.2
58.2
59.3
50.2
57.3
59.2
59.9
50.3
57.8
SpectBGNN
59.6
67.3
51.9
61.4
69.6
71.0
72.2
71.4
64.6
69.2
62.1
66.4
STAGIN
68.5
71.2
69.2
74.1
68.4
67.8
69.1
72.0
68.5
69.5
69.2
73.1
Bolt
67.2
68.4
66.3
71.3
70.2
70.5
71.4
75.0
68.7
69.5
68.9
73.2
Table 1. Comparative experiments on MRI scenarios.
Method
HRM
Kaggle
AVG
Acc
Recall
Pre
AUC
Acc
Recall
Pre
AUC
Acc
Recall
Pre
AUC
RAN
78.4
73.5
75.8
81.0
98.2
97.8
97.5
99.1
88.3
85.7
86.7
90.1
SCN
79.3
74.8
76.5
82.5
94.5
93.2
94.8
95.5
86.9
84.0
85.7
89.0
DMUE
80.5
77.1
77.5
83.9
97.1
96.5
97.8
98.2
88.8
86.8
87.7
91.1
RUL
80.8
76.9
78.4
84.2
97.6
96.8
98.1
98.5
89.2
86.9
88.3
91.4
EAC
81.2
77.5
78.9
84.8
96.3
95.9
96.7
97.5
88.8
86.7
87.8
91.2
Table 2. Comparative experiments on facial expression scenarios.
Method
Modules
ABIDE
ABIDE II
AVG
Word
Phrase
Sent.
Fusion
Acc
Recall
Pre
AUC
Acc
Recall
Pre
AUC
Acc
Recall
Pre
AUC
Baseline
×
×
×
×
76.2
77.1
76.5
80.1
69.1
65.8
70.2
72.3
72.7
71.5
73.4
76.2
w/ Word
✓
×
×
×
76.9
77.6
77.2
80.7
70.0
66.9
71.5
73.1
73.5
72.3
74.4
76.9
w/ Word, Phrase
✓
✓
×
×
77.6
78.3
78.0
81.4
71.0
68.1
72.6
74.2
74.3
73.2
75.3
77.8
w/ Word, Phrase, Sentence
✓
✓
✓
×
78.1
78.8
78.5
82.0
71.8
68.8
73.2
75.0
75.0
73.8
75.9
78.5
Ours
✓
✓
✓
✓
78.5
79.2
78.8
82.5
73.3
69.4
73.7
76.7
75.9
74.3
76.3
79.6
Table 3. Ablation study on MRI scenarios.
Codebook Size
HRM
Kaggle
ABIDE
ABIDE II
Accuracy
Recall
Precision
AUC
Accuracy
Recall
Precision
AUC
Accuracy
Recall
Precision
AUC
Accuracy
Recall
Precision
AUC
64
83.8
80.5
81.1
86.9
98.4
97.6
97.9
98.1
78.1
78.6
78.2
81.9
71.8
68.9
73.2
75.2
128
84.5
81.2
81.8
87.6
98.7
97.9
98.2
98.7
78.5
78.9
78.8
82.5
72.3
69.4
73.7
75.7
256
84.2
81.4
81.5
87.5
98.5
97.7
98.1
98.8
78.3
78.9
78.6
82.6
72.1
69.3
73.5
75.8
Table 4. Experimental results of different codebook sizes on multiple datasets.
Number of Experts
HRM
Kaggle
ABIDE
ABIDE II
Accuracy
Recall
Precision
AUC
Accuracy
Recall
Precision
AUC
Accuracy
Recall
Precision
AUC
Accuracy
Recall
Precision
AUC
2
84.2
80.9
81.5
87.2
98.4
97.7
97.9
98.4
78.2
78.8
78.6
82.2
72.2
69.2
73.5
75.3
3
84.4
81.2
81.7
87.4
98.7
97.7
98.1
98.6
78.4
78.7
78.9
82.3
72.1
69.2
73.6
75.5
4
84.5
81.2
81.8
87.6
98.7
97.9
98.2
98.7
78.5
79.2
78.8
82.5
72.3
69.4
73.7
75.7
Table 5. Comparative experiments with different numbers of experts on multiple datasets.
Figure 4. Qualitative visualization of attention mechanisms and hierarchical semantic alignment.
Figure 5. Data volume statistics of different categories across platforms. Pie charts showing the distribution of normal, mild, moderate, and severe samples across Bilibili, YouTube, and TikTok.
Figure 6. Cross-platform performance comparison on social media datasets. Bar charts comparing UniAR and the EAC baseline on accuracy, F1-score, and recall across three social-media platforms.
Platform
Rouge-1
Rouge-2
Rouge-L
Bilibili
72.3
65.7
70.1
YouTube
78.6
72.4
76.9
TikTok
77.8
71.3
76.2
Table 6. Evaluation metrics on cross-platform datasets.
National Special Education Resource Center for Children with Autism, Zhejiang Normal University, China · School of Computer Science and Technology, Zhejiang Normal University, China · School of Mathematics and Statistics, Xi’an Jiaotong University, China +5
Department of Mechanical Engineering, Stanford University, Stanford, CA 94305, USA · Department of Pediatrics, Stanford University, Stanford, CA 94305, USA · Department of Biomedical Data Science, Stanford University, Stanford, CA 94305, USA +2