MERID: Multimodal Exploration via Recursive Self-Improvement Agents for Major Depression Analysis
Authors: Lei Liu, Zhaokang Liang, Qingcheng Zeng, Chenda Duan, Lu Mi, Zhen Tan, Tianyu Liu
Organizations: Yale University · Zhejiang University · Northwestern University · University of California, Los Angeles · Tsinghua University · Stevens Institute of Technology
Major depressive disorder (MDD) severely impacts daily activities and quality of life. Detecting MDD involves multimodal data, such as interview recordings and sensor measurements. This is particularly challenging, as these heterogeneous modalities often demand distinct, customized prediction pipelines. Existing efforts to address this challenge have explored both manually engineered multimodal architectures and agent-assisted pipeline development. Despite their progress, it remains challenging to autonomously revise pipelines based on experimental feedback and carry verified improvements forward into subsequent designs. To this end, we propose Multimodal Exploration via Recursive Self-Improvement Agents for Major Depression Analysis (MERID). The framework develops depression pipelines through experience-based recursive self-improvement (RSI). Grounded State Construction (GSC) grounds experience by aligning multimodal records with subject-level depression targets. Coupled Pipeline Exploration (CPE) jointly modifies representations, fusion, and predictors to build successor pipelines for classification and severity estimation. Evidence-Guided Evolution (EGE) guides revisions through feedback and verifies gains under uncertainty in small depression cohorts before inheritance. Extensive experiments on depression benchmarks show that MERID achieves the best results on multiple tasks compared with multimodal and agent-based baselines. Further analysis highlights the value of acoustic and linguistic cues for depression detection. Our code is available at https://github.com/DiscoAILab/MERID
Figures & tables
Figure 2: Conventional approaches require repeated experiments and manual pipeline tuning for depression detection. MERID uses experimental feedback to revise pipelines and verify improvements. Retained pipelines and experience guide recursive self-improvement for depression assessment.
Figure 3: The framework of the proposed MERID . GSC aligns experimental records, and CPE jointly explores pipeline components. EGE uses feedback to guide revisions and verifies gains under uncertainty. The retained pipeline and updated experience guide the next RSI round.
Dataset
E-DAIC
MPDD-2025
Mental Health
Kaggle
Method
Binary F1 ↑
PHQ-8 CCC ↑
Elder Track ↑
Young Track ↑
Status F1 ↑
A–V F1 ↑
Avg. rank ↓
MulT
0.5221
-0.0047
0.4549
0.3952
—
—
4.50 (4)
Flex-MoE
0.5993
0.0135
0.5440
0.4483
0.9548
0.6201
2.67 (6)
MCMoE
0.4105
0.0000
0.4511
0.3419
n/a
n/a
5.75 (4)
Gemini 2.5 Flash
0.3978
-0.0917
0.4322
0.3995
0.0843
0.2966
5.33 (6)
MDAgents
0.4697
0.0520
0.2096
0.4069
0.0607
0.3845
4.33 (6)
Table 1: Performance comparison on depression benchmarks and supplementary tasks, with five-seed means for MERID using Claude Opus 4.8 and GPT-6 Astra. Best and second-best results are bold and underlined . See Appendices B.1 and B.2 for per-seed results.
Figure 4: Five-seed score differences from the best baseline in Table 1 (markers: seeds; bars: means). Positive values favor MERID . The axis is linear within ±0.1 and logarithmic beyond.
Figure 5: Ablation visualization and sensitivity analyses with Opus 4.8; see Appendix B.5 .
Figure 7: RSI source evolution (a) and score–cost comparisons (b,c). MERID points are five-seed means. Details: Appendices C.2 and D .
Appendix figures & tables33 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 8: A recorded MPDD-2025 evidence-guided evolution step. Dialogue is condensed from the review record. The final arrow tracks the best development candidate, which is distinct from the retained incumbent.
Aspect
DepressionAgent (text branch)
MERID
Evidence
Interview statements and supporting or countervailing interpretations
Task constraints, candidate scores, prediction diagnostics, and execution checks
Revision target
The current participant’s risk judgment
Representation–fusion recipes and predictor configurations
Depression-specific action
Re-read the transcript for overlooked risk evidence
Check clinical-scale exclusion and propose an ordinal severity predictor
Verification
Reconsider evidence through review and re-evaluation
Evaluate proposed pipelines under a shared protocol and replacement margin
Retained consequence
A final case judgment and its reasoning record
A reusable pipeline, admitted components, and experience for the next round
Appendix
Table 2: How depression-related evidence enters revision. DepressionAgent describes the evaluated text-branch adaptation. The MERID column describes its pipeline-development loop, with the concrete example in Figure 8 .
Task
Metric
Seed 0
Seed 1
Seed 2
Seed 3
Seed 4
Mean ± SD
E-DAIC binary
Macro-F1
0.6655
0.6937
0.6937
0.6655
0.6655
0.6768 ± 0.0154
E-DAIC PHQ-8
CCC
0.3201
0.3201
0.3538
0.3201
0.3201
0.3268 ± 0.0151
MPDD-2025 Elder
Track score
0.5820
0.6071
0.6177
0.6282
0.6335
0.6137 ± 0.0204
MPDD-2025 Young
Track score
0.4116
0.3545
0.3610
0.3755
0.3594
0.3724 ± 0.0233
Mental Health
Macro-F1
0.9704
0.9704
0.9630
0.9527
0.9710
0.9655 ± 0.0079
Kaggle A–V
Macro-F1
0.7811
0.7624
0.7803
0.8074
0.7869
0.7836 ± 0.0162
Appendix
Table 3: Test scores of MERID under seeds 0–4 for every column of Table 1 . Each seed is one complete run (development search, review and a single test evaluation). Best and second-best distinct seed scores within each row are bold and underlined, including ties. The last column reports the mean and sample standard deviation over the five runs.
Figure 10: Test scores under seeds 0–4 with Claude Opus 4.8 (blue) and GPT-6 Astra (orange) as the base model; bars are seed means. Each panel keeps its task’s metric and scale.
Task
Metric
Opus 4.8
GPT-6 Astra
Δ
E-DAIC binary
Macro-F1
0.6768 ± 0.0154
0.7057 ± 0.0118
+0.0289
E-DAIC PHQ-8
CCC
0.3268 ± 0.0151
0.3201 ± 0.0000
-0.0067
MPDD-2025 Elder
Track score
0.6137 ± 0.0204
0.6148 ± 0.0291
+0.0011
MPDD-2025 Young
Track score
0.3724 ± 0.0233
0.3714 ± 0.0261
-0.0010
Mental Health
Macro-F1
0.9655 ± 0.0079
0.9565 ± 0.0042
-0.0090
Kaggle A–V
Macro-F1
0.7836 ± 0.0162
0.6070 ± 0.0000
-0.1766
Appendix
Table 4: Test scores of MERID with Claude Opus 4.8 and GPT-6 Astra as the base model: mean ± sample SD over seeds 0–4 of the same 16 run units, protocol and code. Δ is Astra minus Opus.
Quantity
Opus 4.8
GPT-6 Astra
Run attempts for 80 tasks
80
85
Aborted attempts
0
5
MPDD-2026 proposals audited
112
135
invalid (run aborts)
0
5 (3.7%)
duplicates (re-asked once)
0
0
Chair replies re-asked for format
35
0
Appendix
Table 5: Failure analysis of the two five-seed studies. An attempt is one run of a task; aborted attempts were re-run under the same seed, at most twice. MPDD-2026 audits every proposal: an invalid one ends the run, a duplicate is re-asked once. The chair is re-asked when its reply lacks the required format. Token counts and costs cover completed runs; costs use uncached list prices, an upper bound when part of the input was cached.
Task
Seed
Att.
Prop.
Failure
Offending fragment
MPDD-2026 Elder
0
1
7
out-of-space value
"sev_fusion": "wav2emb:densenet"
MPDD-2026 Elder
2
1
6
out-of-space value
"audio": "wav2span"
MPDD-2026 Elder
2
2
1
out-of-space value
"audio": "wav2.0"
MPDD-2026 Elder
4
1
6
malformed JSON
gp_stats_pz","model":"plsn4":"bad"}}
MPDD-2026 Young
2
1
9
malformed JSON
1.0","sev_model":"gb0.05d1,"sev_fusion":
Appendix
Table 6: The invalid GPT-6 Astra proposals that aborted a run: task, seed, attempt, the proposal’s position in the run, and the offending part of its final JSON line, checked against the action space.
Task / test metric ↑
w/o CPE
w/o EGE
MERID (full)
E-DAIC binary (F1)
0.5151 ± 0.0760
0.5162 ± 0.0345
0.6768 ± 0.0154
E-DAIC PHQ-8 (CCC)
0.0361 ± 0.0452
0.2997 ± 0.0196
0.3268 ± 0.0151
Mental Health (F1)
0.9672 ± 0.0149
0.9530 ± 0.0005
0.9655 ± 0.0079
MPDD-25 Elder 1 s binary
0.7231 ± 0.0846
0.7850 ± 0.0042
0.7593 ± 0.0349
MPDD-25 Elder 1 s ternary
0.6089 ± 0.0822
0.5955 ± 0.0667
0.6295 ± 0.0673
MPDD-25 Elder 1 s quinary
0.4583 ± 0.0341
0.4472 ± 0.0482
0.4738 ± 0.0647
Appendix
Table 7: Module ablations. Test results report mean ± sample SD over seed means: MERID (full) is Table 1 ’s five-seed result (seeds 0–4); w/o CPE and w/o EGE use seeds 0–2 and start without an inherited pipeline. Repeated runs are averaged within each seed. MPDD-2025 uses the official composite score. GSC results use development data with one run per configuration. Best and second-best distinct means within each row are bold and underlined .
Binary classification
Available
Dev Macro-F1
Test Macro-F1 ↑
Test Accuracy ↑
Deployed
A
0.6348 ± 0.0000
0.5336 ± 0.0679
0.5655 ± 0.0412
A
V
0.5754 ± 0.0476
0.6086 ± 0.0092
0.6488 ± 0.0103
V
T
0.6750 ± 0.0531
0.6107 ± 0.0535
0.6607 ± 0.0309
T
AV
0.6348 ± 0.0000
0.5612 ± 0.1064
0.6726 ± 0.0103
A+V
AT
0.7348 ± 0.0000
0.6468 ± 0.0153
0.6964 ± 0.0179
A+T
Appendix
Table 8: Available modalities on E-DAIC. Values are mean ± sample SD over three seeds. Development scores describe the selected candidate; test scores describe the locked deployment. PHQ-8 development RMSE is the negated selection score. Deployed sources are the union across ensemble members within a run; semicolons separate configurations observed across seeds. Best and second-best test means within each task and metric are bold and underlined.
Figure 11: Additional sensitivity scans from the recorded configurations. Top: E-DAIC PHQ CCC versus budget, maximum revision rounds, and seed. Middle: MPDD-2025 1 s ternary official scores versus folds, budget, and maximum rounds. Bottom: MPDD-2026 joint track scores versus budget, maximum rounds, and seed. Points average available runs at each setting, and whiskers show repeated-run ranges where nonzero. Panels use separate metric scales. These are separate runs at different settings, rather than successive states of one trajectory.
Selection rule
Development
Test
Incumbent + one standard error
0.7172
0.6843[0.6655,0.6937]
Classic one-standard-error rule
0.7737
0.4991
Maximum development score
0.7737
0.4876
Appendix
Table 9: Selection-rule sensitivity on E-DAIC binary classification. The incumbent-based rule reports the mean and range of three test runs; alternative rules were each evaluated once. Development and test scores are Macro-F1 ( ↑ ).
Task / metric
Reference + one SE
Winner’s-curse configuration
Search selection (one SE)
E-DAIC binary (Macro-F1)
0.6937
0.6814
0.5492
E-DAIC PHQ-8 (CCC)
0.3201
0.3201
0.3201
Mental Health (Macro-F1)
0.9630
0.9535
0.9535
MPDD-25 Elder 1 s binary
0.7212
0.7212
0.7212
MPDD-25 Elder 1 s ternary
0.6842
0.6842
0.6220
MPDD-25 Elder 1 s quinary
0.3962
0.5490
0.4011
Appendix
Table 10: Historical deployment configurations, with one seed-0 run per cell. MPDD-2025 uses the official composite. The Mental Health winner’s-curse-family record predates that gate and uses only the rank-transfer check. Best and second-best distinct values per row are bold and underlined. Each cell uses the latest logged run for its protocol, configuration, task, and seed.
Figure 12: Selection and verification analyses. (a,b) E-DAIC selection-rule and revision records; the test whisker spans three incumbent-based runs, while each alternative has one run. (c,d) A separate terminal-margin replay of five proposed promotions. E-DAIC uses Macro-F1; MPDD-2025 Elder uses its official score. Each task is shown separately. Panel (d) compares saved promoted single-model predictions with the recorded reverted deployment.
Task
Full
Staged
Fixed recipe
Fixed predictor
1-SE
Random predictor
Random recipe
E-DAIC binary
0.6937
0.4991
0.6814
0.4593
0.4991
0.5564
0.5735
E-DAIC PHQ-8
0.3201
0.3201
0.3201
0.3201
0.3201
0.3307
0.3201
Mental Health
0.9630
0.9711
0.9527
0.9305
0.9535
0.9527
0.9539
MPDD-25 binary
0.7212
0.7543
0.7212
0.7212
0.7565
0.7266
—
MPDD-25 ternary
0.6390
0.6443
0.6922
0.6482
0.6183
0.6619
—
MPDD-25 quinary
0.4115
0.4066
0.4514
0.4317
0.4169
0.4624
—
Appendix
Table 11: Search and selection variants. Ordinary arms retain one seed-0 run per task; random-space arms average three seeds. MPDD-2025 rows then average the 1 s and 5 s Elder tasks. Metrics are Macro-F1 for binary/status classification, CCC for PHQ-8, and the official composite for MPDD-2025. MPDD random recipes are inapplicable because the proposed recipe space already exhausts the 12 eligible combinations. Best and second-best distinct values are bold and underlined. For each protocol, configuration, task, and seed, the latest logged run is used before seed and window aggregation.
Task
Memory off
Evolving
Δ
Calls (off → evolving)
E-DAIC binary (F1)
0.4927
0.5099
+0.0172
19.7 → 33.7
E-DAIC PHQ-8 (CCC)
0.3142
0.3311
+0.0168
28.7 → 29.3
Mental Health (F1)
0.9647
0.9582
-0.0065
27.0 → 33.0
MPDD-25 Elder 1 s binary
0.7802
0.7886
+0.0084
22.0 → 18.0
MPDD-25 Elder 1 s ternary
0.6557
0.6335
-0.0222
18.0 → 39.7
MPDD-25 Elder 1 s quinary
0.4164
0.4164
+0.0000
22.0 → 25.7
Appendix
Table 12: Research-memory ablation. Test scores and LLM calls are means over three seeds. Δ is evolving minus off on the indicated task metric.
Figure 13: Retained development scores across recorded revision steps in the separate research-memory experiment. Colors distinguish memory settings, and line styles and markers identify seeds. Each panel includes all six task-specific runs. Scores are Macro-F1 for E-DAIC binary and Mental Health, CCC for E-DAIC PHQ, and the official composite for MPDD-2025 Elder. Curves stop at their last record, and overlapping trajectories retain their measured values.
Memory
Step
Runs
New evaluations
Retained-score gains
Off
1
23
751
2
Off
2
5
107
0
Off
3
3
52
0
Evolving
1
22
727
3
Evolving
2
12
270
0
Evolving
3
8
157
0
Appendix
Table 13: Candidate evaluations and retained-score gains by recorded revision step. Each run contributes only to steps present in its summary. The step counts cover the 54-run research-memory experiment.
Figure 14: Source and memory measures in the amended eight-round study. (a-c) Six E-DAIC runs per arm; curves average available records with ± 1 standard-error bands. Panel (a) is candidate-weighted L2; it differs from the source-slot proportions in the main figure. Gray lines in (b,c) are binary and PHQ-8 training-label references. (d) Lessons and review minutes from nine runs per review arm. (e) Citation rates use nine memory-enabled runs; action matching uses the six E-DAIC runs with applicable entries. (f) Run counts for panel (a), with rows following the arm legend. Counts describe available task-run records, not additional independent seeds.
Figure 15: Development and test trajectories under the amended eight-round protocol. Rows show E-DAIC binary classification, E-DAIC PHQ-8, and MPDD-2025 Elder 1 s ternary classification. Left: best newly evaluated candidate (solid) and retained incumbent (dashed), using available development records. Right: the locked deployment at each round. After a run ends, its final deployment score is carried forward through round eight; unrecorded intermediate rounds remain excluded from that round’s mean. Each task and arm has three runs; intermediate test means use one to three available records. Bands show ± 1 standard error when at least two records are available. Each task retains its own metric scale.
Task / test metric
Evolving − random
Evolving − no memory
E-DAIC binary / Macro-F1
+0.0004 [-0.0176, 0.0185]
−0.0017 [-0.0206, 0.0173]
E-DAIC PHQ-8 / CCC
+0.0024 [-0.0023, 0.0071]
+0.0045 [0.0013, 0.0079]
MPDD-25 Elder 1 s ternary / official
+0.0076 [-0.0093, 0.0245]
+0.0002 [-0.0121, 0.0125]
Appendix
Table 14: Secondary L4 slope contrasts per round. Each arm has three runs per task. Brackets are unadjusted 95% run-bootstrap intervals; task metrics are kept separate.
Task
Seed
Add / revise / retire
Citation rate
Action-matched use
E-DAIC binary
0
14 / 12 / 2
0.73
1 / 7
E-DAIC binary
1
11 / 15 / 2
0.82
6 / 22
E-DAIC binary
2
17 / 10 / 1
0.72
6 / 21
E-DAIC PHQ-8
0
14 / 15 / 0
0.80
1 / 5
E-DAIC PHQ-8
1
14 / 12 / 1
0.66
7 / 33
E-DAIC PHQ-8
2
8 / 6 / 2
0.81
6 / 14
Appendix
Table 15: Research-memory use in the nine evolving eight-round runs. Operations count entries added, revised and retired over the run. Citation rate is the fraction of active entries whose identifier appears in the next round’s minutes or proposals. Action matching counts activated entries whose next action names a source family and whose next round adds that family, over the activated entries that name one; this source-family proxy is computed for E-DAIC only.
Task
Arm
Seed
Reviews
Failed
Requested / declined
New rounds
Candidates
E-DAIC binary
Memory
0
8
1
3 / 1
6
448
E-DAIC binary
Memory
1
8
1
1 / 0
7
512
E-DAIC binary
Memory
2
8
1
2 / 1
6
512
E-DAIC binary
No memory
0
8
0
6 / 0
8
512
E-DAIC binary
No memory
1
8
1
4 / 0
7
480
E-DAIC binary
No memory
2
8
1
5 / 0
7
480
Appendix
Table 16: Review and exploration counts for every amended eight-round run. Reviews counts scheduled review rounds, including failures. Requested / declined counts explicit follow-up requests for a revision and requests that still produced none. New rounds counts revision rounds with newly evaluated candidates. Random revision exhausts its drawable MPDD-2025 space after five rounds.
Figure 16: LLM efficiency. (a) E-DAIC test Macro-F1 versus recorded calls: MERID uses the means over its five seed runs; inference baselines cover 56 test participants. DepressionAgent uses its scored 268-call run. Non-LLM methods appear at zero calls. (b,c) Diamonds show medians of nine task-level means from the module ablations, with full MERID from the five-seed runs of Table 1 ; vertical lines show their minimum-to-maximum ranges, not uncertainty intervals. Local fitting and feature extraction are outside this accounting.
Task
MERID (full)
w/o CPE
w/o EGE
Runs
E-DAIC binary
11.4 / $0.60
83.0 / $4.67
2.0 / $0.16
5 / 1 / 3
E-DAIC PHQ-8
11.4 / $0.61
41.0 / $2.12
1.0 / $0.06
5 / 1 / 3
Mental Health
10.8 / $0.64
89.0 / $5.09
0.0 / $0.00
5 / 1 / 3
MPDD-25 Elder 1 s binary
12.2 / $0.65
89.0 / $5.29
1.0 / $0.01
5 / 1 / 1
MPDD-25 Elder 1 s ternary
12.4 / $0.70
20.0 / $1.06
1.0 / $0.01
5 / 1 / 1
MPDD-25 Elder 1 s quinary
12.2 / $0.69
33.0 / $1.89
1.0 / $0.01
5 / 1 / 1
Appendix
Table 17: LLM usage per run of the module ablations: mean LLM calls / USD over the logged runs of each task and configuration (Claude Opus 4.8 at 5/25 per million input/output tokens). Local feature extraction and model fitting are excluded. Full MERID uses the five runs of Table 1 ; for the removals, recovered runs without a cost record are not included, so their sample differs from the performance sample of Table 7 . Runs gives the number of logged runs per configuration (full / w/o CPE / w/o EGE).
Memory
Runs
Cap
Total mean
Total range
Revision mean
Memory off
27
320
264.4
139–320
33.7
Evolving memory
27
320
272.9
139–320
42.7
Appendix
Table 18: Realized candidate evaluations in the 54-run research-memory comparison. Each arm covers nine tasks and three seeds. Means are over runs; ranges span their totals. Revision evaluations exclude initialization.
Task
Arm
Cap
Initial mean
Total mean
Total range
E-DAIC binary
Memory
512
254.7
490.7
448–512
E-DAIC binary
No memory
512
256.0
490.7
480–512
E-DAIC binary
Random
512
256.0
512.0
512–512
E-DAIC PHQ-8
Memory
512
177.3
512.0
512–512
E-DAIC PHQ-8
No memory
512
182.0
512.0
512–512
E-DAIC PHQ-8
Random
512
186.7
512.0
512–512
Appendix
Table 19: Exploration effort in the amended eight-round study. Each row covers three seeds. The cap includes initialization and revisions. Counts include all recorded candidate evaluations. The total range is the observed minimum–maximum across the three runs.
Dataset
Cohort
Task
Metric
Official
MulT
Flex-MoE
Gemini 2.5 Flash
MDAgents
Ours
E-DAIC
–
Binary
Accuracy ↑
0.5893
0.7143
0.6607
0.6429
0.7393
Bal. Acc. ↑
0.5226
0.5958
0.4744
0.4947
0.6702
Macro-F1 ↑
0.5221
0.5993
0.3978
0.4697
0.6768
W. F1 ↑
0.5925
0.6836
0.5541
0.5887
0.7326
κ↑
0.0445
0.2209
-0.0683
-0.0127
0.3554
PHQ-8
RMSE ↓
6.9902
7.0394
7.7113
6.6201
5.9098
Appendix
Table 20: Complete performance metrics, continued over three parts. MERID reports means over seeds 0–4 (Kaggle Video: retained run); baselines are unchanged. This part reports E-DAIC, MPDD-2026, Mental Health, and Kaggle. Arrows indicate preferred directions. Best and second-best distinct values among the methods shown are bold and underlined, including ties; dashes denote unavailable results. Bal. Acc. and W. F1 denote balanced accuracy and weighted F1. Avg. Rank uses one primary metric per task: Macro-F1 for classification, RMSE for PHQ regression, and the official score for MPDD-2025. Per-task ranks use midranks among reported methods; averages require complete task coverage and exclude track aggregates. MPDD-2026 denotes MPDD-AVG 2026. Its MulT and Flex-MoE PHQ-9 estimates are derived from ternary-class midpoints. Kaggle evaluates acted-emotion proxy labels; MERID uses actor-disjoint nested cross-validation with review disabled.
Dataset
Cohort
Task
Metric
Official
MulT
Flex-MoE
Gemini 2.5 Flash
MDAgents
Ours
MPDD-2025
Elder
1 s Binary
Accuracy ↑
0.6916
0.7841
0.7416
0.7357
0.1982
0.7789
Macro-F1 ↑
0.5900
0.6318
0.6266
0.5048
0.1809
0.6964
W. Acc. ↑
0.6360
0.6332
0.6597
0.5055
0.5035
—
W. F1 ↑
0.7222
0.7852
0.7604
0.7238
0.1038
0.7987
Task Acc. ↑
0.6638
0.7086
0.7006
0.6206
0.3509
0.7712
Task F1 ↑
0.6561
0.7085
0.6935
0.6143
0.1424
0.7475
Appendix
Table 21: Complete performance metrics (continued): MPDD-2025 Elder. All seven task metrics and all three track aggregates are shown. W. Acc. denotes inverse-class-frequency-weighted accuracy, and W. F1 denotes support-weighted F1. Task Acc. averages Accuracy and W. Acc.; Task F1 averages Macro-F1 and W. F1; Official averages Task Acc. and Task F1. Track metrics average the corresponding task metrics over both windows. Flex-MoE cells report the means of three runs. Ours Track Acc. and Track F1 are computed from the reported task components. Ours W. Acc. was not included in the available main-run summary and is left unreported. Formatting follows the first part.
Dataset
Cohort
Task
Metric
Official
MulT
Flex-MoE
Gemini 2.5 Flash
MDAgents
Ours
MPDD-2025
Young
1 s Binary
Accuracy ↑
0.5644
0.4924
0.6048
0.4848
0.5455
0.4038
Macro-F1 ↑
0.5446
0.4491
0.5749
0.3450
0.4762
0.4018
W. Acc. ↑
0.5644
0.4924
0.6048
0.4848
0.5455
—
W. F1 ↑
0.5446
0.4491
0.5749
0.3450
0.4762
0.4018
Task Acc. ↑
0.5644
0.4924
0.6048
0.4848
0.5455
0.4038
Task F1 ↑
0.5446
0.4491
0.5749
0.3450
0.4762
0.4018
Appendix
Table 22: Complete performance metrics (continued): MPDD-2025 Young. Metric definitions, aggregation, and formatting follow the Elder part. Flex-MoE cells report the means of three runs. Avg. Rank covers all ten MPDD-2025 tasks across Elder and Young, using the official task score once per task. Ours W. Acc. was not included in the available main-run summary and is left unreported.
Method
Macro-F1 ↑
Gemini 2.5 Flash
0.3236
MedAgent-Pro
0.7080
MERID (Ours)
0.7593
Appendix
Table 23: Kaggle raw-video emotion classification. Best and second-best scores are bold and underlined. MERID uses actor-disjoint nested cross-validation with review disabled.
Track
Window/task
Audio
Video
Max. length
Batch
Learning rate
Epochs
Elder
1-s binary
MFCC
OpenFace
26
1
2×10−5
200
Elder
1-s ternary
OpenSMILE
ResNet
26
2
4.5835899×10−6
400
Elder
1-s quinary
OpenSMILE
DenseNet
26
1
1.6570181×10−5
400
Elder
5-s binary
OpenSMILE
ResNet
5
2
1.8×10−5
200
Elder
5-s ternary
wav2vec
OpenFace
5
16
1.4797182×10−4
400
Elder
5-s quinary
MFCC
ResNet
5
16
9.8×10−5
200
Appendix
Table 24: Task-specific settings for the official MPDD 2025 reproduction.
Stage
Operation and retained consequence
Review
Roles analyze current feedback, prior outcomes, and active memory. The Chair issues a revision request or ends review.
Construct
The proposer turns a supported request into recipe or predictor additions. Source exclusions remove dependent candidates before reselection.
Evaluate
CPE evaluates untried pairs under the fixed protocol and remaining candidate budget. Each trial records predictions, its score, or an execution error.
Verify
The configured selection rule determines incumbent replacement. The round record retains the request and measured outcome.
Learn
Distillation proposes conditional lessons. Reconciliation adds, revises, or retires memory entries using the recorded evidence.
Re-enter
The next round receives the expanded space, retained pipeline, revision history, and updated memory. The loop ends on Chair acceptance or a budget limit.
Appendix
Table 25: Execution order of a memory-enabled improvement round. Learning also records nonpromoting and failed revisions.
Field
Meaning
id
Stable identity used to revise or retire an existing entry.
condition
Data and pipeline conditions under which the lesson applies.
lesson
An interpretation of the recorded experiment, marked as model synthesis.
next_action
A proposed experiment for extending or checking the interpretation.
evidence_ids
References to measured candidate records, round outcomes, or diagnostics.
status , versions
Active or retired status, with the previous contents retained after each update.
Appendix
Table 26: Research-memory fields. Numeric evidence remains attached to the underlying execution records.
Depression is a major mental disorder for which diagnosis relies primarily on clinical assessments. Automated methods to support its detection via the psychiatric MADRS scale are getting more and more attention. While existing solutions primarily focus on detecting the disorder from different text sources (e.g., online text, social media), there is still limited support for clinical trials, where clinical assessments are conducted through structured interviews based on standard guidelines such as SIGMA. In this work, we develop a LLM pipeline specifically designed to support clinicians in supporting the assessment of depression in patients enrolled in clinical trials. Our pipeline converts audio interviews into transcripts, maps them into the ten MADRS symptom items, estimates their severity, and identify problematic clinical ratings associated with them. Evaluation on real clinical interviews shows a strong overall correlation of 0.867 with expert ratings, providing interpretable support for future assessments in clinical trials.
Deep learning-based Major Depressive Disorder (MDD) detection using Electroencephalography (EEG) is fundamentally constrained by the "small-sample dilemma." Prevailing generative data augmentation methods not only incur heavy computational overhead but also risk introducing synthetic noise, thereby blurring classification boundaries. To challenge the traditional "data quantity first" convention, we propose a novel framework "Beyond Augmentation": Score-Guided Classification (SGC). SGC does not synthesize pseudo-samples; instead, it utilizes an unsupervised generative network architecture to model the structural and statistical anomaly degrees of samples, serving as the core "Pathological Prior". This prior, after robust normalization, is explicitly fused with deep feature representations, thereby precisely guiding the classifier's decision boundary. Furthermore, to dynamically adapt to varying channel configurations, we propose a Cross-Channel Spatial Adaptation module, utilizing a spatial mapping mechanism to effectively resolve the hardware heterogeneity of mismatched channels in multi-center datasets. Extensive experiments on the Mumtaz2016 and high-density MODMA datasets demonstrate the effectiveness and exceptional generalizability of our method under the challenging "zero data augmentation" setting and at "zero sample synthesis cost". Keywords: Electroencephalography (EEG), Depression Detection, Anomaly Score, Diffusion Models, Few-Shot Learning
Xiaojing Chen, Jingqi Cheng, Xu Zhao +2
School of Computer Science and Technology, Hefei University of Technology Hefei, China · School of Computer Science and Information Engineering, Hefei University of Technology Hefei, China
Audio--visual recordings provide complementary cues for estimating depression severity, but their informativeness varies across time and modalities. Point predictions alone do not express the uncertainty associated with these estimates. We present EviDep, a multimodal evidential regression framework that integrates multi-scale temporal modeling and shared--private representation learning for uncertainty-aware depression estimation. Frequency-aware Feature Extraction decomposes behavioral feature sequences into multiple frequency bands and refines them with scale-specific experts. Disentangled Evidential Learning encourages the disentanglement of cross-modal shared and modality-specific information in the refined features. Multi-branch Evidential Regression maps the resulting shared and private representations to three Normal-Inverse-Gamma (NIG) outputs and uses evidence-weighted aggregation to estimate depression severity and quantify aleatoric and epistemic uncertainty. Experiments on AVEC 2013, AVEC 2014, DAIC-WOZ, and E-DAIC show competitive prediction accuracy, with ablation studies supporting the contributions of frequency-aware refinement and shared--private disentanglement. Further analyses show that estimated epistemic uncertainty helps identify higher-error predictions, while both uncertainty estimates generally increase under controlled feature degradation.
Fangyuan Liu, Sirui Zhao, Yangsong Zhang +5
School of Computer Science and Technology, University of Science and Technology of China, Hefei, China · aSchool of Computer Science and Technology, University of Science and Technology of China, Hefei, China · bSchool of Computer Science and Technology, Laboratory for Brain Science and Artificial Intelligence, Southwest University of Science and Technology, Mianyang, China +1