MERID: Multimodal Exploration via Recursive Self-Improvement Agents for Major Depression Analysis
Organizations: Yale University · Zhejiang University · Northwestern University · University of California, Los Angeles · Tsinghua University · Stevens Institute of Technology
Abstract
Major depressive disorder (MDD) severely impacts daily activities and quality of life. Detecting MDD involves multimodal data, such as interview recordings and sensor measurements. This is particularly challenging, as these heterogeneous modalities often demand distinct, customized prediction pipelines. Existing efforts to address this challenge have explored both manually engineered multimodal architectures and agent-assisted pipeline development. Despite their progress, it remains challenging to autonomously revise pipelines based on experimental feedback and carry verified improvements forward into subsequent designs. To this end, we propose Multimodal Exploration via Recursive Self-Improvement Agents for Major Depression Analysis (MERID). The framework develops depression pipelines through experience-based recursive self-improvement (RSI). Grounded State Construction (GSC) grounds experience by aligning multimodal records with subject-level depression targets. Coupled Pipeline Exploration (CPE) jointly modifies representations, fusion, and predictors to build successor pipelines for classification and severity estimation. Evidence-Guided Evolution (EGE) guides revisions through feedback and verifies gains under uncertainty in small depression cohorts before inheritance. Extensive experiments on depression benchmarks show that MERID achieves the best results on multiple tasks compared with multimodal and agent-based baselines. Further analysis highlights the value of acoustic and linguistic cues for depression detection. Our code is available at https://github.com/DiscoAILab/MERID
Figures & tables
| Dataset | E-DAIC | MPDD-2025 | Mental Health | Kaggle | |||
| Method | Binary F1 | PHQ-8 CCC | Elder Track | Young Track | Status F1 | A–V F1 | Avg. rank |
| MulT | 0.5221 | -0.0047 | 0.4549 | 0.3952 | — | — | 4.50 (4) |
| Flex-MoE | 0.5993 | 0.0135 | 0.5440 | 0.4483 | 0.9548 | 0.6201 | 2.67 (6) |
| MCMoE | 0.4105 | 0.0000 | 0.4511 | 0.3419 | n/a | n/a | 5.75 (4) |
| Gemini 2.5 Flash | 0.3978 | -0.0917 | 0.4322 | 0.3995 | 0.0843 | 0.2966 | 5.33 (6) |
| MDAgents | 0.4697 | 0.0520 | 0.2096 | 0.4069 | 0.0607 | 0.3845 | 4.33 (6) |
Appendix figures & tables33 assets
Supplementary material from the paper’s appendix.
Appendix
| Aspect | DepressionAgent (text branch) | MERID |
|---|---|---|
| Evidence | Interview statements and supporting or countervailing interpretations | Task constraints, candidate scores, prediction diagnostics, and execution checks |
| Revision target | The current participant’s risk judgment | Representation–fusion recipes and predictor configurations |
| Depression-specific action | Re-read the transcript for overlooked risk evidence | Check clinical-scale exclusion and propose an ordinal severity predictor |
| Verification | Reconsider evidence through review and re-evaluation | Evaluate proposed pipelines under a shared protocol and replacement margin |
| Retained consequence | A final case judgment and its reasoning record | A reusable pipeline, admitted components, and experience for the next round |
| Task | Metric | Seed 0 | Seed 1 | Seed 2 | Seed 3 | Seed 4 | Mean SD |
|---|---|---|---|---|---|---|---|
| E-DAIC binary | Macro-F1 | 0.6655 | 0.6937 | 0.6937 | 0.6655 | 0.6655 | 0.6768 0.0154 |
| E-DAIC PHQ-8 | CCC | 0.3201 | 0.3201 | 0.3538 | 0.3201 | 0.3201 | 0.3268 0.0151 |
| MPDD-2025 Elder | Track score | 0.5820 | 0.6071 | 0.6177 | 0.6282 | 0.6335 | 0.6137 0.0204 |
| MPDD-2025 Young | Track score | 0.4116 | 0.3545 | 0.3610 | 0.3755 | 0.3594 | 0.3724 0.0233 |
| Mental Health | Macro-F1 | 0.9704 | 0.9704 | 0.9630 | 0.9527 | 0.9710 | 0.9655 0.0079 |
| Kaggle A–V | Macro-F1 | 0.7811 | 0.7624 | 0.7803 | 0.8074 | 0.7869 | 0.7836 0.0162 |
| Task | Metric | Opus 4.8 | GPT-6 Astra | |
|---|---|---|---|---|
| E-DAIC binary | Macro-F1 | 0.6768 0.0154 | 0.7057 0.0118 | +0.0289 |
| E-DAIC PHQ-8 | CCC | 0.3268 0.0151 | 0.3201 0.0000 | -0.0067 |
| MPDD-2025 Elder | Track score | 0.6137 0.0204 | 0.6148 0.0291 | +0.0011 |
| MPDD-2025 Young | Track score | 0.3724 0.0233 | 0.3714 0.0261 | -0.0010 |
| Mental Health | Macro-F1 | 0.9655 0.0079 | 0.9565 0.0042 | -0.0090 |
| Kaggle A–V | Macro-F1 | 0.7836 0.0162 | 0.6070 0.0000 | -0.1766 |
| Quantity | Opus 4.8 | GPT-6 Astra |
|---|---|---|
| Run attempts for 80 tasks | 80 | 85 |
| Aborted attempts | 0 | 5 |
| MPDD-2026 proposals audited | 112 | 135 |
| invalid (run aborts) | 0 | 5 (3.7%) |
| duplicates (re-asked once) | 0 | 0 |
| Chair replies re-asked for format | 35 | 0 |
| Task | Seed | Att. | Prop. | Failure | Offending fragment |
|---|---|---|---|---|---|
| MPDD-2026 Elder | 0 | 1 | 7 | out-of-space value | "sev_fusion": "wav2emb:densenet" |
| MPDD-2026 Elder | 2 | 1 | 6 | out-of-space value | "audio": "wav2span" |
| MPDD-2026 Elder | 2 | 2 | 1 | out-of-space value | "audio": "wav2.0" |
| MPDD-2026 Elder | 4 | 1 | 6 | malformed JSON | gp_stats_pz","model":"plsn4":"bad"}} |
| MPDD-2026 Young | 2 | 1 | 9 | malformed JSON | 1.0","sev_model":"gb0.05d1,"sev_fusion": |
| Task / test metric | w/o CPE | w/o EGE | MERID (full) |
| E-DAIC binary (F1) | 0.5151 0.0760 | 0.5162 0.0345 | 0.6768 0.0154 |
| E-DAIC PHQ-8 (CCC) | 0.0361 0.0452 | 0.2997 0.0196 | 0.3268 0.0151 |
| Mental Health (F1) | 0.9672 0.0149 | 0.9530 0.0005 | 0.9655 0.0079 |
| MPDD-25 Elder 1 s binary | 0.7231 0.0846 | 0.7850 0.0042 | 0.7593 0.0349 |
| MPDD-25 Elder 1 s ternary | 0.6089 0.0822 | 0.5955 0.0667 | 0.6295 0.0673 |
| MPDD-25 Elder 1 s quinary | 0.4583 0.0341 | 0.4472 0.0482 | 0.4738 0.0647 |
| Binary classification | ||||
|---|---|---|---|---|
| Available | Dev Macro-F1 | Test Macro-F1 | Test Accuracy | Deployed |
| A | 0.6348 0.0000 | 0.5336 0.0679 | 0.5655 0.0412 | A |
| V | 0.5754 0.0476 | 0.6086 0.0092 | 0.6488 0.0103 | V |
| T | 0.6750 0.0531 | 0.6107 0.0535 | 0.6607 0.0309 | T |
| AV | 0.6348 0.0000 | 0.5612 0.1064 | 0.6726 0.0103 | A+V |
| AT | 0.7348 0.0000 | 0.6468 0.0153 | 0.6964 0.0179 | A+T |
| Selection rule | Development | Test |
|---|---|---|
| Incumbent one standard error | 0.7172 | |
| Classic one-standard-error rule | 0.7737 | 0.4991 |
| Maximum development score | 0.7737 | 0.4876 |
| Task / metric | Reference one SE | Winner’s-curse configuration | Search selection (one SE) |
|---|---|---|---|
| E-DAIC binary (Macro-F1) | 0.6937 | 0.6814 | 0.5492 |
| E-DAIC PHQ-8 (CCC) | 0.3201 | 0.3201 | 0.3201 |
| Mental Health (Macro-F1) | 0.9630 | 0.9535 | 0.9535 |
| MPDD-25 Elder 1 s binary | 0.7212 | 0.7212 | 0.7212 |
| MPDD-25 Elder 1 s ternary | 0.6842 | 0.6842 | 0.6220 |
| MPDD-25 Elder 1 s quinary | 0.3962 | 0.5490 | 0.4011 |
| Task | Full | Staged | Fixed recipe | Fixed predictor | 1-SE | Random predictor | Random recipe |
|---|---|---|---|---|---|---|---|
| E-DAIC binary | 0.6937 | 0.4991 | 0.6814 | 0.4593 | 0.4991 | 0.5564 | 0.5735 |
| E-DAIC PHQ-8 | 0.3201 | 0.3201 | 0.3201 | 0.3201 | 0.3201 | 0.3307 | 0.3201 |
| Mental Health | 0.9630 | 0.9711 | 0.9527 | 0.9305 | 0.9535 | 0.9527 | 0.9539 |
| MPDD-25 binary | 0.7212 | 0.7543 | 0.7212 | 0.7212 | 0.7565 | 0.7266 | — |
| MPDD-25 ternary | 0.6390 | 0.6443 | 0.6922 | 0.6482 | 0.6183 | 0.6619 | — |
| MPDD-25 quinary | 0.4115 | 0.4066 | 0.4514 | 0.4317 | 0.4169 | 0.4624 | — |
| Task | Memory off | Evolving | Calls (off evolving) | |
|---|---|---|---|---|
| E-DAIC binary (F1) | 0.4927 | 0.5099 | +0.0172 | 19.7 33.7 |
| E-DAIC PHQ-8 (CCC) | 0.3142 | 0.3311 | +0.0168 | 28.7 29.3 |
| Mental Health (F1) | 0.9647 | 0.9582 | -0.0065 | 27.0 33.0 |
| MPDD-25 Elder 1 s binary | 0.7802 | 0.7886 | +0.0084 | 22.0 18.0 |
| MPDD-25 Elder 1 s ternary | 0.6557 | 0.6335 | -0.0222 | 18.0 39.7 |
| MPDD-25 Elder 1 s quinary | 0.4164 | 0.4164 | +0.0000 | 22.0 25.7 |
| Memory | Step | Runs | New evaluations | Retained-score gains |
|---|---|---|---|---|
| Off | 1 | 23 | 751 | 2 |
| Off | 2 | 5 | 107 | 0 |
| Off | 3 | 3 | 52 | 0 |
| Evolving | 1 | 22 | 727 | 3 |
| Evolving | 2 | 12 | 270 | 0 |
| Evolving | 3 | 8 | 157 | 0 |
| Task / test metric | Evolving random | Evolving no memory |
|---|---|---|
| E-DAIC binary / Macro-F1 | [-0.0176, 0.0185] | [-0.0206, 0.0173] |
| E-DAIC PHQ-8 / CCC | [-0.0023, 0.0071] | [0.0013, 0.0079] |
| MPDD-25 Elder 1 s ternary / official | [-0.0093, 0.0245] | [-0.0121, 0.0125] |
| Task | Seed | Add / revise / retire | Citation rate | Action-matched use |
|---|---|---|---|---|
| E-DAIC binary | 0 | 14 / 12 / 2 | 0.73 | 1 / 7 |
| E-DAIC binary | 1 | 11 / 15 / 2 | 0.82 | 6 / 22 |
| E-DAIC binary | 2 | 17 / 10 / 1 | 0.72 | 6 / 21 |
| E-DAIC PHQ-8 | 0 | 14 / 15 / 0 | 0.80 | 1 / 5 |
| E-DAIC PHQ-8 | 1 | 14 / 12 / 1 | 0.66 | 7 / 33 |
| E-DAIC PHQ-8 | 2 | 8 / 6 / 2 | 0.81 | 6 / 14 |
| Task | Arm | Seed | Reviews | Failed | Requested / declined | New rounds | Candidates |
|---|---|---|---|---|---|---|---|
| E-DAIC binary | Memory | 0 | 8 | 1 | 3 / 1 | 6 | 448 |
| E-DAIC binary | Memory | 1 | 8 | 1 | 1 / 0 | 7 | 512 |
| E-DAIC binary | Memory | 2 | 8 | 1 | 2 / 1 | 6 | 512 |
| E-DAIC binary | No memory | 0 | 8 | 0 | 6 / 0 | 8 | 512 |
| E-DAIC binary | No memory | 1 | 8 | 1 | 4 / 0 | 7 | 480 |
| E-DAIC binary | No memory | 2 | 8 | 1 | 5 / 0 | 7 | 480 |
| Task | MERID (full) | w/o CPE | w/o EGE | Runs |
|---|---|---|---|---|
| E-DAIC binary | 11.4 / $0.60 | 83.0 / $4.67 | 2.0 / $0.16 | 5 / 1 / 3 |
| E-DAIC PHQ-8 | 11.4 / $0.61 | 41.0 / $2.12 | 1.0 / $0.06 | 5 / 1 / 3 |
| Mental Health | 10.8 / $0.64 | 89.0 / $5.09 | 0.0 / $0.00 | 5 / 1 / 3 |
| MPDD-25 Elder 1 s binary | 12.2 / $0.65 | 89.0 / $5.29 | 1.0 / $0.01 | 5 / 1 / 1 |
| MPDD-25 Elder 1 s ternary | 12.4 / $0.70 | 20.0 / $1.06 | 1.0 / $0.01 | 5 / 1 / 1 |
| MPDD-25 Elder 1 s quinary | 12.2 / $0.69 | 33.0 / $1.89 | 1.0 / $0.01 | 5 / 1 / 1 |
| Memory | Runs | Cap | Total mean | Total range | Revision mean |
|---|---|---|---|---|---|
| Memory off | 27 | 320 | 264.4 | 139–320 | 33.7 |
| Evolving memory | 27 | 320 | 272.9 | 139–320 | 42.7 |
| Task | Arm | Cap | Initial mean | Total mean | Total range |
|---|---|---|---|---|---|
| E-DAIC binary | Memory | 512 | 254.7 | 490.7 | 448–512 |
| E-DAIC binary | No memory | 512 | 256.0 | 490.7 | 480–512 |
| E-DAIC binary | Random | 512 | 256.0 | 512.0 | 512–512 |
| E-DAIC PHQ-8 | Memory | 512 | 177.3 | 512.0 | 512–512 |
| E-DAIC PHQ-8 | No memory | 512 | 182.0 | 512.0 | 512–512 |
| E-DAIC PHQ-8 | Random | 512 | 186.7 | 512.0 | 512–512 |
| Dataset | Cohort | Task | Metric | Official | MulT | Flex-MoE | Gemini 2.5 Flash | MDAgents | Ours |
| E-DAIC | – | Binary | Accuracy | 0.5893 | 0.7143 | 0.6607 | 0.6429 | 0.7393 | |
| Bal. Acc. | 0.5226 | 0.5958 | 0.4744 | 0.4947 | 0.6702 | ||||
| Macro-F1 | 0.5221 | 0.5993 | 0.3978 | 0.4697 | 0.6768 | ||||
| W. F1 | 0.5925 | 0.6836 | 0.5541 | 0.5887 | 0.7326 | ||||
| 0.0445 | 0.2209 | -0.0683 | -0.0127 | 0.3554 | |||||
| PHQ-8 | RMSE | 6.9902 | 7.0394 | 7.7113 | 6.6201 | 5.9098 |
| Dataset | Cohort | Task | Metric | Official | MulT | Flex-MoE | Gemini 2.5 Flash | MDAgents | Ours |
| MPDD-2025 | Elder | 1 s Binary | Accuracy | 0.6916 | 0.7841 | 0.7416 | 0.7357 | 0.1982 | 0.7789 |
| Macro-F1 | 0.5900 | 0.6318 | 0.6266 | 0.5048 | 0.1809 | 0.6964 | |||
| W. Acc. | 0.6360 | 0.6332 | 0.6597 | 0.5055 | 0.5035 | — | |||
| W. F1 | 0.7222 | 0.7852 | 0.7604 | 0.7238 | 0.1038 | 0.7987 | |||
| Task Acc. | 0.6638 | 0.7086 | 0.7006 | 0.6206 | 0.3509 | 0.7712 | |||
| Task F1 | 0.6561 | 0.7085 | 0.6935 | 0.6143 | 0.1424 | 0.7475 |
| Dataset | Cohort | Task | Metric | Official | MulT | Flex-MoE | Gemini 2.5 Flash | MDAgents | Ours |
| MPDD-2025 | Young | 1 s Binary | Accuracy | 0.5644 | 0.4924 | 0.6048 | 0.4848 | 0.5455 | 0.4038 |
| Macro-F1 | 0.5446 | 0.4491 | 0.5749 | 0.3450 | 0.4762 | 0.4018 | |||
| W. Acc. | 0.5644 | 0.4924 | 0.6048 | 0.4848 | 0.5455 | — | |||
| W. F1 | 0.5446 | 0.4491 | 0.5749 | 0.3450 | 0.4762 | 0.4018 | |||
| Task Acc. | 0.5644 | 0.4924 | 0.6048 | 0.4848 | 0.5455 | 0.4038 | |||
| Task F1 | 0.5446 | 0.4491 | 0.5749 | 0.3450 | 0.4762 | 0.4018 |
| Method | Macro-F1 |
|---|---|
| Gemini 2.5 Flash | 0.3236 |
| MedAgent-Pro | 0.7080 |
| MERID (Ours) | 0.7593 |
| Track | Window/task | Audio | Video | Max. length | Batch | Learning rate | Epochs |
| Elder | 1-s binary | MFCC | OpenFace | 26 | 1 | 200 | |
| Elder | 1-s ternary | OpenSMILE | ResNet | 26 | 2 | 400 | |
| Elder | 1-s quinary | OpenSMILE | DenseNet | 26 | 1 | 400 | |
| Elder | 5-s binary | OpenSMILE | ResNet | 5 | 2 | 200 | |
| Elder | 5-s ternary | wav2vec | OpenFace | 5 | 16 | 400 | |
| Elder | 5-s quinary | MFCC | ResNet | 5 | 16 | 200 |
| Stage | Operation and retained consequence |
|---|---|
| Review | Roles analyze current feedback, prior outcomes, and active memory. The Chair issues a revision request or ends review. |
| Construct | The proposer turns a supported request into recipe or predictor additions. Source exclusions remove dependent candidates before reselection. |
| Evaluate | CPE evaluates untried pairs under the fixed protocol and remaining candidate budget. Each trial records predictions, its score, or an execution error. |
| Verify | The configured selection rule determines incumbent replacement. The round record retains the request and measured outcome. |
| Learn | Distillation proposes conditional lessons. Reconciliation adds, revises, or retires memory entries using the recorded evidence. |
| Re-enter | The next round receives the expanded space, retained pipeline, revision history, and updated memory. The loop ends on Chair acceptance or a budget limit. |
| Field | Meaning |
|---|---|
| id | Stable identity used to revise or retire an existing entry. |
| condition | Data and pipeline conditions under which the lesson applies. |
| lesson | An interpretation of the recorded experiment, marked as model synthesis. |
| next_action | A proposed experiment for extending or checking the interpretation. |
| evidence_ids | References to measured candidate records, round outcomes, or diagnostics. |
| status , versions | Active or retired status, with the previous contents retained after each update. |