LLM-based Agentic Reasoning Frameworks: A Survey from Methods to Scenarios
Abstract
Recent advances in LLM-based agents highlight the importance of their reasoning frameworks, which guide the problem-solving process in diverse ways. This survey introduces a unified formal language to systematically categorize these frameworks at three compositional levels: single-agent, tool-based, and multi-agent methods. Following our taxonomy, we review key application scenarios across scientific discovery, healthcare, software engineering, society, economics, and general-purpose tasks. It also compares the distinct features and evaluation strategies of each category. Through our taxonomy and comparisons, our survey explores the designs and strengths of LLM-based agentic frameworks in different scenarios, reviewing the fast-paced development of complex agentic systems in the real world.
Figures & tables
| Evidence source | How evidence enters reasoning | Works |
| Literature and reference knowledge | Retrieved passages supply reference knowledge to the reasoning context. | MOOSE-Chem ( Yang et al., 2025d ) ; Stewart et al. ( Stewart and Buehler, 2025 ) ; BioAgents ( Mehandru et al., 2025 ) ; CRISPR-GPT ( Qu et al., 2026 ) ; BioResearcher ( Luo et al., 2025b ) ; DrugAgent ( Inoue et al., 2024 ) . |
| Scientific records and knowledge graphs | Queried records and graph relations are combined as task-specific evidence. | ChatDrug ( Liu et al., 2024b ) , CLADD ( Lee et al., 2026a ) , DrugAgent ( Inoue et al., 2024 ) ; Chemist-X ( Chen et al., 2023 ) ; ChatMOF ( Kang and Kim, 2024 ) , LLaMP ( Chiang et al., 2025 ) . |
| Computational and simulation outputs | Tool outputs provide observations for revising intermediate reasoning and actions. | LIDDIA ( Averly et al., 2025 ) , DrugAgent ( Inoue et al., 2024 ) ; ChatChemTS ( Ishida et al., 2025 ) ; ChemCrow ( M. Bran et al., 2024 ) , CACTUS ( McNaughton et al., 2024 ) ; El Agente Q ( Zou et al., 2025 ) ; AtomAgents ( Ghafarollahi and Buehler, 2025a ) ; ProtAgents ( Ghafarollahi and Buehler, 2024 ) , TourSynbio-Search ( Liu et al., 2024d ) . |
| Execution and workflow feedback | Intermediate results and execution checks guide revision and stage handoffs. | AutoBA ( Zhou et al., 2024a ) ; BioMaster ( Su et al., 2026b ) ; ChemHTS ( Li et al., 2025f ) ; PharmAgents ( Gao et al., 2025b ) , Ruan et al. ( Ruan et al., 2024b ) . |
| Previous experimental observations | Retained observations inform reasoning in subsequent research rounds. | BioDiscoveryAgent ( Roohani et al., 2025 ) . |
| Metrics Level | ||
| Focus | Related Work | Metrics |
| Lab-involved Biomedical Research | BioResearcher ( Luo et al., 2025b ) | Completeness, Level of Detail, Correctness, Logical Soundness, Structural Soundness |
| High-fidelity Materials Knowledge Retrieval | LLaMP ( Chiang et al., 2025 ) | Precision, coefficient of precision, confidence, self-consistency of response, MAE |
| Drug Discovery | LIDDIA ( Averly et al., 2025 ) | drug-likeness ( Bickerton et al., 2012 ) , Lipinski’s Rule of Five ( Lipinski et al., 1997 ) , synthetic accessibility ( Ertl and Schuffenhauer, 2009 ) , binding affinities ( Trott and Olson, 2010 ) |
| Benchmark/Dataset Level | ||
| Focus | Related Work | Benchmark/Dataset |
| Metrics Level | ||||
| Focus | Related Work | Metrics | ||
| AutoSurvey ( Wang et al., 2024h ) | Survey Creation Speed, Content Quality | |||
| Gao et al. ( Gao et al., 2023b ) | Citation Quality | |||
| SurveyX ( Liang et al., 2025 ) | Insertion over Union, semantic-based reference relevance | |||
| SurveyForge ( Yan et al., 2025 ) | Reference, Outline, and Content Quality (SAM Metrics) | |||
| Agent Laboratory ( Schmidgall et al., 2025 ) | Inference cost, Inference time, Success Rate | |||
| Benchmark/Dataset Level | |||
| Focus | Related Work | Benchmark/Dataset | |
| Clinical Consultation Flow | Wang et al. ( Wang et al., 2025e ) | MVME ( Fan et al., 2025 ) | |
| Zero-shot Medical Reasoning | MedAgents ( Tang et al., 2024 ) | Jin et al. ( Jin et al., 2021 ) ,Pal et al. ( Pal et al., 2022 ) ,PubMedQA ( Jin et al., 2019 ) , Hendrycks et al. ( Hendrycks et al., 2021a ) | |
| Automated supervision of Healthcare Safety | TAO ( Kim et al., 2025b ) | Safetybench ( Zhang et al., 2024i ) ,Medsafetybench ( Han et al., 2024b ) , Chang et al. ( Chang et al., 2025 ) , Hu et al. ( Hu et al., 2024a ) ,Wang et al. ( Wang et al., 2025f ) | |
| Evolvable Medical Agents | MDAgents ( Kim et al., 2024c ) | MedQA ( Jin et al., 2021 ) , PubMedQA ( Jin et al., 2019 ) ,MedBullets ( Chen et al., 2025a ) , JAMA ( Chen et al., 2025a ) ,DDXPlus ( Tchango et al., 2022 ) ,SymCat ( Al-Ars et al., 2023 ) , PathVQA ( He et al., 2020 ) ,PMC-VQA ( Zhang et al., 2024d ) ,MedVidQA ( Gupta et al., 2023 ) ,MIMIC-CXR ( Bae et al., 2023 ) | |
| Healthcare Intent Awareness | MedAide ( Yang et al., 2026 ) | Pre-Diagnosis, Diagnosis, Medicament, Post-Diagnosis Bench ( Yang et al., 2026 ) | |
| Method / model | Source | HE ( Chen et al., 2021 ) | HE-ET ( Dong et al., 2025 ) | HE+ ( Liu et al., 2023a ) | MBPP ( Austin et al., 2021 ) | MBPP-ET ( Dong et al., 2025 ) | MBPP+ ( Liu et al., 2023a ) | DS-1000 ( Lai et al., 2023 ) | LCB S112 ( Jain et al., 2025 ) | LCB V1 ( Jain et al., 2025 ) |
| Direct baselines | ||||||||||
| GPT-3.5 | ( Huang et al., 2023b ) | 57.3 | 42.7 | – | 52.2 | 36.8 | – | – | – | – |
| GPT-3.5 | ( Islam et al., 2024 ) | 48.1 | 37.2 | 66.5 | 49.8 | 37.7 | – | – | – | – |
| GPT-4 | ( Huang et al., 2023b ) | 67.6 | 50.6 | – | 68.3 | 52.2 | – | – | – | – |
| GPT-4 | ( Islam et al., 2024 ) | 80.1 | 73.8 | 81.7 | 81.1 | 54.7 | – | – | – | – |
| GPT-4o | ( Anthropic, 2024 ) | 90.2 | – | – | – | – | – | – | – | – |
| Defects4J v1.2 ( Just et al., 2014 ) | |||
| Method | Source | Loc. | Repair |
| GPT-3.5 | |||
| AgentFL | ( Qin et al., 2025a ) | 157/395 | – |
| ChatGPT + Ochiai | ( Qin et al., 2025a ) | 121/395 | – |
| RepairAgent | ( Bouzenia et al., 2025 ) | – | 74/395 |
| SWE-bench ( Jimenez et al., 2024 ) | |||
| Related work | Simulation focus | Data source | # |
| GOVSIM ( Piatti et al., 2024 ) | Sustainable cooperation | Simulated | – |
| MetaAgents ( Li et al., 2025c ) | Job fair | Simulated | – |
| Generative Agents ( Park et al., 2023 ) | Interactive human behavior | Simulated | 25 |
| RecAgent ( Wang et al., 2025b ) | Recommendation system | Simulated | 20 |
| BotSim ( Qiao et al., 2025 ) | Malicious social botnet | Simulated | 3k |
| S3 ( Gao et al., 2023a ) | Population-level interaction | Simulated | 10k |
| Method | Model | Source | TB 2.0 ( Merrill et al., 2026 ) | TB 4.0 ( Marten, 2026 ) | GAIA ( Mialon et al., 2024 ) | BrowseComp ( Wei et al., 2025b ) | HLE ( Center for AI Safety et al., 2026 ) | xBench-DS ( Chen et al., 2025b ) | ALE ( Sun et al., 2026 ) |
| FRIDAY / OS-Copilot ( Wu et al., 2024b ) | GPT-4-turbo-1106 | ( Wu et al., 2024b ) | – | – | 40.86/20.13/6.12/- | – | – | – | – |
| AutoGen ( Wu et al., 2024a ) | GPT-4o-2024-08-06 | ( Lu et al., 2026b ) | – | – | -/-/-/6.3 § | – | – | – | – |
| OctoTools ( Lu et al., 2026b ) | GPT-4o-2024-08-06 | ( Lu et al., 2026b ) | – | – | -/-/-/18.4 § | – | – | – | – |
| AutoAgent ( Tang et al., 2026b ) | 3.5 Sonnet | ( Tang et al., 2026b ) | – | – | 71.70/53.49/26.92/55.15 | – | – | – | – |
| DeepVerifier ( Wan et al., 2026 ) | 3.7 Sonnet | ( Wan et al., 2026 ) | – | – | -/-/-/58.93 | 9.0 | – | 44.0 | – |
| Workforce / OWL ( Hu et al., 2025a ) | 3.7 Sonnet | ( Hu et al., 2025a ) | – | – | 84.91/68.60/42.31/69.70 | – | – | – | – |
| Method | Task model | Source | Split | TB * ( Merrill et al., 2026 ) | SWE-V ( OpenAI, 2024 ) | SWE-Pro ( Deng et al., 2025 ) | Gaia2-mini ( Froger et al., 2026 ) | AppWorld ( Trivedi et al., 2024 ) | ARC-AGI-2 ( ARC Prize Foundation, 2025 ) |
| GEPA ( Agrawal et al., 2026 ) | GPT-5 mini | ( Chen et al., 2026b ) | T | 17.6–19.6 | 45.3–46.7 | – | – | 36.7–37.4 | – |
| HarnessFix ( Chen et al., 2026b ) | GPT-5 mini | ( Chen et al., 2026b ) | T | 17.6–26.5 | 45.3–57.3 | – | – | 36.7–43.0 | – |
| Meta-Harness ( Lee et al., 2026b ) | GPT-5 mini | ( Chen et al., 2026b ) | T | 17.6–23.5 | 45.3–54.7 | – | – | 36.7–40.4 | – |
| ReCreate ( Hao et al., 2026 ) | GPT-5 mini | ( Chen et al., 2026b ) | T | 17.6–21.6 | 45.3–51.7 | – | – | 36.7–39.3 | – |
| SCOPE ( Pei et al., 2025 ) | GPT-5 mini | ( Chen et al., 2026b ) | T | 17.6–20.6 | 45.3–48.3 | – | – | 36.7–38.9 | – |
| HarnessFix ( Chen et al., 2026b ) | Sonnet 4.5 | ( Chen et al., 2026b ) | T | 34.3–47.1 | 61.3–71.7 | – | – | 61.1–70.4 | – |
| Method | Task model | Source | Split | GAIA ( Mialon et al., 2024 ) | HLE ( Center for AI Safety et al., 2026 ) | ALFWorld ( Shridhar et al., 2021 ) | WebShop ( Yao et al., 2022 ) |
| AFlow ( Zhang et al., 2025c ) | GPT-5-Chat | ( Cheng et al., 2026 ) | U | 18.47–19.75 | – | 86.87–93.40 | 25.10–37.90 |
| Alita ( Qiu et al., 2025b ) | GPT-5-Chat | ( Cheng et al., 2026 ) | U | 18.47–72.73 | – | 86.87–86.13 | 25.10–30.21 |
| Mem 2 Evolve ( Cheng et al., 2026 ) | GPT-5-Chat | ( Cheng et al., 2026 ) | U | 18.47–76.31 | – | 86.87–94.31 | 25.10–39.20 |
| GEPA ( Agrawal et al., 2026 ) | GPT-5 mini | ( Chen et al., 2026b ) | T | 43.3–46.7 | – | – | – |
| HarnessFix ( Chen et al., 2026b ) | GPT-5 mini | ( Chen et al., 2026b ) | T | 43.3–61.7 | – | – | – |
| Meta-Harness ( Lee et al., 2026b ) | GPT-5 mini | ( Chen et al., 2026b ) | T | 43.3–56.7 | – | – | – |
| Scenario | Task constraints | Organization | Tool use | Oversight |
| Scientific research | Computational checks; experimental feedback | Tool loops; specialist review | Retrieval and computation | Goal setting; proposal review |
| Healthcare | Incomplete evidence; high cost of error | Specialist teams; deliberation | Records and diagnostic tools | Tiered and clinician review |
| Software engineering | Artifact dependencies; executable feedback | Controllers or teams; artifact handoffs | Execution and repair search | Sandboxes; independent tests |
| Social/economic simulation | Private objectives; interacting behavior | Individuals or groups; cooperation/competition | Memory and environment tools | Behavioral validation; risk control |
| General-purpose agents | Varied tasks; uneven feedback | Planner–worker teams; verification | Search, code and reusable skills | Permissions; update gates |