cs.CLSep 5, 2026

SurveyAgent-HKA: A multi-agent framework for scientific survey generation with LLMs and human knowledge augmentation

Authors: Tong Bao, Mir Tafseer Nayeem, Yi Zhao, Davood Rafiei, Chengzhi Zhang

Organizations: Department of Information Management, Nanjing University of Science and Technology, Nanjing, China · Department of Computing Science, University of Alberta, Edmonton, Canada · School of Management, Anhui University, Hefei, China

Abstract

Automatic scientific survey generation has become an important task in scientific document processing. The common approach of retrieving literature from a single source (e.g., arXiv) and generating surveys through a one-pass large language model (LLM) call often leads to limited reference coverage and, more importantly, fails to replicate the expert-driven revision process that is crucial for writing high-quality surveys. In this paper, we introduce SurveyAgent-HKA, a multi-agent framework that improves end-to-end scientific survey generation by incorporating knowledge derived from published surveys and peer-review comments. The framework decomposes survey generation into well-defined sub-tasks handled by LLM-powered agent. It first retrieves relevant papers from multiple sources and identifies key topics through clustering to construct an initial outline, which is then refined using outlines from related human-written surveys. Based on the refined outline, topic-focused papers are retrieved and re-ranked to select for drafting a well-grounded survey. Then, we identify common issues raised by experts in peer-review comments from published surveys to guide the revisions and finalize the survey. Experiments on two domains show that our approach outperforms mainstream baselines in citation quality, structural consistency, and content quality. Furthermore, our framework is efficient in both time and cost, making it a practical solution for broader AI-assisted scientific writing applications.

Explore similar work

May 28, 2026cs.AI

DeepSurvey: Agent-Oriented Automated Survey Generation with Analytical Depth and Citation Reliability

As scientific literature grows rapidly and research increasingly involves AI agents, automated survey generation has become a key capability for both agents and human researchers. For such agents, a survey serves as a primary knowledge source of a field prior to research. Since pretrained models already encode broad knowledge of established work, these consumers benefit more from analytical depth than from breadth alone; moreover, unsupported claims, once ingested as knowledge, can propagate into downstream research. However, existing systems tend to overemphasize coverage and presentation; they suffer from limited analytical depth due to reliance on abstracts and isolated paper processing, and from unreliable citations due to imprecise retrieval and post-hoc grounding. We present DeepSurvey, an agentic generation system that addresses both limitations. To enhance depth, DeepSurvey extracts structured keynotes, models cross-paper relationships through clustering and comparative analysis, and integrates a code-agent subsystem to recover implementation-level details. To fortify reliability, it combines citation-graph expansion with hybrid filtering for topic-focused retrieval, enforces evidence-constrained analysis and writing, and deploys multi-granularity agentic refinement to validate citation--claim alignment. Experiments show that DeepSurvey achieves the highest content score (8.34/10) and citation quality (recall and precision gains of 25.3% and 35.2% over the strongest baseline), generalizes more robustly across domains, and is preferred by domain experts over human-written surveys (83.3% in overall quality, 100% in content depth). Moreover, when a coding agent uses a survey as its only literature source, the agent equipped with DeepSurvey achieves the best performance among human-written and baseline-generated surveys.
Aug 18, 2026cs.CV

Deep Academic Survey: Stateful Agentic Closed-Loop Paradigm for Academic Survey Automation

Academic surveys play a central role in organizing rapidly expanding scholarly literature, yet their construction requires extensive paper analysis, coherent knowledge organization, fine-grained citation support, and reliable manuscript assembly. Existing Deep Research and automated survey generation systems address parts of this process, but typically do not coordinate paper understanding, literature organization, evidence-grounded drafting, and manuscript validation through a shared, revisable state. We introduce DAS, a stateful agentic framework for generating publication-oriented academic surveys. Its key idea is to separate reusable paper analysis from topic-specific manuscript construction. DAS builds on DAS-2M, a dynamically updated metadata lake containing survey-oriented representations of approximately two million papers. Its agents maintain explicit literature, organization, writing, and finalization states through candidate-grounded taxonomy planning, reverse paper-to-section routing, and hierarchical claim and citation planning. Semantic review reactivates only the affected writing states for repair and reevaluation, forming a scoped closed loop with deterministic validation. We further introduce DAS-Bench, a 30-topic benchmark, together with DAS-Eval, which assesses scholarly citation quality, taxonomic synthesis, hierarchical discourse, and manuscript assembly reliability through 16 criteria. Among systems evaluated on all 30 topics, DAS achieves the highest average in all four dimensions, with an overall score of 4.34 compared with 4.03 for the strongest competitor, and the same ordering is preserved on the matched 21-topic CS subset. Blinded expert evaluation further prefers DAS to Naive RAG on 27 of 30 topics and to AutoSurvey on 19 of 21 shared CS topics. The project page is available at https://zhikaixu24.github.io/projects/DAS/.
Jun 11, 2026cs.CL

From Passive Generation to Investigation: A Proactive Scientific Peer Review Agent

Large language models (LLMs) have shown promise in automating scientific peer review. However, existing approaches often struggle to generate in-depth reviews supported by concrete evidence. We argue that a key limitation is the lack of flexibility to proactively investigate suspicious parts of a paper based on accumulated evidence, as human reviewers do. In this paper, we explore how to enable an LLM-based review agent to perform such proactive investigation. We find that this can be naturally formulated as a Markov Decision Process (MDP), and propose ProReviewer, a scientific peer review agent that proactively reviews a paper guided by a maintained, structured review log. The structured review log serves as a workspace for the agent to track evidence and intermediate findings collected during review. Experiments show that ProReviewer with an 8B backbone, trained by supervised fine-tuning and optimized by reinforcement learning, achieves the highest average score across five quality dimensions, outperforming prompt-based methods with much larger frontier LLMs by up to 39% and the strongest fine-tuned baseline by 16% relatively. It also attains the highest win rates against baselines in human evaluation.