cs.IR · 2605.26400 Copy arXiv ID · May 26, 2026 Save Plans for Evaluating Structured Generative Search Summaries Authors: Tetsuya Sakai , Jina Lee , Hanpei Fang , Young-In Song
Organizations: Waseda University/Naver Corporation · Tokyo, Japan · Waseda University · Naver Corporation · Seongnam, Korea
Abstract We propose a framework for evaluating structured generative search summaries that are placed atop organic web search results. A structured summary, generated by a large language model, typically consists of an overview, several sections with section titles, and a list of source documents that are cited within the summary. We then describe our plans for implementing and evaluating the framework.
Explore similar work Apr 28, 2026 · Huyen Nguyen, Haoxuan Zhang, Yang Zhang +2 Summarization Long-Form Generation
Apr 28, 2026 · Huyen Nguyen, Haoxuan Zhang, Yang Zhang +2 Summarization Large Language Model Evaluation
Sep 18, 2026 · Nikhil Reddy Pottanigari, Ramin Fahimi, Noah Bolger +2 Information Retrieval Metrics
Apr 28, 2026 · cs.CL J/K move · Enter open · S save
Huyen Nguyen, Haoxuan Zhang, Yang Zhang, Junhua Ding +1
Evaluating long document summaries remains the primary bottleneck in summarization research. Existing metrics correlate weakly with human judgments and produce aggregate scores without explaining deficiencies or guiding improvement, preventing effective refinement in applications requiring verifiable accuracy. We introduce LongSumEval, a unified framework bridging evaluation and generation through structured question-answering feedback. The framework operationalizes summary quality as answerability and factual alignment of question-answer pairs, generating interpretable scores and actionable feedback that identifies coverage gaps and factual inconsistencies. This resolves the misalignment where evaluation operates independently of generation objectives. Meta-evaluation of our QA-based evaluation module across seven benchmarks demonstrates substantially stronger agreement with human judgments compared to established metrics. Structured feedback enables significant quality improvements through self-refinement without retraining. By demonstrating that evaluation feedback can serve as executable instructions for generation, this work establishes a generalizable paradigm for aligning assessment with improvement, with direct implications for controllable text generation requiring verifiable accuracy and transparent quality control. All code and datasets will be released in GitHub for reproducibility.