Text Summarization

Momentum

8 papers in the last four weeks, up 33% on the four weeks before. 0.1% of all new papers.

Jul 13Week of Sep 28

Latest papers 97

Jan 18, 2026cs.SD

Multi-Task Instruction Tuning via Data Scheduling for Low-Resource Arabic SpeechLLMs

Audio large language models (LLMs) enable unified speech understanding and generation, but adapting them to linguistically complex and dialect-rich settings such as Arabic-English remains challenging. We present a controlled study of multi-task instruction tuning for an Arabic-centric audio LLM across generative tasks, including automatic speech recognition (ASR) and speech and text summarization, as well as discriminative tasks, including dialect identification (DID) and speech emotion recognition (SER), in a resource-constrained setting. To support end-to-end Arabic speech summarization, we introduce AraMega-SSum, the first Arabic speech summarization dataset designed for training and benchmarking Arabic-centric audio LLMs. We compare four training strategies: (i) Uniform Mixing (UM), (ii) Task-Progressive Curriculum (TPC), (iii) Aligner-Based Diverse Sampling (ADS) for training-time batch construction, and (iv) a two-stage TPC->ADS strategy. Our results reveal a clear efficiency-robustness trade-off. TPC achieves the strongest performance on generative tasks, including ASR and summarization. ADS improves paralinguistic tasks but reduces generative stability when used alone. The two-stage TPC->ADS strategy provides the best overall balance, achieving the strongest DID and SER performance while outperforming large proprietary models such as Gemini-2.5-Pro on discriminative tasks. We will publicly release AraMega-SSum together with all experimental resources to support future research in Arabic speech understanding.
Jan 7, 2026cs.CL

Gavel: Agent Meets Checklist for Evaluating LLMs on Long-Context Legal Summarization

Large language models (LLMs) now support contexts of up to 1M tokens, but their strengths and weaknesses on complex long-context tasks remain unclear. To study this, we focus on multi-document legal case summarization, where a single case often spans many documents exceeding 100K tokens. We systematically evaluate 12 frontier LLMs with Gavel, which consists of Gavel-Ref, a reference-based evaluation framework with checklist, residual-fact, and writing-style evaluations, and Gavel-Agent, a reference-free agent for evaluating factual coverage directly from source documents. Our results show that current models are more prone to omitting key information than hallucinating. They all perform well on simple checklist items, such as filing date, but struggle with rare and complex items, such as settlements. Performance also declines as case length increases. To meta-evaluate Gavel, we collect 160 hours of human annotations. Gavel-Agent reduces token usage by at least 36% compared to end-to-end and chunk-by-chunk methods while achieving competitive performance. Gavel-Agent also generalizes to the medical domain, performing the best with at least 77% fewer tokens.
Jan 6, 2026cs.CL

MAPLE: Medical Aspect-Based Summarization with Phrase-Level Evidence

Trustworthy clinical summarization requires every claim to be traceable to its evidence, yet existing attribution often resolves only to the sentence or document, leaving clinicians to scan surrounding text for the few words that matter. We argue that the unit of attribution should match the unit of verification: the precise phrase the reader's eye must land on. We present MAPLE (Medical Aspect-Based Summarization with Phrase-Level Evidence), a human-annotated benchmark that grounds each summarized claim in both cited sentences and contributory phrases within them. Spanning 152 randomized controlled trial (RCT) abstracts and 16 clinically motivated aspects, MAPLE comprises 1,799 aspect-based summaries with two-level evidence. We further introduce a decoupled evaluation framework that separately scores content, traceability, and locatability, together with a proxy for the amount of source text a clinician must inspect to verify a claim. Benchmarking eleven LLMs shows that sentence-level citation is consistently strong (C-F1 up to 90.9%), while phrase-level grounding remains less stable and the most discriminative axis across models (P-F1 66.1-84.5%). These results suggest that the key challenge is not only producing accurate summaries, but localizing their supporting evidence precisely enough for efficient clinical verification. Data and code are available at https://github.com/chubohao/maple.
Dec 19, 2025cs.CL

Narrative Consolidation: Formulating a New Task for Unifying Multi-Perspective Accounts

Processing overlapping narrative documents, such as legal testimonies or historical accounts, often aims not for compression but for a unified, coherent, and chronologically sound text. Standard Multi-Document Summarization (MDS), with its focus on conciseness, fails to preserve narrative flow. This paper formally defines this challenge as a new NLP task, Narrative Consolidation, focusing on chronological integrity, completeness, and the fusion of complementary details. We establish the resources needed to study it: a formal task definition, an evaluation paradigm, the Gospel Consolidation Language Resource -- a benchmark built from the four Biblical Gospels with 169 canonical events, cross-document alignments, and a reference consolidation -- and a suite of reference systems, ranging from timeline-agnostic heuristics to the Temporal Alignment Event Graph (TAEG). Benchmarking yields three findings. First, the explicit temporal backbone is the dominant factor: every system granted the canonical timeline raises ROUGE-L F1 from 0.206 to at least 0.81. Second, on a fusion-style reference, a simple length heuristic is remarkably strong (0.947 ROUGE-L F1), outperforming graph-based selection (0.846). Third, ablations show the discriminative signal resides in temporal edges, while intra-cluster lexical similarity is uninformative. These results establish Narrative Consolidation as a distinct task and pose an open challenge: to surpass that heuristic with a principled selection mechanism and true fusion of complementary details.
Oct 8, 2025cs.CL

CARPAS: Towards Content-Aware Refinement of Provided Aspects for Summarization in Large Language Models

Aspect-based summarization has attracted significant attention for its ability to generate more fine-grained and user-aligned summaries. While most existing approaches assume a set of predefined aspects as input, real-world scenarios often present challenges where these given aspects may be incomplete, irrelevant, or entirely missing from the document. Users frequently expect systems to adaptively refine or filter the provided aspects based on the actual content. In this paper, we initiate this novel task setting, termed Content-Aware Refinement of Provided Aspects for Summarization (CARPAS), with the aim of dynamically adjusting the provided aspects based on the document context before summarizing. We construct three new datasets and conduct pilot experiments with four representative prompting strategies. Our analysis reveals that LLMs often over-generate aspects, resulting in excessively long and misaligned summaries. Building on this observation, we introduce a two-stage framework that derives lightweight scope guidance before aspect refinement and summarization, improving focus and reducing over-generation. Our extensive experiments show that the proposed approach significantly improves performance across all datasets. Moreover, our deeper analyses uncover LLMs' compliance when the requested number of aspects differs from their own estimations, establishing a crucial insight for the deployment of LLMs in similar real-world applications.
Mar 21, 2025cs.CL

Evaluating the Evaluator: Summarization Metrics and LLM-Judges beyond English

Automatic text summarization relies on automatic evaluation to quickly determine the quality of summarization models via automatic metrics and LLM-as-a-Judge models. However, these techniques require meta-evaluation to ensure that they capture human judgments correctly. In this paper, we explore this meta-evaluation beyond English by generating a new multilingual summary meta-evaluation dataset (BASSE), which comprises human judgments on 2,040 abstractive summaries, generated either manually or by five Large Language Models (LLMs) with four different prompts. For each summary, annotators evaluate five criteria on a 5-point Likert scale: coherence, consistency, fluency, relevance, and 5W1H. We then benchmark automatic summarization metrics and LLM-as-a-Judge models. Our results show that currently proprietary judge LLMs have the highest correlation with human judgments, followed by criteria-specific automatic metrics, while open-sourced judge LLMs perform poorly.
Date pendingcs.CL

Attributable by Construction: Claim-Anchored Provenance for Multi-Document Summarization

Large language models produce fluent multi-document summaries, but their attributions are typically coarse---whole documents or passages---and generated post hoc, leaving each statement hard to verify. We argue that attribution should be a structural property of generation rather than a downstream prediction. We present CAMS, a Claim-Anchored Multi-document Summarization framework that decomposes every source document into atomic claims whose provenance is resolved deterministically from verbatim quotes to token spans, clusters equivalent claims across documents while flagging inter-source conflicts, selects a support-aware and salient subset, and rewrites it so that every summary sentence terminates in claim identifiers resolving back to source spans. This yields a separation we make explicit: provenance is an invariant holding for every emitted sentence independently of model accuracy, whereas faithfulness is an objective that selection, constrained rewriting, and verification only encourage---a distinction end-to-end and post-hoc systems conflate. We evaluate on MultiNews, DiverseSumm, and zero-shot on WCEP under a two-regime protocol separating reference-free citation quality from gold-aligned localization, audited by a support model never used for selection or verification. CAMSmatches strong end-to-end and span-attribution baselines on summary quality while improving faithfulness and citation precision, raising multi-source attribution accuracy from 38% to 64% without inflating the number of cited sources, and cutting human verification time per claim by 3.4×3.4\times. We release code and ∼320{\sim}320K claim--quote--span annotations over MultiNews as a reusable fine-grained attribution resource.