There are more than 2,000 listed companies on the UK's London Stock Exchange, divided into 11 sectors who are required to communicate their financial results at least twice in a single financial year. UK annual reports are very lengthy documents with around 80 pages on average. In this study, we aim to benchmark a variety of summarisation methods on a set of different pre-trained transformers with different extraction techniques. In addition, we considered multiple evaluation metrics in order to investigate their differing behaviour and applicability on a dataset from the Financial Narrative Summarisation (FNS 2020) shared task, which is composed of annual reports published by firms listed on the London Stock Exchange and their corresponding summaries. We hypothesise that some evaluation metrics do not reflect true summarisation ability and propose a novel BRUGEscore metric, as the harmonic mean of ROUGE-2 and BERTscore. Finally, we perform a statistical significance test on our results to verify whether they are statistically robust, alongside an adversarial analysis task with three different corruption methods.
Figures & tables
Metric
Embeddings
Language Model
n-gram
ROUGE
No
N.A
n-gram
BERTScore
Yes
Roberta Large
1-gram
BARTScore
Yes
Bart Large
1-gram
METEOR
No
N.A
1-gram
Bleurt
Yes
BERT-lg
Sequence
BRUGEscore
Yes
N.A
2-gram
Table 1: Summary of the features of the evaluation metrics used in this study
Data Type
Train
Validate
Test
Report full text
3,000
363
500
Gold summaries
9,873
1,250
1,673
Table 2: FNS 2021 Shared Task Dataset
Transformer
model_name
max_input
max_target
batch_size
train_epochs
T5
t5-small
4096
512
4
5
LED base
allenai/led-base-16384
8000
1000
4
5
LED large
allenai/led-large-16384
4096
512
4
5
Pegasus
google/pegasus-large
1024
256
4
5
BART
facebook/bart-base
1024
128
4
5
Table 3: description of hyperparameters during training on the FNS dataset
Figure 1: Correlation Matrix of Scores Produced using T5
System / Metric
R-1/F
R-2/F
R-L/F
R-SU4/ F
BE/F
BA/F
bleurt
meteor
BR
T5-Small-96
0.496
0.374
0.487
0.417
0.910
0.830
-0.836972
0.184
0.530
LED-base-128
0.492
0.370
0.484
0.413
0.899
0.816
-0.849750
0.182
0.524
Pegasus
0.476
0.350
0.467
0.394
0.847
0.759
-0.925372
0.174
0.495
BART
0.453
0.317
0.440
0.365
0.852
0.774
-0.928474
0.176
0.462
Lead-1000
0.443
0.307
0.431
0.356
0.774
0.694
-1.039358
0.162
0.440
RNN-LSTM-RL
0.459
0.270
0.431
0.268
0.761
0.647
-1.027724
0.175
0.399
Table 4: F-measure scores for Rouge-1, Rouge-2, Rouge-L, SU4, BERTScore, BARTScore, Bleurt, and Meteor, ranked based on Rouge-2 F1 measure. The abbreviations used are BE for BERT score (roberta-large-mnli), BA for BART score (bart-large-mnli), and BR for BRUGEscore.
Metric
Word dropping_10 (%)
Word Permutation_10 (%)
Bert Mask filling_10(%)
ROUGE-1
0.826
0.000
0.982
ROUGE-2
0.958
1
0.99
ROUGE-3
0.968
0.998
0.992
ROUGE-S1
0.958
1
0.99
ROUGE-S2
0.946
0.996
0.992
ROUGE-L
0.922
0.978
0.99
Table 5: Mean accuracy by metric on the three corruption tasks. We apply three types of corruptions on the system generated summaries. We create a corruption every 10 chunks. Each metric is used to score the original and the corrupted versions of these summaries. This task should give the uncorrupted version a higher score to make sure that the metric is sensitive to corrupted summaries. The results reported shows the accuracy by metric on this task. All standard deviations were small (less than 0.2%). The experiments were performed on the FNS dataset using the best performing system which is the small version of T5 transformer
T5-Small-96
LED-BASE-128
0.0448
LED-BASE-256
0.0161
0.3352
LED-BASE-1000
0.0042
0.1930
0.3210
BART
0.0000
0.0000
0.0000
0.0000
mBART
0.0000
0.0000
0.0000
0.0000
0.4943
PEGASUS
0.0000
0.0000
0.0000
0.0000
0.2440
0.2429
Table 6: The p-values of the BERT score results using the Bootstrap test are presented in each column, where column i includes the p-values of system i and the p-values of the remaining n-i systems.
T5-Small-96
LED-BASE-128
0.2622
LED-BASE-256
0.1766
0.3579
LED-BASE-1000
0.0774
0.2204
0.3232
PEGASUS
0.0000
0.0001
0.0004
0.0015
BART
0.0000
0.0000
0.0003
0.0011
0.4284
mBART
0.0000
0.0001
0.0001
0.0009
0.4227
0.4961
Table 7: The p-values of the Bleurt score results using the Bootstrap test are presented in each column, where column i includes the p-values of system i and the p-values of the remaining n-i systems.
T5-Small-96
LED-BASE-128
0.1243
LED-BASE-256
0.0777
0.1649
LED-BASE-1000
0.0453
0.0874
0.2068
PEGASUS
0.0000
0.0000
0.0000
0.0000
T5-MULTI-REFERENCES
0.0000
0.0000
0.0000
0.0000
0.0003
T5-Small-256
0.0000
0.0000
0.0000
0.0000
0.0001
0.2906
Table 8: The p-values of the Rouge-2 score results using the Bootstrap test are presented in each column, where column i includes the p-values of system i and the p-values of the remaining n-i systems.
T5-Small-96
LED-BASE-256
0.1075
LED-BASE-128
0.1330
0.7523
LED-BASE-1000
0.0520
0.1245
0.0740
PEGASUS
0.0000
0.0000
0.0000
0.0000
T5-MULTI-REFERENCES
0.0000
0.0000
0.0000
0.0000
0.0000
PEGASUS-MULTI-REFERENCES
0.0000
0.0000
0.0000
0.0000
0.0000
0.2039
Table 9: The p-values of the Rouge-L score results using the Bootstrap test are presented in each column, where column i includes the p-values of system i and the p-values of the remaining n-i systems.
Summarization ships in countless production systems, making model selection a routine decision that depends on measuring summary quality. Existing metrics struggle to support this: ROUGE captures only surface overlap, while LLM-as-judge scores saturate to near-identical values that fail to rank models effectively. We observe this saturation across three public datasets, two proprietary datasets, and multilingual settings. Motivated by this, we introduce Semantic Scaffold, an evaluation framework that extracts a hierarchical representation of facts, questions, and entity attributes from a source text, labeling each as a main point or supporting detail, and reusing this structure as a fixed reference for scoring summaries. From this representation, we derive three diagnostic metrics: Fact Preservation Score (FPS), Question Preservation Score (QPS), and Entity Preservation Score (EPS), designed to reward the preservation of essential information while penalizing detail overload, and position them as interpretable diagnostics that remain informative where holistic axes collapse. Finally, we analyze four recurring failure modes of ROUGE and LLM-as-judge scores, demonstrating that scaffold-based evaluation remains informative where conventional metrics collapse.
Finance reporting is a natural proving ground for large language models, and the very-long-context capabilities of recent models across all sizes make rigorous evaluation in this domain an increasingly pressing need. Yet most public financial resources reduce the task to plain-text SEC 10-K filings paired with a handful of question-answer items. We release LEDGER (Long-context Evaluation of Documents for Grounded Extraction and Retrieval), a corpus of 4,999 digitized corporate annual reports - full documents with figures, tables, and narrative, not just regulatory filings. Each report is labeled with 31 consolidated financial KPIs to be extracted and linked to the market's reaction at the earnings date. From this data we derive three evaluation benchmarks spanning the difficulty spectrum: a pure page-level KPI retrieval task with TREC-style relevance judgments over 118,048 questions in natural language, a conversational "needle-in-a-haystack" single-value lookup, and a full KPI extraction task, both from long, numerically dense reports. We additionally provide human OCR-quality annotations with inter-annotator agreement and the complete extraction, validation, and scoring toolchain. We further demonstrate the dataset's research utility with a case study linking CEO-letter rhetoric to post-publication market impact.
Charles Moslonka, Amaury de Vitry, Arthur Garnier +2
Artefact Research Center Paris, France · MICS, CentraleSupélec, Université Paris-Saclay Gif-sur-Yvette, France · Ardian Paris, France
Long documents often distribute important information across extensive narrative passages and multiple tables, making faithful summarization particularly challenging. Existing methods may generate individually supported quantitative facts and analytical statements yet associate them incorrectly, producing quantitatively plausible yet analytically unfaithful summaries. In this work, we propose LOOMSUM, a training-free framework that extracts source-grounded atomic evidence, explicitly links table-derived facts with supporting narrative analyses, and plans the discourse structure before generation. We also introduce Table-Grounded Faithfulness (TGF), a claim-level metric that separately evaluates Numeric Grounding, Analysis Support, and Relation Consistency. Experiments on the text--table summarization benchmarks FINDSum and USTT show that LOOMSUM improves analytical faithfulness while maintaining strong summarization quality. Human evaluation finds positive component-level associations with the corresponding human judgments. Our Relation Consistency metric further shows stronger agreement with human relation judgments than generic factuality metrics, indicating that explicit cross-modal linking helps reduce errors in which supported quantities are paired with incorrect narrative interpretations. Together, these findings show that faithful long text--table summarization requires not only grounding individual facts, but also preserving the relations between them.
Meng Zhou, Wenhao You, Wei Yuan
University of Toronto · University of Waterloo · Independent Researcher