Can Classical Semantic-Extractive Summarization Be Evaluated in Hindi? A Replication Study
Authors: Showket Ahmad Khan, Mudasir Mohd, Nasrullah Sheikh, Mohsin Altaf Wani, Abid Hussain Wani, Hilal Ahmad Khanday, Niyaz Ahmad Wani
Organizations: Department of Computer Science, South Campus, University of Kashmir, Anantnag, India · IBM Research, San Jose, CA, USA · Manipal University Jaipur, Dehmi Kalan, Jaipur 303007, Rajasthan, India
We replicate the distributional-semantics extractive summarisation method of Mohd, Jan and Shah (2020) and adapt it to Hindi, substituting a Devanagari-appropriate component at every language-specific step. The system is evaluated on two independent corpora --- the Hindi portion of XL-Sum and FIRE ILSUM 2.0 Hindi --- under a Devanagari-aware ROUGE implementation validated against the XL-Sum authors' own multilingual scorer, with all comparisons drawn as 1000-resample paired bootstraps. In its published equal-weight configuration the replicated system is significantly worse than a three-sentence lead baseline on both corpora, trailing Lead-3 by 0.042 ROUGE-1 Fon XL-Sum and by 0.265 on ILSUM. A feature ablation shows that sentenceposition is the only feature that contributes: position alone reproduces the lead baseline exactly, removing position gives the weakest configuration,and a validation-tuned weighting can at best equal Lead-3 and never exceed it. TextRank fails identically, making this a class-level rather than an implementation-level result. A selection analysis shows the remaining features steer extraction towards long, entity-dense body sentences while the references reuse the article lead.Current Hindi benchmarks therefore cannot reward non-lead content selection, motivating purpose-built evaluation resources.
Figures & tables
System
R1-F [95% CI]
R2-F
RL-F
SU4-F
Ours-full (7 feat, cluster-rr)
0.186 [0.176, 0.196]
0.048
0.126
0.062
Ours-global-topn (7 feat)
0.185 [0.175, 0.195]
0.048
0.124
0.062
Ours-position-only
0.228 [0.218, 0.239]
0.058
0.160
0.076
Ours-minus-position
0.182 [0.172, 0.191]
0.047
0.121
0.061
Lead-3
0.228 [0.218, 0.239]
0.058
0.160
0.076
TextRank-3
0.204 [0.193, 0.214]
0.053
0.140
0.068
Table 1: XL-Sum Hindi, 200 test documents, three-sentence extracts. ROUGE F-measure; ROUGE-1 with 95% bootstrap CI.
System
Δ R1-F vs Lead-3
95% CI
Real gap?
Ours-full
− 0.042
[ − 0.051, − 0.033]
yes
Ours-global-topn
− 0.043
[ − 0.052, − 0.034]
yes
Ours-position-only
+ 0.000
[ + 0.000, + 0.000]
no
Ours-minus-position
− 0.046
[ − 0.056, − 0.037]
yes
TextRank-3
− 0.025
[ − 0.034, − 0.014]
yes
Random-3
− 0.029
[ − 0.038, − 0.019]
yes
Table 2: XL-Sum Hindi, paired ROUGE-1 F gap against Lead-3 (1000-resample paired bootstrap). A gap is “real” when its 95% interval excludes zero.
System
R1-F [95% CI]
R2-F
RL-F
SU4-F
all7 (equal weight)
0.256 [0.237, 0.277]
0.148
0.201
0.153
Lead-1
0.441 [0.407, 0.476]
0.380
0.428
0.367
Lead-3
0.522 [0.488, 0.557]
0.451
0.499
0.449
TextRank-3
0.268 [0.245, 0.291]
0.146
0.214
0.152
Random-3
0.233 [0.212, 0.254]
0.102
0.175
0.115
Table 3: ILSUM 2.0 Hindi, 200 held-out test documents, three-sentence extracts. ROUGE F-measure; ROUGE-1 with 95% bootstrap CI.
System
Δ R1-F vs Lead-3
95% CI
Real gap?
all7 (equal weight)
− 0.265
[ − 0.296, − 0.235]
yes
Lead-1
− 0.080
[ − 0.113, − 0.049]
yes
TextRank-3
− 0.254
[ − 0.289, − 0.224]
yes
Random-3
− 0.289
[ − 0.323, − 0.259]
yes
Table 4: ILSUM 2.0 Hindi, paired ROUGE-1 F gap against Lead-3 (1000-resample paired bootstrap).
Position weight
XL-Sum, all7
XL-Sum, pos+tfidf
ILSUM, pos+tfidf
1
0.191
0.211
0.347
2
0.198
0.223
0.426
4
0.209
0.234
0.490
8
0.227
0.238
0.543
16
0.235
0.239
0.560
Table 5: Validation-set ROUGE-1 F under a position-weight sweep (three-sentence budget). XL-Sum figures are the seven-feature and position+TF–IDF families; ILSUM figures are the position+TF–IDF family.
Corpus (test slice)
Tuned R1-F [95% CI]
Lead-3 R1-F [95% CI]
Δ R1-F [95% CI]
Verdict
XL-Sum (docs 201–400)
0.231 [0.221, 0.241]
0.231 [0.221, 0.242]
− 0.000 [ − 0.002, + 0.001]
statistical tie
ILSUM (test split)
0.516 [0.483, 0.552]
0.522 [0.488, 0.557]
− 0.005 [ − 0.011, − 0.000]
detectable but negligible; never above
Table 6: Tuned configuration (position+TF–IDF, position weight sixteen) against Lead-3 on each held-out test slice; ROUGE-1 F with 95% bootstrap CIs and the paired gap. Statistically detectable but negligible ( ≤ 0.011 R1-F) on ILSUM and a tie on XL-Sum: at best equal, never above.
System
D1
D2
D3
D4
D5
D6
D7
D8
D9
D10
mean
Ours-full
0.20
0.14
0.11
0.10
0.09
0.07
0.08
0.04
0.07
0.08
0.391
Ours-global-topn
0.20
0.12
0.12
0.11
0.09
0.08
0.07
0.05
0.07
0.08
0.394
Ours-position-only
0.71
0.20
0.08
0.01
0.01
0.00
0.00
0.00
0.00
0.00
0.070
Ours-minus-position
0.12
0.09
0.10
0.10
0.09
0.09
0.09
0.06
0.12
0.14
0.501
Lead-3
0.71
0.20
0.08
0.01
0.01
0.00
0.00
0.00
0.00
0.00
0.070
TextRank-3
0.15
0.10
0.11
0.10
0.09
0.12
0.10
0.08
0.09
0.06
0.441
Table 7: Fraction of selected sentences by source-position decile on the XL-Sum ablation set (D1 = first tenth of the article, D10 = last tenth), with mean normalised position.
Faithfulness is a central concern in legal text summarization, which motivates extractive approaches that select verbatim content traceable to its source. Such methods typically rank paragraphs or other structural units in isolation, yet give little attention to consolidating evidence that is distributed across, and shares salience between, distant parts of a document. We introduce LexLattice, an extractive summarizer that reifies a legal act's hierarchy as a two-dimensional semantic lattice and consolidates over it with a masked 2D neural cellular automata before selection. LexLattice attains state-of-the-art ROUGE across all 24 languages of EUR-Lex-Sum in both multilingual and cross-lingual settings, surpassing instruction-tuned baselines with billions of parameters, despite concentrating all trainable capacity in a 1.8M parameter consolidator over a frozen multilingual encoder. A consolidator trained only on high-resource languages further transfers to unseen languages with near-lossless retention (0.99), indicating that the model operates on language-agnostic semantic geometry rather than surface form. Our results position explicit consolidation over document structure as a compact and traceable alternative to scale for multilingual legal summarization.
Sujay Uday Rittikar, Sheela Ramanna
Applied Computer Science The University of Winnipeg Winnipeg, MB
Summarization ships in countless production systems, making model selection a routine decision that depends on measuring summary quality. Existing metrics struggle to support this: ROUGE captures only surface overlap, while LLM-as-judge scores saturate to near-identical values that fail to rank models effectively. We observe this saturation across three public datasets, two proprietary datasets, and multilingual settings. Motivated by this, we introduce Semantic Scaffold, an evaluation framework that extracts a hierarchical representation of facts, questions, and entity attributes from a source text, labeling each as a main point or supporting detail, and reusing this structure as a fixed reference for scoring summaries. From this representation, we derive three diagnostic metrics: Fact Preservation Score (FPS), Question Preservation Score (QPS), and Entity Preservation Score (EPS), designed to reward the preservation of essential information while penalizing detail overload, and position them as interpretable diagnostics that remain informative where holistic axes collapse. Finally, we analyze four recurring failure modes of ROUGE and LLM-as-judge scores, demonstrating that scaffold-based evaluation remains informative where conventional metrics collapse.
Quantifying abstractiveness in generated summaries is essential for evaluating summarization models beyond surface-level metrics like ROUGE. We introduce Reference Abstraction (RA), Summary Abstraction (SA), and Abstraction Ratio (AR) -- a set of principled heuristic metrics that measure how much a summary diverges from extractive copying of the source text. The formulation uses the harmonic mean of document lengths modulated by a cubic non-overlap factor, yielding dimensionally consistent, bounded output with non-linear sensitivity to the extractive-abstractive boundary. Evaluation on 100 XSUM documents across four summarization models (BART-large-cnn, Pegasus-xsum, DistilBart, MT5-small) demonstrates that the metrics successfully discriminate between extractive models (SA ~ 0.12-0.26) and abstractive models (SA ~ 0.96-1.77), and that the Abstraction Ratio identifies summaries requiring manual evaluation for potential hallucination. Code and results are available at https://github.com/katweNLP/AbstractionStudy.
Praveenkumar Katwe, Rakesh Chandra Balabantaray, Kali Prasad Vittala
Department of Computer Science and Engineering, International Institute of Information Technology, Bhubaneswar, India · Salesforce India Pvt Ltd, Bengaluru, India