Can Classical Semantic-Extractive Summarization Be Evaluated in Hindi? A Replication Study
Authors: Showket Ahmad Khan, Mudasir Mohd, Nasrullah Sheikh, Mohsin Altaf Wani, Abid Hussain Wani, Hilal Ahmad Khanday, Niyaz Ahmad Wani
Organizations: Department of Computer Science, South Campus, University of Kashmir, Anantnag, India · IBM Research, San Jose, CA, USA · Manipal University Jaipur, Dehmi Kalan, Jaipur 303007, Rajasthan, India
We replicate the distributional-semantics extractive summarisation method of Mohd, Jan and Shah (2020) and adapt it to Hindi, substituting a Devanagari-appropriate component at every language-specific step. The system is evaluated on two independent corpora --- the Hindi portion of XL-Sum and FIRE ILSUM 2.0 Hindi --- under a Devanagari-aware ROUGE implementation validated against the XL-Sum authors' own multilingual scorer, with all comparisons drawn as 1000-resample paired bootstraps. In its published equal-weight configuration the replicated system is significantly worse than a three-sentence lead baseline on both corpora, trailing Lead-3 by 0.042 ROUGE-1 Fon XL-Sum and by 0.265 on ILSUM. A feature ablation shows that sentenceposition is the only feature that contributes: position alone reproduces the lead baseline exactly, removing position gives the weakest configuration,and a validation-tuned weighting can at best equal Lead-3 and never exceed it. TextRank fails identically, making this a class-level rather than an implementation-level result. A selection analysis shows the remaining features steer extraction towards long, entity-dense body sentences while the references reuse the article lead.Current Hindi benchmarks therefore cannot reward non-lead content selection, motivating purpose-built evaluation resources.
Figures & tables
System
R1-F [95% CI]
R2-F
RL-F
SU4-F
Ours-full (7 feat, cluster-rr)
0.186 [0.176, 0.196]
0.048
0.126
0.062
Ours-global-topn (7 feat)
0.185 [0.175, 0.195]
0.048
0.124
0.062
Ours-position-only
0.228 [0.218, 0.239]
0.058
0.160
0.076
Ours-minus-position
0.182 [0.172, 0.191]
0.047
0.121
0.061
Lead-3
0.228 [0.218, 0.239]
0.058
0.160
0.076
TextRank-3
0.204 [0.193, 0.214]
0.053
0.140
0.068
Table 1: XL-Sum Hindi, 200 test documents, three-sentence extracts. ROUGE F-measure; ROUGE-1 with 95% bootstrap CI.
System
Δ R1-F vs Lead-3
95% CI
Real gap?
Ours-full
− 0.042
[ − 0.051, − 0.033]
yes
Ours-global-topn
− 0.043
[ − 0.052, − 0.034]
yes
Ours-position-only
+ 0.000
[ + 0.000, + 0.000]
no
Ours-minus-position
− 0.046
[ − 0.056, − 0.037]
yes
TextRank-3
− 0.025
[ − 0.034, − 0.014]
yes
Random-3
− 0.029
[ − 0.038, − 0.019]
yes
Table 2: XL-Sum Hindi, paired ROUGE-1 F gap against Lead-3 (1000-resample paired bootstrap). A gap is “real” when its 95% interval excludes zero.
System
R1-F [95% CI]
R2-F
RL-F
SU4-F
all7 (equal weight)
0.256 [0.237, 0.277]
0.148
0.201
0.153
Lead-1
0.441 [0.407, 0.476]
0.380
0.428
0.367
Lead-3
0.522 [0.488, 0.557]
0.451
0.499
0.449
TextRank-3
0.268 [0.245, 0.291]
0.146
0.214
0.152
Random-3
0.233 [0.212, 0.254]
0.102
0.175
0.115
Table 3: ILSUM 2.0 Hindi, 200 held-out test documents, three-sentence extracts. ROUGE F-measure; ROUGE-1 with 95% bootstrap CI.
System
Δ R1-F vs Lead-3
95% CI
Real gap?
all7 (equal weight)
− 0.265
[ − 0.296, − 0.235]
yes
Lead-1
− 0.080
[ − 0.113, − 0.049]
yes
TextRank-3
− 0.254
[ − 0.289, − 0.224]
yes
Random-3
− 0.289
[ − 0.323, − 0.259]
yes
Table 4: ILSUM 2.0 Hindi, paired ROUGE-1 F gap against Lead-3 (1000-resample paired bootstrap).
Position weight
XL-Sum, all7
XL-Sum, pos+tfidf
ILSUM, pos+tfidf
1
0.191
0.211
0.347
2
0.198
0.223
0.426
4
0.209
0.234
0.490
8
0.227
0.238
0.543
16
0.235
0.239
0.560
Table 5: Validation-set ROUGE-1 F under a position-weight sweep (three-sentence budget). XL-Sum figures are the seven-feature and position+TF–IDF families; ILSUM figures are the position+TF–IDF family.
Corpus (test slice)
Tuned R1-F [95% CI]
Lead-3 R1-F [95% CI]
Δ R1-F [95% CI]
Verdict
XL-Sum (docs 201–400)
0.231 [0.221, 0.241]
0.231 [0.221, 0.242]
− 0.000 [ − 0.002, + 0.001]
statistical tie
ILSUM (test split)
0.516 [0.483, 0.552]
0.522 [0.488, 0.557]
− 0.005 [ − 0.011, − 0.000]
detectable but negligible; never above
Table 6: Tuned configuration (position+TF–IDF, position weight sixteen) against Lead-3 on each held-out test slice; ROUGE-1 F with 95% bootstrap CIs and the paired gap. Statistically detectable but negligible ( ≤ 0.011 R1-F) on ILSUM and a tie on XL-Sum: at best equal, never above.
System
D1
D2
D3
D4
D5
D6
D7
D8
D9
D10
mean
Ours-full
0.20
0.14
0.11
0.10
0.09
0.07
0.08
0.04
0.07
0.08
0.391
Ours-global-topn
0.20
0.12
0.12
0.11
0.09
0.08
0.07
0.05
0.07
0.08
0.394
Ours-position-only
0.71
0.20
0.08
0.01
0.01
0.00
0.00
0.00
0.00
0.00
0.070
Ours-minus-position
0.12
0.09
0.10
0.10
0.09
0.09
0.09
0.06
0.12
0.14
0.501
Lead-3
0.71
0.20
0.08
0.01
0.01
0.00
0.00
0.00
0.00
0.00
0.070
TextRank-3
0.15
0.10
0.11
0.10
0.09
0.12
0.10
0.08
0.09
0.06
0.441
Table 7: Fraction of selected sentences by source-position decile on the XL-Sum ablation set (D1 = first tenth of the article, D10 = last tenth), with mean normalised position.
Department of Computer Science and Engineering, International Institute of Information Technology, Bhubaneswar, India · Salesforce India Pvt Ltd, Bengaluru, India