We study retrieval-augmented generation (RAG) for questions about French PDF documents when both the answer and its supporting document pages are evaluated. Five system variants add dense retrieval, rank fusion, reranking, and query decomposition to a BM25 baseline. On 595 challenge questions, the complete system scores 0.4450 MRR@10 and 0.4013 Recall@10, compared with 0.3430 and 0.2994 for BM25. Dense retrieval alone and a simple lexical--dense fusion both underperform BM25. Reranking improves the hybrid system, whereas adding query decomposition produces the largest further gain, with higher latency and more detected output artifacts. The complete system slightly exceeds the reported anonymous overall mean on two answer metrics but falls below it on most page-retrieval metrics. These results identify accurate page selection, rather than semantic retrieval in isolation, as the main opportunity for improvement in this setting.
Figures & tables
Run
Retrieval
RRF
Rerank
Decompose
A
BM25
–
–
–
B
Dense
–
–
–
C
BM25 + dense
Yes
–
–
D
BM25 + dense
Yes
Yes
–
E
BM25 + dense
Yes
Yes
Yes
Table 1: Five component variants (corresponding to runs 1–5 in the supplied results).
Run
LLMaaJ
MRR
Recall
Top1
nDCG
QBERT
QPARA
@10
@10
@10
Albert
A: BM25
.4820
.3430
.2994
.2913
.2916
.6935
.7087
B: dense
.2383
.0874
.0703
.0777
.0716
.6444
.5534
C: hybrid
.4204
.2168
.2201
.1650
.1966
.6826
.6796
D: reranked
.4262
.2573
.2431
.2039
.2265
.6708
.7282
E: decomposed
.5940
.4450
.4013
.3592
.3801
.7093
.7961
Table 2: Reported challenge scores. Bold identifies the best of our five runs in each column. The last two rows are anonymous challenge-wide summary statistics, not additional runs.
Run
Non-answers
Very short
Script intrusions
Residues
A
88
9
11
82
B
156
30
12
54
C
111
14
14
82
D
100
12
11
89
E
66
4
21
92
Table 3: Reported counts from automated output checks over 595 questions per run.
Run
Runtime (min)
Client energy (kWh)
Client emissions (g CO 2 e)
A
70.3
.01759
.879
B
65.2
.01630
.815
C
77.5
.01937
.968
D
78.3
.01957
.978
E
101.1
.02526
1.263
Table 4: Reported runtime and partial client-side estimates under the stated constant-power assumptions.