cs.DLDec 14, 2025

MTRACE: Multilingual Retrieval-Augmented Generation for Temporally Diverse Text Corpora

Authors: Souhail Bakkali, Anthony Mudet

Organizations: Univ Rennes, CNRS, IRISA - UMR 6074, Rennes, France · L3i-lab, La Rochelle Universit´e, France

Abstract

Large multilingual knowledge bases expose temporally diverse information, yet retrieval quality remains sensitive to lexical variation and cross-lingual terminology shifts. We develop and evaluate MTRACE (Multilingual Temporal Retrieval-Augmented Generation with evidence grounding), a pipeline designed to test whether query expansion and multi-query fusion mitigate vocabulary mismatch in temporally layered text corpora, on the French and English subsets of MIRACL. Our approach integrates: (i) semantic query expansion (SQE) and multi-query fusion via Reciprocal Rank Fusion (RRF), targeting retrieval stability under query variation; (ii) a generation prompt enforcing strict grounding in retrieved evidence and explicit abstention when evidence is insufficient; and (iii) a modular architecture enabling systematic component evaluation. Ablation studies on Named Entity Recognition (NER) and embedding model selection demonstrate the importance of syntactic coherence in entity extraction and of self-retrieval and efficiency measurements for retriever selection. Our end-to-end evaluation over 50 constructed queries shows faithful answers for well-supported queries, correct abstention on unanswerable questions, and no re-scored similarity gains from multi-query fusion over single-query dense retrieval. By scoping our claims to a clean, text-only baseline, we separate these effects from OCR-noise confounds; direct measurement of diachronic lexical drift is left to future work. Code and configurations are available at \url{https://anonymous.4open.science/r/MIRAGE-8EAA/

Figures & tables

Appendix figures & tables7 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

Apr 28, 2026cs.CL

CORAL: Adaptive Retrieval Loop for Culturally-Aligned Multilingual RAG

Multilingual retrieval-augmented generation (mRAG) is often implemented within a fixed retrieval space, typically via query or document translation or multilingual embedding vector representations. However, this approach may be inadequate for culturally grounded queries, in which retrieval-condition misalignment may occur. Even strong retrievers and generators may struggle to produce culturally relevant answers when sourcing evidence from inappropriate linguistic or regional contexts. To this end, we introduce CORAL (COntext-aware Retrieval with Agentic Loop, an adaptive retrieval methodology for mRAG that enables iterative refinement of both the retrieval space (corpora) and the retrieval probe (query) based on the quality of the evidence. The overall process includes: (1) selecting corpora, (2) retrieving documents, (3) critiquing evidence for relevance and cultural alignment, and (4) checking sufficiency. If the retrieved documents are insufficient to answer the query correctly, the system (5) reselects corpora and rewrites the query. Across two cultural QA benchmarks, CORAL achieves up to a 3.58%p accuracy improvement on low-resource languages relative to the strongest baselines.
Sep 23, 2026cs.CL

TEMPS: Temporal Sentence Embeddings for Temporal Information Retrieval

Modern information retrieval (IR) systems rarely represent time, yet many information needs depend on it: in clinical, journalistic, and legal search, when an event occurred can decide whether a document is relevant. Dense retrievers and Retrieval-Augmented Generation (RAG) pipelines match queries to documents well on topic but poorly on time, so they surface content that is on-topic yet temporally wrong. We introduce Temporal Textual Similarity (TTS), a task that measures how well two anchored texts align in time, independent of their topical similarity. We then present TEMPS (Temporal Embedding Model for Precise Search), a modular temporal branch that attaches to a frozen semantic retriever and trains on that signal. It resolves anchored temporal expressions to intervals and moment-matches each one to a Gaussian; the resulting ordering supervises an anchor-date-conditioned encoder, whose score we fuse with the semantic score at inference. Grounding supplies the supervision, so training uses no hand-labeled temporal data. The temporal score itself is the Gaussian-KL inclusion measure from distributional embeddings; what TEMPS adds is the grounding and the moment-matched supervision. On three temporal benchmarks, TEMPS improves MRR for every semantic backbone tested and, on TS- Retriever, lifts R@1 from 19.92 to 25.39 over the prior temporal state of the art.
Apr 22, 2026cs.CL

All Languages Matter: Understanding and Mitigating Language Bias in Multilingual RAG

Multilingual Retrieval-Augmented Generation (mRAG) leverages cross-lingual evidence to ground Large Language Models (LLMs) in global knowledge. However, we show that current mRAG systems suffer from a language bias during reranking, systematically favoring English and the query's native language. By introducing an estimated oracle evidence analysis, we quantify a substantial performance gap between existing rerankers and the achievable upper bound. Further analysis reveals a critical distributional mismatch: while optimal predictions require evidence scattered across multiple languages, current systems systematically suppress such ``answer-critical'' documents, thereby limiting downstream generation performance. To bridge this gap, we propose \textit{\textbf{L}anguage-\textbf{A}gnostic \textbf{U}tility-driven \textbf{R}eranker \textbf{A}lignment (LAURA)}, which aligns multilingual evidence ranking with downstream generative utility. Experiments across diverse languages and generation models show that LAURA effectively mitigates language bias and consistently improves mRAG performance.