cs.AISep 26, 2026

READ-Bench: Benchmarking Historical Instance Retrieval for Time-Series Diagnosis

Authors: Gerardo Pastrana, Haojun Li, Dhruv Mehta, Anoushka Vyas, Sina Khoshfetrat Pakazad, Henrik Ohlsson, John Paparrizos

Organizations: C3 AI · The Ohio State University

Abstract

Time-series diagnostic systems rarely rely on retrieving relevant historical cases, and when they do, retrieval is evaluated only indirectly through downstream prediction. We introduce READ-Bench, a benchmark for historical-case retrieval across 12 diagnostic datasets, centered on multivariate time series, that defines relevance by shared fault or event type rather than signal shape, so visually different traces of the same fault count as relevant while similar-looking traces of different faults do not. Treating retrieval as a base retriever followed by a reranker, we evaluate classical distances, symbolic retrievers, self-supervised and foundation-model embedders, and their fusion, plus label-aware and language-model rerankers, under one protocol that varies supervision, pollution, and corpus scale with significance testing. Under a common channel-independent interface, pretrained representations offer no statistically detectable advantage over strong classical and symbolic baselines for search alone. The decisive factor is a small amount of resolved-case supervision at reranking, namely a Gaussian-process reranker that propagates a few neighbor labels in embedding space, which helps far more than more sophisticated representations or language-model reasoning and holds under pollution and at full corpus scale. Guided by these findings, we fuse a normal-residual-scored embedder with a dynamic time warping leg via reciprocal-rank fusion, then rerank with the Gaussian-process reranker, improving NDCG@10 over its own search stage on all 12 datasets, by +0.11 from reranking and +0.16 over the strongest single base retriever.

Figures & tables

Appendix figures & tables74 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Which Histories Matter for Time Series Forecasting? Learning Predictive Relevance with Future Supervision

    Aug 24, 2026Yong-Hoon Choi, Kwang-Hyun Park, Youngjin ChoLearning to RankTime Series Forecasting

  2. Revitalizing Medical Time Series with Vision-Informed Retrieval: A Vision-Language Perspective

    Sep 28, 2026Guoqi Yu, Juncheng Wang, Shujun WangMultivariate Time Series ClassificationTime Series Classification

  3. TS-Haystack: A Multi-Task Retrieval Benchmark for Long-Context Time-Series Reasoning

    Feb 15, 2026Nicolas Zumarraga, Thomas Kaar, Ning Wang +12Time Series ReasoningTime Series QA