cs.CLOct 6, 2026

Leveraging a four-quadrant approach for evaluating Redpine Science

Authors: Filip Dorm, Leonora Vesterbacka

Organizations: Redpine, Sweden.

Abstract

Redpine Science gives models and agents a single access point to a wide range of peer-reviewed literature, queried directly through the Model Context Protocol (MCP) and an API. This report evaluates Redpine Science on two levels: the relevance of the retrieved chunks, and a model's answer when it has access to Redpine Science compared to web search. Both public and expert-validated benchmarks are used. Public benchmarks are a widely accepted way to test model development and are comparable across labs, but risk saturation and memorization. To address this, we complement them with an expert-validated question set. In total, this report presents four evaluations. On ScholarQABench SciFact, the public answer-quality benchmark reported here, an agent with Redpine Science answers 94.4% of claims correctly against 87.6% with no retrieval. On the expert-validated question set, an agent with Redpine Science states 80.1% of the required claims against 70.2% for an agent restricted to web search. On the 668 queries of a public retrieval benchmark whose gold paper Redpine holds, stripped of any model reasoning, Redpine Science places the correct source paper in its top ten results for 83.1% of queries (Recall@10), against 79.3% for the benchmark's creator. A blinded expert relevance panel places Redpine Science's Precision@5 at 75.2% against 39.8% for the PubMed search tool. We release the expert-validated question set and instructions to reproduce every headline result above, at https://github.com/redpine-ai/benchmarks.

Figures & tables

Appendix figures & tables2 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

Jul 13, 2026cs.CL

ResearchQA: Benchmarking Citation-Grounded Question-Answering on Scientific Papers

Large language models are increasingly used to assist scientific reading, but existing evaluation methods often fail to detect whether answers are supported by verifiable citations. We introduce ResearchQA, a benchmark of 6,211 single-paper question-answer pairs from 494 open-access papers spanning eight domains and four question types: lookup, comprehension, multi-hop, and adversarial. ResearchQA is designed for citation-grounded evaluation: it permits multiple valid supporting passages for a claim and rewards grounded refusal when the source paper does not support an answer. We evaluate eight leading closed- and open-weight models in a citation-grounded chat-with-paper setting using a deterministic citation matcher and an LLM-based rubric evaluator. Citation-based metrics separate systems more clearly than LLM-evaluator scores: section coverage and citation accuracy vary substantially across models, while evaluator scores remain tightly compressed. We further find that open-weight models approach the best closed-model citation accuracy while achieving 3 to 6 times lower per-example latency. We release the benchmark, evaluation harness, and evaluator prompt.
Sep 23, 2026cs.AI

Large Knowledge Model: A Knowledge Foundation for Agentic Science at Scale

Agentic science envisions many autonomous agents investigating concurrently while building on a shared, evolving body of scientific knowledge. This requires a knowledge foundation that supports high-concurrency access, preserves traceable and reusable reasoning, and grows incrementally. We propose the Large Knowledge Model (LKM), a growing, agent-native knowledge foundation that provides a general representation of scientific knowledge across disciplines. LKM organizes the scientific literature into reasoning graphs, with claims as the core nodes and associated reasoning chains that make explicit how premises and evidence support conclusions. These source-grounded objects are persistent and addressable; cross-paper links organize them into aligned question, workflow, and evidence views. Newly extracted papers extend the foundation incrementally while preserving existing object identities. Building on this foundation, we develop an agent-native, reasoning-aware scientific retrieval system that retrieves claims together with their reasoning chains and sources, enabling agents to inspect and reuse the evidence underlying scientific conclusions. Across benchmarks, agents using LKM retrieve more evidence, cite more faithfully, and answer scientific questions more accurately: LKM nearly doubles the known supporting and contradicting evidence retrieved on SciFact-Open (818 versus 443 claim-paper pairs), reasoning graphs raise citation F1 on ScholarQABench by more than 5 points over the same retrieved papers, and LKM retrieval improves a fixed answering model by 9.3, 4.2, and 14.7 points over no retrieval on ChemBench, PubMedQA, and SciBench. LKM lays the foundation for a scientific ecosystem in which AI scientists not only recall accumulated knowledge but also extend it, returning new questions, workflows, and evidence to a memory that every subsequent investigation can build on.
Aug 7, 2026cs.CL

LitTraceQA: A Benchmark for Multi-Stage Grounding and Verification in Scientific Question Answering

Scientific literature is increasingly used as a knowledge source for language models, retrieval-augmented generation systems, and research assistants, but answering research questions from papers requires more than fluent generation. A reliable system must identify the relevant papers, locate the concrete evidence that supports the answer, and produce a response that is faithful to that evidence. We present LitTraceQA, a benchmark for literature-grounded question answering over scientific papers. Given a research question and a metadata pool of papers, a system must return three connected outputs: canonical paper identifiers, supporting evidence locations, and answers in one or more requested formats, including free-form text, multiple-choice answers, and structured tables. LitTraceQA targets evidence types common in scientific reading: tables, figures, text spans, equations or algorithms, and citation contexts. The public development split contains 55 examples, including 26 hidden-source single-paper questions and 29 multi-paper questions, and provides gold papers, evidence annotations, and answers for local validation. We also analyze a larger final annotation collection with 4,978 unique-question records over 4,859 unique gold papers. By evaluating paper retrieval, evidence grounding, and answer accuracy separately, LitTraceQA provides a testbed for scientific QA systems that produce verifiable answers rather than unsupported summaries.