cs.CLSep 17, 2026

F2^{2}DR: A Fine-Grained Full-Pipeline Reward Framework for DeepSearch Workflows

Authors: Bojian Xiong, Wentao Ding, Yujing Lu, Shaowei Zhang, Ling Shi, Jing Liao, Yan Wang, Yueyang Zhang, +6 more

Organizations: Tianjin University · TJUNLP Lab, Tianjin University, Tianjin, China · Baidu Inc. · Baidu Inc., Beijing, China

Abstract

With the widespread industrial deployment of Large Language Models (LLMs), DeepSearch has emerged as the dominant paradigm for resolving complex user queries. It typically operates through an iterative closed-loop workflow consisting of planning and reflection, information retrieval, and answer generation. However, existing reward models (RMs) and evaluation benchmarks are primarily designed for static single-turn tasks, failing to capture the full-pipeline complexity of DeepSearch workflows. To address this limitation, we propose F2DR, a fine-grained full-pipeline DeepSearch reward framework. F2DR evaluates DeepSearch workflows across three dimensions: Content, Trajectory, and Answer, enabling comprehensive process-level assessment. We further construct DeepSearch RM-Bench, a dedicated benchmark for evaluating RMs in DeepSearch scenarios. Extensive experiments demonstrate that F2DR achieves significantly higher evaluation consistency than self-evaluation-based baselines, while DeepSearch RM-Bench exhibits strong discriminative capability across existing open-source RMs. We will publicly release the complete DeepSearch RM-Bench dataset soon.

Explore similar work

CardsList
  1. G-ReAct: Graph-Guided Deep Search via Structure-State Co-Evolution

    Aug 2, 2026Shaoxiong Yang, Mengyuan Zhang, Shaojun Lin +4DeepseekInference-Time