Dr.Credit: Rubric-Grounded Process Credit Assignment for Deep Research Agents
Authors: Yingjian Zhu, Zhenyi Wang, Jiaxin Guo, Kun Ding, Ying Wang, Shen Huang, Xunjie Zhu, Pengjun Xie, +1 more
Organizations: School of Artificial Intelligence, University of Chinese Academy of Sciences · State Key Laboratory of Multimodal Artificial Intelligence Systems (MAIS), Institute of Automation, Chinese Academy of Sciences · Alibaba Token Hub, Alibaba Group · Peking University
Rubric-based tasks are increasingly addressed through reinforcement learning (RL), with rubric scores used as training rewards. However, these rewards typically supervise final answers without distinguishing the contributions of intermediate decisions. Many existing credit assignment methods rely on ground-truth answers to define process rewards, limiting their applicability to open-ended tasks without canonical solutions. To address this limitation, the proposed rubric-grounded credit uses task requirements as a shared reference for final answer evaluation and process supervision. The information returned by tools is assessed for the additional support it provides toward satisfying each rubric relative to that rubric's history of accepted support. By referencing these histories, credit distinguishes new support from evidence already present in the trajectory while recognizing partial support for each rubric. Dr.Credit uses rubric-grounded credit to supervise intermediate tool turns in an RL framework for deep research agents. The resulting process advantages are combined with GRPO outcome advantages to guide research decisions while retaining supervision of final-report quality. Evaluations on four in-domain and out-of-domain benchmarks show that Dr.Credit outperforms the evaluated open deep research baselines on every primary metric and submetric. Meanwhile, with an 8B-parameter backbone, the trained agent achieves average performance competitive with the evaluated frontier proprietary models. Further analyses suggest more efficient evidence acquisition and higher-quality reports under limited research-turn budgets, motivating the extension of rubric-grounded process supervision to a broader range of rubric-based tasks.
Figures & tables
Figure 1: Credit assignment for rubric-based agent training. (a) Intermediate turns share an outcome advantage. (b) Ground-truth answers guide turn-level credit. (c) Task rubrics ground turn-level credit in additional evidence support while outcome supervision is retained. In (c), each rubric’s track darkens as evidence accumulates across turns. Turn colors denote normalized advantages.
Figure 2: Overall framework of Dr.Credit , which incorporates rubric-grounded process credit into reinforcement learning for deep research agents. Left: the policy generates a group of rollouts for each query through ReAct interactions with the web environment. Center: an LLM judge assesses the additional support each tool turn provides relative to per-rubric evidence histories within its rollout, then updates the histories with accepted support points. Right: process credits are normalized across research turns in the rollout group and added to GRPO outcome advantages. Bottom: an example rollout Oi . Black dashed boxes mark tool responses masked out of the loss.
Figure 3: Composite credit for Search turns in Dr.Credit. Upper green branch: history-aware Visit credits are attributed to the search through URL matching. Lower yellow branch: an independent judge matches snippets from results without later visits to rubrics, without reading or updating Visit histories. Navigation and snippet credits are added to obtain Search credit.
Methods
DeepRubric (val)
ResearchQA
DeepResearchBench
ResearchRubrics
Average
Overall
Factual
Logical
Overall
Comp.
Depth
Instr.
Read.
Frontier Proprietary Models
Kimi K3 + Our Tools
71.5
55.5
90.2
76.7
48.4
47.2
47.5
50.2
49.5
58.0
63.6
GPT-5.6-luna + Our Tools
73.2
55.6
93.9
73.2
48.8
47.3
48.4
50.4
49.8
56.6
63.0
DeepSeek-V4-Pro + Our Tools
68.4
52.7
87.0
78.0
44.5
43.5
43.7
46.1
45.6
50.8
60.4
Opus 4.8 + Our Tools
71.0
55.2
89.7
72.9
45.5
43.6
44.8
47.7
47.6
51.5
60.2
Table 1: Overall performance on four deep research benchmarks. Dr.Credit outperforms all open deep research baselines on every metric, including all submetrics, and is competitive with proprietary systems. Average is the unweighted mean of the four benchmark-level scores. Bold denotes the best scores across the lower three model groups. ∗ and † indicate results taken from DeepRubric ( Zhu et al., 2026a ) and ResearchRubrics ( Sharma et al., 2026 ) , respectively.
Visit
Navigation
Snippets
Discount
DRub
RQA
DRB
RR
Avg
✓
✗
✗
✗
68.8
76.8
44.7
48.6
59.7
✓
✓
✗
✗
70.8
78.4
45.0
47.2
60.4
✓
✓
✓
✗
70.6
79.8
46.1
50.3
61.7
✓
✓
✗
✓
68.2
77.3
44.5
45.9
59.0
✓
✓
✓
✓
68.0
78.0
45.3
48.1
59.9
Table 2: Ablation of process credit designs. Visit denotes history-aware credit for Visit turns; gray shading marks the final Dr.Credit configuration. DRub, RQA, DRB, and RR denote DeepRubric (val), ResearchQA, DeepResearchBench, and ResearchRubrics, respectively. Avg is the arithmetic mean of the four benchmark scores; bold marks the best score per benchmark and the highest Avg.
Removed
Δ DRub
Δ RQA
Δ DRB
Δ Avg 3
(a) Outcome and Process Supervision
Process Arg
-3.2
-4.4
-2.9
-3.5
Outcome AO
-1.2
-3.6
-2.1
-2.3
(b) Evidence History
History
-3.0
-3.9
-2.1
-3.0
Table 3: Ablations of Dr.Credit: (a) removing either outcome or process advantage from research-turn supervision; (b) removing evidence history when computing process credit. Each ablation is relative to full Dr.Credit. Values are score changes, with Δ Avg 3 averaged across the three benchmarks.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Stage
Hyperparameter
Value
SFT
Optimizer
AdamW
Learning rate
1×10−5
Learning-rate schedule
Cosine; 30% warmup
Per-device batch size / accumulation steps
1 / 4
Training epochs / selected epoch
3 / 2
Precision
BF16
Appendix
Table 4: SFT and RL hyperparameters. The cosine schedule spans three SFT epochs; both RL methods start from the checkpoint after epoch two. Interaction limits are specified in Appendix A.3 .
Parameter
Setting
W1↓
RMSE ↓
Sign agreement ↑
Process
Fused
Visit levels
(0,0.25,0.5)
0.044
0.136
97.1
97.5
(0,0.2,1.0)
0.012
0.178
98.5
98.4
Shared cap
0.75
0.066
0.179
96.1
96.6
1.25
≤0.077
≤0.165
≥96.9
≥96.9
Snippet η
0.025
0.007
0.066
99.2
99.0
Appendix
Table 5: Sensitivity of normalized credit signals to parameter changes, with settings selected for signal stability from a 53-configuration sweep. The reference uses Visit levels (0,0.2,0.5) , cap 1 , and snippet weight 0.1 ; each row changes only the indicated parameters across the same 519 rollout groups and 32,966 turns. Inequalities denote conservative bounds from missing snippet judgments, while all other values are exact and sign agreement across corresponding turns is reported in percent.
Stage
GRPO
Dr.Credit
Reduction
Rollout collection
49.41
42.89
13.2%
Policy update
17.12
13.91
18.8%
Total
75.14
63.96
14.9%
Appendix
Table 6: Recorded training time over steps 1–200 in hours. Rollout collection includes scoring; the total covers all work within the training-step timer, including stages beyond those listed. Row-wise reductions use GRPO as the reference, under the same accelerator configuration.
Measurement
GRPO
Dr.Credit
Full scoring latency (s/trajectory)
2.40
23.03
Process Judge tasks per trajectory
0
29.09
Appendix
Table 7: Scoring workload over steps 1–200. Latency is averaged over trajectories within each step, then over steps. Task counts exclude deterministic zeros and do not count retries separately.
Deep research agents synthesize long-form reports by searching and reasoning over retrieved evidence. Reinforcement learning with rubric-based rewards improves these agents by optimizing them against checkable criteria that translate report quality into reward signals, but its efficiency depends on whether those criteria reliably capture the task scope and evidence needs. Most existing studies ask an LLM to generate rubrics for a given query, but when the model fails to infer the underlying information needs, the generated rubrics may be incomplete and reduce RL efficiency. To obtain more reliable query--rubric supervision, we introduce DeepRubric, a data construction framework that reverses this process: instead of inferring evaluation criteria for a given query, it first determines what an evidence-backed report should be evaluated on and then synthesizes aligned query--rubric pairs from those evaluation targets. Starting from a sampled seed topic, DeepRubric builds an evidence tree by recursively expanding evidence-backed sub-questions, whose leaves serve as atomic and verifiable evaluation targets. It then uses the evidence tree to synthesize the training query and rubrics, ensuring that the reward evaluates exactly the information requested by the query. Using DeepRubric, we construct 9K query--rubric supervision examples and train DeepRubric-8B with rubric-based GRPO, achieving comparable performance to prior open state-of-the-art deep research models across three benchmarks with roughly 13x fewer RL GPU-hours.
Minghang Zhu, Chuyang Wei, Junhao Xu +3
Shandong University, Qingdao, China · Zhongguancun Academy, Beijing, China · Fudan University, Shanghai, China
Open-ended reasoning and long-form generation tasks lack reliable automatic verification signals for reward-based policy optimization. Rubrics offer a promising alternative, but existing approaches treat them as given artifacts -- either hand-crafted or prompt-generated -- and often miss the task-specific, knowledge-intensive dimensions that matter most, distorting the reward signal. Our key observation is that rubric construction is itself a research problem: identifying what makes a response correct or insightful requires discovering and synthesizing external knowledge. We propose Deep Research as Rubric (DR-rubric), a two-stage framework for constructing such rubrics. Stage I elicits domain facts, structural constraints, and failure modes through iterative multi-turn agentic search; Stage II distills this evidence into atomic, independently verifiable constraints for GRPO-based policy optimization. Because the model under training can serve as its own rubric generator, DR-rubric-8B supports bootstrap rubric generation without frontier-model assistance. We evaluate on 6 benchmarks spanning agentic research and expert reasoning. Experiments show that DR-Rubric achieves strong competitive performance with only 1K -- 3K training instances, where GPT-5-generated rubrics particularly benefit breadth coverage on agentic tasks, Gemini-generated rubrics yield the most balanced performance across agentic and expert reasoning tasks, and bootstrap rubrics exhibit a specialization-to-rebalancing evolution achieving the best overall performance at the third iteration. Results demonstrate that reframing rubric construction from static evaluation templates into an evidence-driven research process yields more scalable, fine-grained reward signals for open-ended tasks.
Wangyi Mei, Zhouhong Gu, Zhenhan Bai +9
Fudan University · Xiaohongshu Inc. · Beijing University of Posts and Telecommunications
Training deep research agents, namely systems that plan, search, evaluate evidence, and synthesize long-form reports, pushes reinforcement learning beyond the regime of verifiable rewards. Their outputs lack ground-truth answers, their trajectories span many tool-augmented decisions, and standard post-training offers little mechanism for turning past attempts into reusable experience. In this work, we argue that rubrics should serve not merely as final-answer evaluators, but as the shared interface that structures policy execution, judge feedback, and agent memory. Based on this view, we introduce RubricEM, a rubric-guided reinforcement learning framework that combines stagewise policy decomposition with reflection-based meta-policy evolution. RubricEM first makes research trajectories stage-aware by conditioning planning, evidence gathering, review, and synthesis on self-generated rubrics. It then assigns credit with Stage-Structured GRPO, which uses stagewise rubric judgments to provide denser semantic feedback for long-horizon optimization. In parallel, RubricEM trains a shared-backbone reflection meta-policy that distills judged trajectories into reusable rubric-grounded guidance for future attempts. The resulting RubricEM-8B achieves strong performance across four long-form research benchmarks, outperforming comparable open models and approaching proprietary deep-research systems. Beyond final performance, we perform thorough analyses to understand the key ingredients of RubricEM.
Gaotang Li, Bhavana Dalvi Mishra, Zifeng Wang +9
University of Illinois Urbana-Champaign · Google Cloud AI Research