Dr.Credit: Rubric-Grounded Process Credit Assignment for Deep Research Agents
Authors: Yingjian Zhu, Zhenyi Wang, Jiaxin Guo, Kun Ding, Ying Wang, Shen Huang, Xunjie Zhu, Pengjun Xie, +1 more
Organizations: School of Artificial Intelligence, University of Chinese Academy of Sciences · State Key Laboratory of Multimodal Artificial Intelligence Systems (MAIS), Institute of Automation, Chinese Academy of Sciences · Alibaba Token Hub, Alibaba Group · Peking University
Rubric-based tasks are increasingly addressed through reinforcement learning (RL), with rubric scores used as training rewards. However, these rewards typically supervise final answers without distinguishing the contributions of intermediate decisions. Many existing credit assignment methods rely on ground-truth answers to define process rewards, limiting their applicability to open-ended tasks without canonical solutions. To address this limitation, the proposed rubric-grounded credit uses task requirements as a shared reference for final answer evaluation and process supervision. The information returned by tools is assessed for the additional support it provides toward satisfying each rubric relative to that rubric's history of accepted support. By referencing these histories, credit distinguishes new support from evidence already present in the trajectory while recognizing partial support for each rubric. Dr.Credit uses rubric-grounded credit to supervise intermediate tool turns in an RL framework for deep research agents. The resulting process advantages are combined with GRPO outcome advantages to guide research decisions while retaining supervision of final-report quality. Evaluations on four in-domain and out-of-domain benchmarks show that Dr.Credit outperforms the evaluated open deep research baselines on every primary metric and submetric. Meanwhile, with an 8B-parameter backbone, the trained agent achieves average performance competitive with the evaluated frontier proprietary models. Further analyses suggest more efficient evidence acquisition and higher-quality reports under limited research-turn budgets, motivating the extension of rubric-grounded process supervision to a broader range of rubric-based tasks.
Figures & tables
Figure 1: Credit assignment for rubric-based agent training. (a) Intermediate turns share an outcome advantage. (b) Ground-truth answers guide turn-level credit. (c) Task rubrics ground turn-level credit in additional evidence support while outcome supervision is retained. In (c), each rubric’s track darkens as evidence accumulates across turns. Turn colors denote normalized advantages.
Figure 2: Overall framework of Dr.Credit , which incorporates rubric-grounded process credit into reinforcement learning for deep research agents. Left: the policy generates a group of rollouts for each query through ReAct interactions with the web environment. Center: an LLM judge assesses the additional support each tool turn provides relative to per-rubric evidence histories within its rollout, then updates the histories with accepted support points. Right: process credits are normalized across research turns in the rollout group and added to GRPO outcome advantages. Bottom: an example rollout Oi . Black dashed boxes mark tool responses masked out of the loss.
Figure 3: Composite credit for Search turns in Dr.Credit. Upper green branch: history-aware Visit credits are attributed to the search through URL matching. Lower yellow branch: an independent judge matches snippets from results without later visits to rubrics, without reading or updating Visit histories. Navigation and snippet credits are added to obtain Search credit.
Methods
DeepRubric (val)
ResearchQA
DeepResearchBench
ResearchRubrics
Average
Overall
Factual
Logical
Overall
Comp.
Depth
Instr.
Read.
Frontier Proprietary Models
Kimi K3 + Our Tools
71.5
55.5
90.2
76.7
48.4
47.2
47.5
50.2
49.5
58.0
63.6
GPT-5.6-luna + Our Tools
73.2
55.6
93.9
73.2
48.8
47.3
48.4
50.4
49.8
56.6
63.0
DeepSeek-V4-Pro + Our Tools
68.4
52.7
87.0
78.0
44.5
43.5
43.7
46.1
45.6
50.8
60.4
Opus 4.8 + Our Tools
71.0
55.2
89.7
72.9
45.5
43.6
44.8
47.7
47.6
51.5
60.2
Table 1: Overall performance on four deep research benchmarks. Dr.Credit outperforms all open deep research baselines on every metric, including all submetrics, and is competitive with proprietary systems. Average is the unweighted mean of the four benchmark-level scores. Bold denotes the best scores across the lower three model groups. ∗ and † indicate results taken from DeepRubric ( Zhu et al., 2026a ) and ResearchRubrics ( Sharma et al., 2026 ) , respectively.
Visit
Navigation
Snippets
Discount
DRub
RQA
DRB
RR
Avg
✓
✗
✗
✗
68.8
76.8
44.7
48.6
59.7
✓
✓
✗
✗
70.8
78.4
45.0
47.2
60.4
✓
✓
✓
✗
70.6
79.8
46.1
50.3
61.7
✓
✓
✗
✓
68.2
77.3
44.5
45.9
59.0
✓
✓
✓
✓
68.0
78.0
45.3
48.1
59.9
Table 2: Ablation of process credit designs. Visit denotes history-aware credit for Visit turns; gray shading marks the final Dr.Credit configuration. DRub, RQA, DRB, and RR denote DeepRubric (val), ResearchQA, DeepResearchBench, and ResearchRubrics, respectively. Avg is the arithmetic mean of the four benchmark scores; bold marks the best score per benchmark and the highest Avg.
Removed
Δ DRub
Δ RQA
Δ DRB
Δ Avg 3
(a) Outcome and Process Supervision
Process Arg
-3.2
-4.4
-2.9
-3.5
Outcome AO
-1.2
-3.6
-2.1
-2.3
(b) Evidence History
History
-3.0
-3.9
-2.1
-3.0
Table 3: Ablations of Dr.Credit: (a) removing either outcome or process advantage from research-turn supervision; (b) removing evidence history when computing process credit. Each ablation is relative to full Dr.Credit. Values are score changes, with Δ Avg 3 averaged across the three benchmarks.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Stage
Hyperparameter
Value
SFT
Optimizer
AdamW
Learning rate
1×10−5
Learning-rate schedule
Cosine; 30% warmup
Per-device batch size / accumulation steps
1 / 4
Training epochs / selected epoch
3 / 2
Precision
BF16
Appendix
Table 4: SFT and RL hyperparameters. The cosine schedule spans three SFT epochs; both RL methods start from the checkpoint after epoch two. Interaction limits are specified in Appendix A.3 .
Parameter
Setting
W1↓
RMSE ↓
Sign agreement ↑
Process
Fused
Visit levels
(0,0.25,0.5)
0.044
0.136
97.1
97.5
(0,0.2,1.0)
0.012
0.178
98.5
98.4
Shared cap
0.75
0.066
0.179
96.1
96.6
1.25
≤0.077
≤0.165
≥96.9
≥96.9
Snippet η
0.025
0.007
0.066
99.2
99.0
Appendix
Table 5: Sensitivity of normalized credit signals to parameter changes, with settings selected for signal stability from a 53-configuration sweep. The reference uses Visit levels (0,0.2,0.5) , cap 1 , and snippet weight 0.1 ; each row changes only the indicated parameters across the same 519 rollout groups and 32,966 turns. Inequalities denote conservative bounds from missing snippet judgments, while all other values are exact and sign agreement across corresponding turns is reported in percent.
Stage
GRPO
Dr.Credit
Reduction
Rollout collection
49.41
42.89
13.2%
Policy update
17.12
13.91
18.8%
Total
75.14
63.96
14.9%
Appendix
Table 6: Recorded training time over steps 1–200 in hours. Rollout collection includes scoring; the total covers all work within the training-step timer, including stages beyond those listed. Row-wise reductions use GRPO as the reference, under the same accelerator configuration.
Measurement
GRPO
Dr.Credit
Full scoring latency (s/trajectory)
2.40
23.03
Process Judge tasks per trajectory
0
29.09
Appendix
Table 7: Scoring workload over steps 1–200. Latency is averaged over trajectories within each step, then over steps. Task counts exclude deterministic zeros and do not count retries separately.