We study the ability of LLMs to generate edit-inducing questions whose answer will improve a paper draft. On a dataset of paired submission and camera-ready papers from ICLR and NeurIPS, we compare the helpfulness of questions from GPT models with or without full paper context to that of human reviewers. GPT produces more edit-inducing questions and its questions are associated with more extensive edits and cover a broader range of edited content compared to questions from reviewers. However, a much smaller percentage of the GPT questions are edit-inducing. Our analyses confirm that automated questions can be beneficial to authors and highlight an example task where proper attending to long context deteriorates reasoning model ability to produce helpful output.
Figures & tables
Figure 1: We seek to generate edit-inducing questions like the example shown above. A question is asked based on content in the initial submission. When answered only in the camera ready version, we can expect content to have been edited to address such a question.
Figure 2: Measures of edit worthiness of questions asked by peer reviewers (PeerQA) and GPT models.
Name
Overlap Proportion
o3
0.2337
o3 Linear
0.2337
o3 Segment
0.2306
GPT-4o
0.3020
GPT-4o Linear
0.3069
GPT-4o Segment
0.3337
Table 1: Proportion of words in the question overlapping with words found from the span of text that elicited it
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Model
Total Qs
Avg Qs/ Paper
Max Qs
Min Qs
PeerQA
362
3.12
11
1
4o full
17,734
152.88
320
38
4o linear
18,158
156.53
321
32
4o segment
16,018
138.09
416
28
o3 full
18,940
163.28
433
47
o3 linear
21,339
183.96
455
43
Appendix
Table 2: Number of questions by source.
Metric
Score
Precision
0.836
Recall
0.586
F1
0.689
Appendix
Table 3: Precision, recall, and F1 @ 5 (among the top 5 answers) for gpt-4o deciding if a segment is an answer to a given question. Evaluated against ground truth PeerQA answers.
Edit-Inducing
Unanswered
Model: o3
Method detail Could you explicitly show how a single layer of a k-GNN is encoded as a TL^(t)_{k+1}(?) expression, in particular how the higher–order neighborhood aggregation and the permutation–invariant updates are handled within only k+1 index variables?
Method detail FIND_HP requires its input point to be δ -non-degenerate, yet Step 1 only guarantees that x_i is a critical point. What mechanism or additional test ensures that each x_i returned by FIND_CP actually satisfies the δ -non-degeneracy condition needed by FIND_HP?
Insights nuance Why do poisoned and benign frames stay close in the SiamFC++ feature space whereas they separate in SiamFC and SiamRPN++? Could you elaborate on what architectural or training differences lead to this behavior?
Method detail You refer to an additional invariant layer that has a "similar (simpler) representation as given in equation 1 (Maron et al., 2019c)". Could you provide the explicit TL representation of this invariant layer and explain why it is simpler?
Method detail For the min-max and z-score baselines, are the normalization statistics computed per instance, per batch, or over the whole training set, and are they recomputed during inference?
Method detail How exactly are the adversarially generated ‘confusing samples’ produced (e.g., number of optimization steps, step size, loss function, and whether this generation is performed on-the-fly during training or pre-computed)?
Insights soundness On what empirical or theoretical grounds do you characterise δ -regularity as a “mild” general position assumption? Have you measured how frequently trained networks satisfy δ -regularity without the artificial perturbation you propose?
Method detail The sentence states that edge weights depend on both distance and direction, yet the Gaussian kernel you provide only uses the squared distance. How is the direction information actually incorporated into the adjacency weights?
Measurement more Can you provide ablation results quantifying the benefit of the parallel exploration relative to a single-process baseline to substantiate the claim of "efficient exploration"?
Method detail You claim HyperDQN measures uncertainty through "infinite" ensembles, but in implementation only finitely many z-samples can be used. How many z vectors are actually drawn per training step and per episode, and how did you decide on this number?
Appendix
Table 4: Side-by-side comparison of edit-inducing and unanswered questions.
PeerQA
Q: Do you evaluate playing strength of agents by restricting them by MCTS iteration counts or by time limits?
4o segment
Q: What motivates the choice to approximate the number of MCTS simulations ( T ) as the maximum number allowed, and could this approximation influence results in scenarios with substantial late-game positions where the full game tree may already be mapped? Originating Span: “We also approximate the number of MCTS simulations T to be the maximum number of simulations allowed, since the maximum is reached at all game positions except late-game positions, where the remaining game tree is already fully mapped.”
o3 full
Q: You approximate T by the maximum number of MCTS simulations allowed because that maximum is supposedly reached at almost all positions. Can you quantify what fraction of positions actually hit the maximum in practice and how deviations from the maximum affect the computed FLOPs and the fitted αC ? Originating Span: “We also approximate the number of MCTS simulations T to be the maximum number of simulations allowed, since the maximum is reached at all game positions except late-game positions, where the remaining game tree is already fully mapped.”
Appendix
Table 5: Table comparing similar questions asked by three different systems
Type: High Edit Rate Edit Rate: 0.617
Q: What are the global convergence criteria used when solving PESNet for a continuous subset of structures? Initial: For H 4+ and cyclobutadiene, we train on discrete sets of geometries from the literature (Scherbela et al., 2021; Kinal & Piecuch, 2007). Final: For H 4+ and cyclobutadiene… [New Paragraph]… To still access convergence, we use the fact that the local energy EL of any eigenfunction (including the ground-state) has 0 variance…
Type: Low Edit Rate Edit Rate: 0.247
Q: What is the relationship between the p∗() label and the one-hot encoding for the hard sample in Figure 3? Initial: we define base difficulty as ∥ey−p∗(x)∥2 . This will be high for ambiguous points, but especially points where the sampled y had low probability under p∗ . Final: we define base difficulty as ∥ey−p∗(x)∥2 , which is large if: ∙x is ambiguous: p∗ has several large components, so there is no one-hot label near p∗ .
Type: Invalid
Q: How was the fine tuning done for the step sizes in the experiments? Initial: except for the step sizes – we fine-tune them using a set of powers of two {2i∣i∈[−10,10]} – Final: except for the step sizes – we fine-tune them using a set of powers of two {2i∣i∈[−10,10]} –
Type: Unanswered
Q: Is there a plan to open-source the proprietary medical knowledge base and the telemedicine software? Initial: nil Final: nil
Appendix
Table 6: Examples drawn from human written questions in the PeerQA dataset highlighting the different types of questions and how they are assessed.
May 27, 2026·Hans Ole Hatzel, Sebastian Steindl, Jan StrichPeer ReviewReviewer
Language Technology Group, University of Hamburg, Germany · OTH Amberg-Weiden, Germany · Hub of Computing and Data Science (HCDS), University of Hamburg, Germany