We study the ability of LLMs to generate edit-inducing questions whose answer will improve a paper draft. On a dataset of paired submission and camera-ready papers from ICLR and NeurIPS, we compare the helpfulness of questions from GPT models with or without full paper context to that of human reviewers. GPT produces more edit-inducing questions and its questions are associated with more extensive edits and cover a broader range of edited content compared to questions from reviewers. However, a much smaller percentage of the GPT questions are edit-inducing. Our analyses confirm that automated questions can be beneficial to authors and highlight an example task where proper attending to long context deteriorates reasoning model ability to produce helpful output.
Figures & tables
Figure 1: We seek to generate edit-inducing questions like the example shown above. A question is asked based on content in the initial submission. When answered only in the camera ready version, we can expect content to have been edited to address such a question.
Figure 2: Measures of edit worthiness of questions asked by peer reviewers (PeerQA) and GPT models.
Name
Overlap Proportion
o3
0.2337
o3 Linear
0.2337
o3 Segment
0.2306
GPT-4o
0.3020
GPT-4o Linear
0.3069
GPT-4o Segment
0.3337
Table 1: Proportion of words in the question overlapping with words found from the span of text that elicited it
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Model
Total Qs
Avg Qs/ Paper
Max Qs
Min Qs
PeerQA
362
3.12
11
1
4o full
17,734
152.88
320
38
4o linear
18,158
156.53
321
32
4o segment
16,018
138.09
416
28
o3 full
18,940
163.28
433
47
o3 linear
21,339
183.96
455
43
Appendix
Table 2: Number of questions by source.
Metric
Score
Precision
0.836
Recall
0.586
F1
0.689
Appendix
Table 3: Precision, recall, and F1 @ 5 (among the top 5 answers) for gpt-4o deciding if a segment is an answer to a given question. Evaluated against ground truth PeerQA answers.
Edit-Inducing
Unanswered
Model: o3
Method detail Could you explicitly show how a single layer of a k-GNN is encoded as a TL^(t)_{k+1}(?) expression, in particular how the higher–order neighborhood aggregation and the permutation–invariant updates are handled within only k+1 index variables?
Method detail FIND_HP requires its input point to be δ -non-degenerate, yet Step 1 only guarantees that x_i is a critical point. What mechanism or additional test ensures that each x_i returned by FIND_CP actually satisfies the δ -non-degeneracy condition needed by FIND_HP?
Insights nuance Why do poisoned and benign frames stay close in the SiamFC++ feature space whereas they separate in SiamFC and SiamRPN++? Could you elaborate on what architectural or training differences lead to this behavior?
Method detail You refer to an additional invariant layer that has a "similar (simpler) representation as given in equation 1 (Maron et al., 2019c)". Could you provide the explicit TL representation of this invariant layer and explain why it is simpler?
Method detail For the min-max and z-score baselines, are the normalization statistics computed per instance, per batch, or over the whole training set, and are they recomputed during inference?
Method detail How exactly are the adversarially generated ‘confusing samples’ produced (e.g., number of optimization steps, step size, loss function, and whether this generation is performed on-the-fly during training or pre-computed)?
Insights soundness On what empirical or theoretical grounds do you characterise δ -regularity as a “mild” general position assumption? Have you measured how frequently trained networks satisfy δ -regularity without the artificial perturbation you propose?
Method detail The sentence states that edge weights depend on both distance and direction, yet the Gaussian kernel you provide only uses the squared distance. How is the direction information actually incorporated into the adjacency weights?
Measurement more Can you provide ablation results quantifying the benefit of the parallel exploration relative to a single-process baseline to substantiate the claim of "efficient exploration"?
Method detail You claim HyperDQN measures uncertainty through "infinite" ensembles, but in implementation only finitely many z-samples can be used. How many z vectors are actually drawn per training step and per episode, and how did you decide on this number?
Appendix
Table 4: Side-by-side comparison of edit-inducing and unanswered questions.
PeerQA
Q: Do you evaluate playing strength of agents by restricting them by MCTS iteration counts or by time limits?
4o segment
Q: What motivates the choice to approximate the number of MCTS simulations ( T ) as the maximum number allowed, and could this approximation influence results in scenarios with substantial late-game positions where the full game tree may already be mapped? Originating Span: “We also approximate the number of MCTS simulations T to be the maximum number of simulations allowed, since the maximum is reached at all game positions except late-game positions, where the remaining game tree is already fully mapped.”
o3 full
Q: You approximate T by the maximum number of MCTS simulations allowed because that maximum is supposedly reached at almost all positions. Can you quantify what fraction of positions actually hit the maximum in practice and how deviations from the maximum affect the computed FLOPs and the fitted αC ? Originating Span: “We also approximate the number of MCTS simulations T to be the maximum number of simulations allowed, since the maximum is reached at all game positions except late-game positions, where the remaining game tree is already fully mapped.”
Appendix
Table 5: Table comparing similar questions asked by three different systems
Type: High Edit Rate Edit Rate: 0.617
Q: What are the global convergence criteria used when solving PESNet for a continuous subset of structures? Initial: For H 4+ and cyclobutadiene, we train on discrete sets of geometries from the literature (Scherbela et al., 2021; Kinal & Piecuch, 2007). Final: For H 4+ and cyclobutadiene… [New Paragraph]… To still access convergence, we use the fact that the local energy EL of any eigenfunction (including the ground-state) has 0 variance…
Type: Low Edit Rate Edit Rate: 0.247
Q: What is the relationship between the p∗() label and the one-hot encoding for the hard sample in Figure 3? Initial: we define base difficulty as ∥ey−p∗(x)∥2 . This will be high for ambiguous points, but especially points where the sampled y had low probability under p∗ . Final: we define base difficulty as ∥ey−p∗(x)∥2 , which is large if: ∙x is ambiguous: p∗ has several large components, so there is no one-hot label near p∗ .
Type: Invalid
Q: How was the fine tuning done for the step sizes in the experiments? Initial: except for the step sizes – we fine-tune them using a set of powers of two {2i∣i∈[−10,10]} – Final: except for the step sizes – we fine-tune them using a set of powers of two {2i∣i∈[−10,10]} –
Type: Unanswered
Q: Is there a plan to open-source the proprietary medical knowledge base and the telemedicine software? Initial: nil Final: nil
Appendix
Table 6: Examples drawn from human written questions in the PeerQA dataset highlighting the different types of questions and how they are assessed.
Local factual edits in scientific manuscripts often create non-local revision obligations. If a dataset changes from 215 to 80 documents, claims such as 'medium-scale' or 'a few hundred items' may also become stale, even though they do not repeat the edited number. In an audit of recent arXiv cs.CL benchmark and dataset papers, we find fact-dependent qualitative claims in 37.2% of papers, suggesting that this dependency pattern is common in the target genre. We introduce EditPropBench, a benchmark for measuring whether LLM editors propagate factual edits through dependent manuscript claims. Each item contains an ML/NLP-style synthetic manuscript, a targeted edit, and a controlled fact graph with sentence-level labels for direct targets, required downstream updates, and unrelated text that should remain unchanged. We summarize cascade success with Edit-Ripple Adherence (ERA), the fraction of required downstream updates correctly revised, and validate the metric with adversarial probes and stress-test variants. On the hardest cases, where dependent claims use implicit or free-form wording rather than repeating the edited value, five LLM editing systems span ERA 0.148-0.705. Even the strongest misses roughly 30% of required cascade updates. This advantage persists in a mixed evaluation that includes easy cases solvable by deterministic substitution. EditPropBench shows that current LLM editors can repair many implicit consequences of factual edits, but reliable scientific revision still requires cascade-aware checking.
As researchers increasingly adopt LLMs as writing assistants, generating high-quality research paper introductions remains both challenging and essential. We introduce Scientific Introduction Generation (SciIG), a task that evaluates LLMs' ability to produce coherent introductions from titles, abstracts, and related works. Curating new datasets from NAACL 2025 and ICLR 2025 papers, we assess five state-of-the-art models, including both open-source (DeepSeek-v3, Gemma-3-12B, LLaMA 4-Maverick, MistralAI Small 3.1) and closed-source GPT-4o systems, across multiple dimensions: lexical overlap, semantic similarity, content coverage, faithfulness, consistency, citation correctness, and narrative quality. Our comprehensive framework combines automated metrics with LLM-as-a-judge evaluations. Results demonstrate LLaMA-4 Maverick's superior performance on most metrics, particularly in semantic similarity and faithfulness. Moreover, three-shot prompting consistently outperforms fewer-shot approaches. These findings provide practical insights into developing effective research writing assistants and set realistic expectations for LLM-assisted academic writing. To foster re- producibility and future research, we publicly release all code and datasets.
LLM-generated reviews for scientific papers are gaining considerable traction and are even being officially piloted by major conferences. We have to assume that not only reviewers are using LLM-assistance, but also that authors use LLMs to revise their papers before submitting. In this work, we perform empirical experiments on papers from the 2025 ACL Rolling Review (ARR) to evaluate LLM reviews from both the author and the reviewer perspective. First, we identify a limited alignment of LLM reviews with human ones. In the best-case scenario, the alignment is reasonable. However, we also find that LLM-human alignment varies substantially across prompts and models. Finally, we investigate the scenario in which the author uses an iterative draft-revise workflow to improve the submission according to the LLM review. We find that this "gaming" of LLM reviews can be effective in specific scenarios, leading to a statistically significant increase of overall scores for up to 35% of papers. We publish our code: https://github.com/uhh-hcds/reviewarcade.
Hans Ole Hatzel, Sebastian Steindl, Jan Strich
Language Technology Group, University of Hamburg, Germany · OTH Amberg-Weiden, Germany · Hub of Computing and Data Science (HCDS), University of Hamburg, Germany