RINI: Seeing the Prior Is Not Enough
Organizations: University of Illinois Urbana-Champaign · Harvey Mudd College · National University of Singapore · Wellesley College · Cornell University
Abstract
A research proposal can describe an established mechanism correctly while claiming to introduce it. We study whether providing the earlier paper corrects such contribution claims. Three controlled experiments compare proposals generated with a contribution-bearing prior and a same-topic control. Providing the prior yields no clear aggregate reduction in unsupported novelty. Human analysis of 175 interpretable exposed proposals finds that 137 recognize the prior's relevance, but 61 correctly attribute the established contribution. Of 71 proposed remaining distinctions, 37 are covered by the same prior. We introduce Research Idea Novelty Inspection (RINI), which audits contribution claims against evidence, checks the remaining distinction, and applies local revisions. Five human annotators evaluate 1,080 original-revision pairs across three methods. On the same 240 originals judged to require correction, successful repair is 11.7% for Self-Revision, 39.1% for Retrieve-and-Revise, and 72.2% for RINI, with research tasks weighted equally. The improvement over same-evidence direct revision is 33.0 percentage points. The revised proposals retain their research questions and technical methods. These results motivate explicit contribution attribution when using literature to generate and revise research proposals.
Figures & tables
| Stage | Main work |
| Candidate generation | Model drafts 376 paper-anchored seeds; deduplication retains 315. |
| Human seed review | Three reviewers and adjudication retain 314 seeds. |
| Evidence curation | Select reference papers, quote evidence, and cross-review 62 bundles. |
| Author finalization | Form a 32-seed pool; hash ordering selects 30 for Study 1. |
| Round | Seeds | estimate | 95% confidence interval |
| Study 1 | 30 | ||
| Study 2 | 30 | ||
| Study 3 | 30 | ||
| Combined | 71 |
| Behavior | Count | Seed-balanced rate |
| Recognizes prior–target relation | 137/175 | 78.3% |
| Correctly attributes target | 61/175 | 34.8% |
| Contracts target claim | 69/175 | 39.2% |
| Valid relocation | 23/165 | 13.8% |
| Coverage of the 71 proposed remainders | ||
| Covered by the same prior | 37/71 | — |
| Policy | Added prior-work evidence | Revision procedure |
| Self-Revision | None | Direct revision from the original proposal and target. |
| Retrieve-and-Revise | Decisive prior and the – relation | Direct evidence-grounded revision of the full proposal. |
| RINI | Same and – relation | Contribution audit, remainder check, and bounded local repair. |
| Method | Successful repair | Correct target attribution | Still needs revision |
| Self-Revision | 11.7% [-0.1ex] [7.4,16.2] | 12.0% [-0.1ex] [7.5,16.7] | 87.2% [-0.1ex] [82.6,91.6] |
| Retrieve-and-Revise | 39.1% [-0.1ex] [32.9,45.7] | 53.2% [-0.1ex] [47.2,59.3] | 60.9% [-0.1ex] [54.3,67.1] |
| RINI | 72.2% [-0.1ex] [67.8,76.6] | 75.3% [-0.1ex] [70.7,80.0] | 25.9% [-0.1ex] [21.1,30.7] |
| RINI RR | +33.0 pp [-0.1ex] [24.5,41.4] | +22.1 pp [-0.1ex] [14.6,29.8] | -35.0 pp [-0.1ex] [-43.4,-26.4] |
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
| Count | Unit | Derivation | Use in analysis |
|---|---|---|---|
| 30 | Study 3 research seeds | Frozen evidence-anchored cohort | The 30 resampling clusters used for seed-balanced analyses. |
| 3 / 2 | Source generators / repeats | Three generator families, two repeats each | Qwen, GLM, and Kimi; repeated observations within each seed. |
| 180 | Matched H/E pairs | One H and one E proposal for each seed–generator–repeat block. | |
| 360 | Frozen original proposals | (H and E) | Both originals in every pair enter the repair benchmark. |
| 540 | Machine-scored proposals | (W, H, E) | The neutral W arm is scored in the exposure study but is not a repair source. |
| 1,080 | Formal repair items | revision policies | Self-Revision, RR, and RINI; one human annotation for each original–revision item. |
| Corpus field | Recorded content |
|---|---|
| Paper count | 50,938 collected; six title duplicates dropped (lowest identifier kept); 50,932 frozen. |
| Corpus scope | cs.CV 18,233, cs.LG 17,147, cs.CL 15,552; 2018: 9,282, each of 2019–2023: 8,328–8,331. |
| Retrieval | Reciprocal-rank fusion of BM25 and SPECTER2 over titles and abstracts of all 50,932 papers. |
| Full-text availability | Fetched for 3,024 curation candidates: 2,797 converted (92.5%); 108 HTTP 404/406, one network failure, 118 too short after conversion. |
| Version and date | Record the exact source version used for contribution coverage and its eligibility under the case cutoff. |
| Duplicates | Preserve a canonical paper identity and a deterministic treatment of versions; avoid giving one work multiple ranking opportunities. |
| Flow stage | Recorded preparation and execution |
|---|---|
| Candidate frame | 376 model-generated candidates; SPECTER2 deduplication removes 61 and retains 315 for human review. |
| Neutrality review | Three reviewers; 711 review rows. Of 315 seeds, 296 pass directly and 18 are retained after adjudicated revision, giving 314. The 236-row packet is one reviewer’s assignment. |
| Evidence construction | Two human curation rounds produce 62 bundles. Author finalization yields 32 seeds; fixed-key ordering selects 30. Study 2 selects 19 original-pool and 11 machine-curated seeds. Study 3 separately supplies 30 cases and 60 pinned-prior fetches. |
| Generation | Study 3: 540 W/H/E outputs in 180 complete blocks, including 360 H/E repair sources. The three revision methods produce 1,080 final conditions. The revision driver records 273 failed attempts across retries. |
| Assessment | Five human annotators complete 216 shuffled repair items each. One separate colleague pilots 90 items before the formal rubric is fixed. |
| Analysis | Repair uses 30 seeds, three source generators, two repeats, and H/E sources. Equal seed weights sum to one. The primary repair cohort contains 240 original proposals. |
| Threshold | Canonical predicate |
|---|---|
| : contribution | Any novelty or contribution beyond repeating known work. |
| : beyond application | More than ordinary application, adaptation, or setting-specific variation. |
| : substantive advance | Substantive methodological, mechanistic, empirical, or evaluative novelty. |
| : strongest force | Firstness, superiority, uniqueness, or field-level novelty. |
| Boundary case | Interpretation |
|---|---|
| Absent from the supplied papers | Retain the explicitly bounded comparison; do not expand it to the entire field. |
| First application of a known method to dataset X | Local firstness may coexist with routine application. Preserve the dataset and application scope when interpreting the machine score. |
| A first retrieval mechanism and a new controller | Separate independently challengeable contributions; evidence covering the mechanism need not cover the controller. |
| Higher accuracy on benchmark X | Distinguish a bounded performance assertion from a new mechanism or general superiority. |
| Insufficient evidence about a distinction | Keep the evidence judgment unresolved; missing counterevidence does not prove priority. |
| Condition | Operational specification |
|---|---|
| Stable generator setting | Same recorded model identifier, instruction, tools, and decoding settings within each block. |
| No cross-arm carryover | Independent generation calls with separate outputs and evaluation feedback. |
| Valid evidence roles | Study 3 paper selection, source-span checks, and semantic evidence vetting precede generation. |
| Matched context | Rendered inputs are checked for document count, normalized word budget, format, and focal position. |
| Stable outcome target | The same decisive reference and threshold predicates are used across arms. |
| Outcome-blind inclusion | Seed and evidence eligibility precede outcomes; incomplete blocks follow the frozen disposition rule. |
| Field / category | Self | RR | RINI |
|---|---|---|---|
| Original contribution-revision need | |||
| yes | 267 | 340 | 255 |
| no | 87 | 17 | 81 |
| unclear | 5 | 2 | 23 |
| unjudgeable | 1 | 1 | 1 |
| Revised contribution-revision need | |||
| Label | Meaning |
|---|---|
| successful_repair | An original substantive contribution problem is resolved, with no new substantive contribution error and useful research content retained. |
| partial_repair | A meaningful contribution correction occurs, but an important contribution problem remains. |
| no_repair | The central contribution problem is essentially unchanged. |
| original_already_acceptable | The original needed no substantive contribution repair, and the revision preserves that acceptability. |
| regressed | Revision introduces a new substantive contribution error or causes serious loss of useful research content. |
| unclear | The supplied evidence cannot settle the outcome. |
| Field | Allowed values |
|---|---|
| Original contribution-revision need | yes , no , unclear , unjudgeable . |
| Revised contribution-revision need | yes , no , unclear , unjudgeable . |
| x_attribution | correct , partial , incorrect , unclear , not_applicable . |
| has_y | yes , no . |
| y_vs_d | covered_by_d , not_covered_by_d , unclear , not_applicable . |
| Unsupported novelty remains | yes , no , unclear . |