SIVIA-RSI: Source-Grounded Adaptation of Diagramming Skills
Authors: Feng Yuan, Yifan Gao, Haoyue Li, Xin Gao
Organizations: School of Biomedical Engineering (Suzhou), Division of Life Science and Medicine, University of Science and Technology of China, Hefei, China · Suzhou Institute of Biomedical Engineering and Technology, Chinese Academy of Sciences, Suzhou, China · Shanghai Innovation Institute, Shanghai, China
Scientific method diagrams express computations through entities, dependencies, and conditional routes. Although generated figures can be improved through repeated editing, it is less clear whether experience from one paper improves the first figure of another. We present \sys, a framework for source-grounded adaptation of reusable diagramming skills, and study transfer through a complete-candidate evaluation. The framework links critiques to source passages, proposes bounded edits to a persistent skill library, and separates candidate competition from skill acceptance. We evaluate the original skill and four learned candidates on two NLP and language-agent papers, with two fresh generations per condition. The strongest candidate attains 87.50% required-relation accuracy compared with 83.33% for the original, while candidate behavior differs across papers. Local improvements on a separate development paper and automatic selector preferences do not establish consistent transfer. Tracing all 22 non-correct relation judgments to their production prompts reveals both incomplete conditional specifications and ambiguities despite explicit instructions. All ten planning diagrams leave an already-terminal selected leaf's route unclear; none of their prompts explicitly binds that route. Our findings show why evaluating reusable diagram skills requires source-grounded relation assessment, complete candidate coverage, and inspection of both prompts and images. We provide all twenty transfer outputs, skill snapshots, assessment records, and executable analyses.
Figures & tables
Figure 1: Source-grounded skill adaptation and complete-candidate auditing. (a) Discovery evidence informs bounded edits to a reusable skill and, optionally, its updater policy. Separate papers support candidate competition and parent acceptance. The two learning runs produce four proposals, all retained regardless of selection. (b) The original skill and four proposals are frozen for generation on RAG and RAP: five skills, two papers, and two fresh repetitions yield twenty first outputs. Source-verified relations and additional mechanism errors are assessed after generation, without target-specific repair. The document and diagram miniatures are schematic; measured outputs appear in Figures 3 and 4 .
Skill
Tokens
Candidate origin and emphasis
S0
24,240
Original nine-file skill.
S1
24,348
Fixed run, first proposal; state ownership and consumers.
S2
24,330
Fixed run, second proposal; state interfaces and layout numbers.
Table 1: All frozen skill identities. The automatic selector advanced S1 and S3 ; the transfer study retains all four proposals. Token counts use a shared proxy serialization.
Required relations correct
Central error: yes / uncertain
Skill
RAG
RAP
Macro (%)
Δ (pp)
RAG (2 images)
RAP (2 images)
Extra unsupported
Pass gate
S0
10/12
10/12
83.33
+0.00
0 / 1
0 / 0
0
—
S1
11/12
10/12
87.50
+4.17
1 / 0
0 / 0
0
No
S2
9/12
10/12
79.17
-4.17
1 / 0
1 / 0
0
No
S3
11/12
9/12
83.33
+0.00
0 / 1
0 / 0
0
No
S4
11/12
7/12
75.00
-8.33
0 / 1
0 / 2
1
No
Table 2: Complete frozen transfer matrix. Macro averages repetitions within each paper and then the two paper families equally. Extra unsupported counts cover four images per skill. Uncertain central-error labels remain unknown. All candidates fail the primary 10-percentage-point continuation gate, independently of additional-error constraints. This is a descriptive development comparison with two paper families.
Figure 2: First-output relation accuracy for every frozen skill. Bars show the mean of two independent design-and-render repetitions; circles and diamonds show the individual images. Dashed lines mark the original skill’s paper-level mean. The same candidate can improve RAG while weakening RAP. The plot summarizes observed outputs, with no confidence intervals or independent-paper interpretation of the repetitions.
Reflexion pair
Relations
Fidelity
Author choice
Fixed run
2/4→4/4
2→4
Updated
Evolving run
3/4→4/4
3→4
Updated
Table 3: Local parent-to-candidate comparisons on Reflexion under the initial rubric. Fidelity uses a 1–5 scale. Both updated skills were proposed by the original updater; the pairs concern one paper.
Prompt-side diagnosis
RAG
RAP
Total
Explicit requirement
6
2
8
Partial / unbound requirement
1
11
12
Absent critical requirement
0
1
1
Generation–evaluation scope mismatch
1
0
1
Table 4: Assistant-authored post-hoc trace of all 22 non-correct human relation judgments. Counts partition these observed items; they are neither additional human labels nor causal error probabilities.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Paper
Required relation (short name)
Correct / 10
RAG
Query to retrieval
10/10
Input and document operands
7/10
Weighted marginalization
9/10
Sequence formulation
10/10
Token formulation
8/10
Training ownership
8/10
Appendix
Table 5: Descriptive relation profile across all five skills and two repetitions per paper. Counts retain the original compound propositions and are not independent paper samples.
Paper
Skill
Rep.
Correct
Relations 1–6
Central error
Unsupported
Rank
RAG
S0
1
5/6
CACCCC
no
0
5
RAG
S0
2
5/6
CCCCCA
uncertain
0
8
RAG
S1
1
6/6
CCCCCC
no
0
2
RAG
S1
2
5/6
CCCCOC
yes
0
10
RAG
S2
1
5/6
CCACCC
no
0
4
RAG
S2
2
4/6
CACCCA
yes
0
9
Appendix
Table 6: Every transfer unit. C: correct; A: ambiguous; O: omitted. Relation positions correspond to Appendix B . Rank is local to each paper, with 1 preferred. Unknown central errors are not relabeled as absent.
Unit
Relation
Human label
Prompt diagnosis
rag-S0-r1
rag-02
ambiguous
E: explicit requirement
rag-S0-r2
rag-06
ambiguous
E: explicit requirement
rag-S1-r2
rag-05
omitted
T: task-scope mismatch
rag-S2-r1
rag-03
ambiguous
E: explicit requirement
rag-S2-r2
rag-02
ambiguous
E: explicit requirement
rag-S2-r2
rag-06
ambiguous
E: explicit requirement
Appendix
Table 7: Complete post-hoc trace inventory. Human labels are unchanged. Prompt categories are assistant-authored diagnostic judgments, not independently human-verified classifications.
Figure 3: Two independent first outputs from the same frozen skill. The lower output chooses a Sequence-only scope permitted by the generator instruction, which differs from the evaluation scope; the author separately identifies a decoding inconsistency. The upper output meets all required relations while retaining a local gradient-direction issue. These post-assessment examples illustrate the difference between relation coverage and overall preference. Original pixels are retained.
Figure 4: RAP under S2 , repetition 2. Five required relations are correct, yet the author records a central control-flow error: the expansion-to-next-selection route can bypass simulation and backup. This post-assessment example motivates keeping extra-mechanism errors separate from checklist coverage. Its unsupported-relation count is zero.
Scientific diagrams are essential for communicating complex methodologies in academic papers. A natural way for researchers to specify such diagrams is through rough sketches, where text labels, connectors, and spatial arrangements express early semantic and topological intentions. However, sketches are usually incomplete, making them insufficient for directly producing publication-quality diagrams. Existing sketch-based generation methods mainly reconstruct the sketch itself, while recent text-driven diagram generation frameworks rely on textual semantics and do not fully exploit the topological structure contained in sketches. In this paper, we introduce DiagramRAG, a lightweight retrieval-augmented framework for sketch-based scientific diagram completion. Given a user sketch, DiagramRAG retrieves reference diagrams that are both semantically relevant to the sketch content and topologically compatible with its structure, and uses them to guide downstream diagram generation. To enable efficient structure-aware retrieval, we represent diagrams as knowledge graphs, synthesize sketch variants at different simplification levels, and train an embedding model to align sketches with compatible diagrams in a shared space. The retrieved references further provide content, topology, and visual priors for completing and rendering the final diagram. Experiments show that DiagramRAG achieves F1-scores of 0.848 and 0.802 on DiagramBank and FigureBench, respectively, and improves generation quality with the best VLM-as-a-Judge score of 7.170, while reducing inference latency to 35.48 seconds per sample. Our code and data are available at https://anonymous.4open.science/r/DiagramRAG-A262 and https://huggingface.co/datasets/anonymous-review-a262/DiagramSketch.
Diagram question answering (Diagram QA) requires reasoning-level attribution that links each question-answer pair to all visual regions needed to derive the answer, rather than only the region containing the final response. Creating such structured evidence across diagrams, charts, maps, circuits, and infographics is time-consuming, and existing annotation tools tightly couple their interfaces to dataset-specific formats. We present DIAGRAMS, a lightweight, schema-driven review framework that decouples interface logic from dataset-specific JSON structures through an internal meta-schema and dataset adapters. Given an image and QA pair with optional candidate regions, the system performs QA-conditioned evidence selection and proposes the regions required for reasoning. When QA pairs or candidate regions are missing, it generates them and supports human verification and refinement. Across six Diagram QA datasets, model-suggested evidence achieves 85.39% precision and 75.30% recall against reviewer-final selections (micro-averaged). These results indicate that the review-first framework reduces manual region creation while maintaining high agreement with final reasoning-level attributions. We release a public demo and installable package to support dataset auditing, grounded supervision creation, and grounded evaluation.
Anirudh Iyengar Kaniyar Narayana Iyengar, Tampu Ravi Kumar, Manan Suri +4
Arizona State University · University of Maryland · IIITDM +1
Recent efforts have aimed to automate scientific diagram generation from paper content (Lin et al., 2026; Zhu et al., 2026a). However, fully satisfying an author's visual and communicative preferences in a single turn is challenging: in our formative user study (N = 14), all participants requested further revisions after viewing an initial draft, and 86% of them rated the refined diagrams as more satisfactory. Despite the clear demand, the multi-turn workflow remains largely underexplored. To bridge this gap, we present MTPaperBananaBench, a benchmark for multi-turn diagram generation containing 292 images annotated with 3,518 user requirements. To reduce expensive human studies and enable scalable benchmarking, we construct a user simulator that, at each turn, identifies unsatisfied requirements and converts k of them into natural language feedback. Evaluating both requirement satisfaction and overall diagram quality reveals two key failure modes shared across baseline multiturn systems: (1) quality drift, where diagram quality progressively declines over turns, and (2) forgetting, where previously implemented features are lost in subsequent turns. To address these issues, we introduce PaperBanana-Interact, a multi-agent system that refines diagrams via an internal critique-and-refine loop. PaperBanana-Interact consistently improves rather than degrades diagram quality across turns, outperforming baselines by 11.9-18.6 points in quality score and reducing forgetting by 3.7-6.2 points.
Xueqing Wu, Ashwin Balasubramanian, Bingxuan Li +7
University of California, Los Angeles · Google · University of Illinois Urbana-Champaign +1