SIVIA-RSI: Source-Grounded Adaptation of Diagramming Skills
Authors: Feng Yuan, Yifan Gao, Haoyue Li, Xin Gao
Organizations: School of Biomedical Engineering (Suzhou), Division of Life Science and Medicine, University of Science and Technology of China, Hefei, China · Suzhou Institute of Biomedical Engineering and Technology, Chinese Academy of Sciences, Suzhou, China · Shanghai Innovation Institute, Shanghai, China
Scientific method diagrams express computations through entities, dependencies, and conditional routes. Although generated figures can be improved through repeated editing, it is less clear whether experience from one paper improves the first figure of another. We present \sys, a framework for source-grounded adaptation of reusable diagramming skills, and study transfer through a complete-candidate evaluation. The framework links critiques to source passages, proposes bounded edits to a persistent skill library, and separates candidate competition from skill acceptance. We evaluate the original skill and four learned candidates on two NLP and language-agent papers, with two fresh generations per condition. The strongest candidate attains 87.50% required-relation accuracy compared with 83.33% for the original, while candidate behavior differs across papers. Local improvements on a separate development paper and automatic selector preferences do not establish consistent transfer. Tracing all 22 non-correct relation judgments to their production prompts reveals both incomplete conditional specifications and ambiguities despite explicit instructions. All ten planning diagrams leave an already-terminal selected leaf's route unclear; none of their prompts explicitly binds that route. Our findings show why evaluating reusable diagram skills requires source-grounded relation assessment, complete candidate coverage, and inspection of both prompts and images. We provide all twenty transfer outputs, skill snapshots, assessment records, and executable analyses.
Figures & tables
Figure 1: Source-grounded skill adaptation and complete-candidate auditing. (a) Discovery evidence informs bounded edits to a reusable skill and, optionally, its updater policy. Separate papers support candidate competition and parent acceptance. The two learning runs produce four proposals, all retained regardless of selection. (b) The original skill and four proposals are frozen for generation on RAG and RAP: five skills, two papers, and two fresh repetitions yield twenty first outputs. Source-verified relations and additional mechanism errors are assessed after generation, without target-specific repair. The document and diagram miniatures are schematic; measured outputs appear in Figures 3 and 4 .
Skill
Tokens
Candidate origin and emphasis
S0
24,240
Original nine-file skill.
S1
24,348
Fixed run, first proposal; state ownership and consumers.
S2
24,330
Fixed run, second proposal; state interfaces and layout numbers.
Table 1: All frozen skill identities. The automatic selector advanced S1 and S3 ; the transfer study retains all four proposals. Token counts use a shared proxy serialization.
Required relations correct
Central error: yes / uncertain
Skill
RAG
RAP
Macro (%)
Δ (pp)
RAG (2 images)
RAP (2 images)
Extra unsupported
Pass gate
S0
10/12
10/12
83.33
+0.00
0 / 1
0 / 0
0
—
S1
11/12
10/12
87.50
+4.17
1 / 0
0 / 0
0
No
S2
9/12
10/12
79.17
-4.17
1 / 0
1 / 0
0
No
S3
11/12
9/12
83.33
+0.00
0 / 1
0 / 0
0
No
S4
11/12
7/12
75.00
-8.33
0 / 1
0 / 2
1
No
Table 2: Complete frozen transfer matrix. Macro averages repetitions within each paper and then the two paper families equally. Extra unsupported counts cover four images per skill. Uncertain central-error labels remain unknown. All candidates fail the primary 10-percentage-point continuation gate, independently of additional-error constraints. This is a descriptive development comparison with two paper families.
Figure 2: First-output relation accuracy for every frozen skill. Bars show the mean of two independent design-and-render repetitions; circles and diamonds show the individual images. Dashed lines mark the original skill’s paper-level mean. The same candidate can improve RAG while weakening RAP. The plot summarizes observed outputs, with no confidence intervals or independent-paper interpretation of the repetitions.
Reflexion pair
Relations
Fidelity
Author choice
Fixed run
2/4→4/4
2→4
Updated
Evolving run
3/4→4/4
3→4
Updated
Table 3: Local parent-to-candidate comparisons on Reflexion under the initial rubric. Fidelity uses a 1–5 scale. Both updated skills were proposed by the original updater; the pairs concern one paper.
Prompt-side diagnosis
RAG
RAP
Total
Explicit requirement
6
2
8
Partial / unbound requirement
1
11
12
Absent critical requirement
0
1
1
Generation–evaluation scope mismatch
1
0
1
Table 4: Assistant-authored post-hoc trace of all 22 non-correct human relation judgments. Counts partition these observed items; they are neither additional human labels nor causal error probabilities.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Paper
Required relation (short name)
Correct / 10
RAG
Query to retrieval
10/10
Input and document operands
7/10
Weighted marginalization
9/10
Sequence formulation
10/10
Token formulation
8/10
Training ownership
8/10
Appendix
Table 5: Descriptive relation profile across all five skills and two repetitions per paper. Counts retain the original compound propositions and are not independent paper samples.
Paper
Skill
Rep.
Correct
Relations 1–6
Central error
Unsupported
Rank
RAG
S0
1
5/6
CACCCC
no
0
5
RAG
S0
2
5/6
CCCCCA
uncertain
0
8
RAG
S1
1
6/6
CCCCCC
no
0
2
RAG
S1
2
5/6
CCCCOC
yes
0
10
RAG
S2
1
5/6
CCACCC
no
0
4
RAG
S2
2
4/6
CACCCA
yes
0
9
Appendix
Table 6: Every transfer unit. C: correct; A: ambiguous; O: omitted. Relation positions correspond to Appendix B . Rank is local to each paper, with 1 preferred. Unknown central errors are not relabeled as absent.
Unit
Relation
Human label
Prompt diagnosis
rag-S0-r1
rag-02
ambiguous
E: explicit requirement
rag-S0-r2
rag-06
ambiguous
E: explicit requirement
rag-S1-r2
rag-05
omitted
T: task-scope mismatch
rag-S2-r1
rag-03
ambiguous
E: explicit requirement
rag-S2-r2
rag-02
ambiguous
E: explicit requirement
rag-S2-r2
rag-06
ambiguous
E: explicit requirement
Appendix
Table 7: Complete post-hoc trace inventory. Human labels are unchanged. Prompt categories are assistant-authored diagnostic judgments, not independently human-verified classifications.
Figure 3: Two independent first outputs from the same frozen skill. The lower output chooses a Sequence-only scope permitted by the generator instruction, which differs from the evaluation scope; the author separately identifies a decoding inconsistency. The upper output meets all required relations while retaining a local gradient-direction issue. These post-assessment examples illustrate the difference between relation coverage and overall preference. Original pixels are retained.
Figure 4: RAP under S2 , repetition 2. Five required relations are correct, yet the author records a central control-flow error: the expansion-to-next-selection route can bypass simulation and backup. This post-assessment example motivates keeping extra-mechanism errors separate from checklist coverage. Its unsupported-relation count is zero.