When Can Text Replace Vision? Structural Bottlenecks in Diagram Reasoning
Organizations: Tulane University · Virginia Tech
Abstract
Can structured text replace vision for diagram reasoning? A wrong answer after textualization can arise because the representation omits information the question needs, or because the solver fails to use information that is present. We introduce a diagnostic protocol to distinguish these explanations. Using the same solver model and generation settings, we compare three input conditions: the original image, question-blind structure extracted by a vision-language model, or gold structure derived from the diagram source. Validity-triggered recovery tests truncation and schema failure, question-relevant fidelity measures preservation of answer-critical structure, and matched edge interventions test the effect of error location. On a reserved holdout of 240 public FlowGen diagrams, evaluated under a frozen protocol, gold structure reaches 87% accuracy while direct vision and learned text both remain below 30%. The aggregate comparison includes source-derived relation labels that may not be printed in the image and uses different learned and gold graph encodings, so it does not isolate extraction error alone. Retrying only invalid extractions makes nearly every public representation schema-valid yet leaves accuracy essentially unchanged. The public learned-text deficit relative to gold more than doubles with structural difficulty. Question-relevant topology predicts correctness better than whole-graph topology. In an exposed intervention study, a single answer-relevant edge edit reduces the primary solver's original-answer accuracy to near zero, while matched irrelevant edits largely preserve it. Supplied structure requires fewer solving tokens than vision, but learned acquisition removes this advantage at single use. These comparisons motivate evaluating acquired text by the answer-relevant evidence it preserves and by the solver's ability to use that representation.
Figures & tables
| Input / solver | Valid / 240 | Correct | Missing | QA (%) |
|---|---|---|---|---|
| Direct vision | – | 196 | 11 | 27.2–28.8 |
| Fixed learned | 196 | 164 | 0 | 22.8 |
| Recovered learned | 238 | 169 | 0 | 23.5 |
| Gold structure | – | 627 | 0 | 87.1 |
| Fixed learned 27B | 196 | 157 | 0 | 21.8 |
| Extraction | Fidelity metric | Whole | Relevant | Paired CI | |||
|---|---|---|---|---|---|---|---|
| Fixed (588 Q) | Topology exact | 0.1406 | 0.1147 | 0.0259 [ | 0.0479, | 0.0047] | † |
| Edge F1 | 0.1305 | 0.1240 | 0.0065 [ | 0.0162, | 0.0032] | ||
| Recovered (714 Q) | Topology exact | 0.1242 | 0.1058 | 0.0184 [ | 0.0346, | 0.0029] | |
| Edge F1 | 0.1165 | 0.1134 | 0.0031 [ | 0.0115, | 0.0054] | ||
| 4B | 27B | 122B | |||||||
|---|---|---|---|---|---|---|---|---|---|
| Edit | Charts | Q | Rel. | Irrel. | Rel. | Irrel. | Rel. | Irrel. | |
| Reverse | 1 | 60 | 172 | 5.2 | 78.5 | 4.7 | 93.6 | 3.5 | 92.4 |
| Reverse | 2 | 58 | 122 | 3.3 | 68.9 | 1.6 | 92.6 | 0.0 | 91.0 |
| Drop | 1 | 50 | 148 | 0.0 | 79.1 | 0.0 | 93.9 | 0.0 | 93.9 |
| Drop | 2 | 49 | 100 | 0.0 | 78.0 | 0.0 | 93.0 | 0.0 | 94.0 |
| Redirect | 1 | 50 | 148 | 0.7 | 74.3 | 0.0 | 92.6 | 0.0 | 93.2 |
| Input | QA (%) | Tokens/Q, | Tokens/Q, | Comparable QA? |
|---|---|---|---|---|
| Direct vision | 27.2–28.8 | 6,423 | 6,423 | – |
| Fixed learned | 22.8 | 7,551 | 3,201 | No |
| Recovered learned | 23.5 | 10,413 | 4,397 | No |
| Gold structure | 87.1 | 1,169 | 1,169 | Yes |
Appendix figures & tables21 assets
Supplementary material from the paper’s appendix.
Appendix
| Cohort | Charts | Questions | Role |
|---|---|---|---|
| Public FlowGen, reserved | 240 | 720 | Confirmatory core |
| Public FlowGen, exposed | 240 | 720 | Recovery / mechanism controls |
| Controlled OLD271 | 271 | 271 | Mixed controlled confirmation |
| Controlled NEW270 | 270 | 270 | Harder controlled extension |
| QZhou, exposed | 200 | 600 | Secondary counterpoint |
| Cohort / comparison | Solver | Extractor | Solve cap | Extraction cap / input |
| Public holdout, primary | 122B-A10B | 122B-A10B | 8,192 | Fixed 1,024 |
| Public holdout, recovery | 122B-A10B | 122B-A10B | 8,192 | 1,024 2,048 8,192 |
| Public holdout, compatibility | 27B | 122B-A10B | 8,192 | Same fixed text |
| Exposed public recovery | 122B-A10B | 122B-A10B | 8,192 | Fixed / recovered |
| Controlled diagnostic | 122B-A10B | 122B-A10B | 8,192 | Fixed / recovered |
| Exposed perturbations | 4B / 27B / 122B | – | 512 | Edited source structure |
| QA (%) | Fixed Gold (pp) | ||||||
| Bin | Charts | Fixed | Recovered | Gold | Gap | 95% CI | |
| Node count | |||||||
| 10 | 49 | 55.8 | 55.8 | 94.6 | 38.8 [ | 49.7, | 27.2] |
| 11–20 | 137 | 20.0 | 20.4 | 85.2 | 65.2 [ | 71.0, | 59.1] |
| 21–40 | 54 | 0.0 | 1.9 | 85.2 | 85.2 [ | 90.1, | 80.2] |
| 40 | 0 | – | – | – | – | – | |
| Outcome / contrast | Axis | Estimate | CI | |
|---|---|---|---|---|
| Primary tests (multiplicity-adjusted 98.33% CI) | ||||
| Fixed Gold QA | Node count | 23.1 [ | 30.7, | 15.6] |
| Fixed Gold QA | Branch depth | 11.2 [ | 14.8, | 7.6] |
| Relevant whole Brier | – | 0.0259 [ | 0.0479, | 0.0047] |
| Secondary difficulty slopes (95% CI) | ||||
| Direct QA | Node count | 25.0 [ | 29.3, | 20.7] |
| Fixed learned | ||||||
|---|---|---|---|---|---|---|
| Domain | Charts | Direct | 1,024 cap | 2,048 cap | Recovered | Gold |
| Circuits | 135 | 99.3 | 85.2 | 89.6 | 99.3 | 94.1 |
| Flowcharts | 406 | 99.0 | 27.3 | 75.1 | 99.5 | 99.8 |
| Calls per question | Seconds per question | |||
|---|---|---|---|---|
| Input / solver | ||||
| Direct vision | 1.00 | 1.00 | 12.63 | 12.63 |
| Fixed learned | 1.82 | 1.15 | 19.35 | 10.71 |
| Recovered learned | 2.20 | 1.39 | 25.86 | 13.86 |
| Gold structure | 1.00 | 1.00 | 7.12 | 7.12 |
| Fixed learned 27B | 1.82 | 1.15 | 52.80 | 44.16 |
| Input / solver | Accepted | Cap hits | Hit rate (%) | Invalid skips | Missing |
|---|---|---|---|---|---|
| Direct vision | 709 | 35 | 4.9 | 0 | 11 |
| Fixed learned | 588 | 11 | 1.9 | 132 | 0 |
| Recovered learned | 714 | 16 | 2.2 | 6 | 0 |
| Gold structure | 720 | 8 | 1.1 | 0 | 0 |
| Fixed learned 27B | 588 | 19 | 3.2 | 132 | 0 |
| Adjudicated outcome | Items |
|---|---|
| QZhou source–image consistency (30 charts) | |
| Match | 28 |
| Cosmetic difference only | 1 |
| Material mismatch | 0 |
| Unresolved | 1 |
| QZhou answer relevance (90 questions) | |