Harness Compilation: Which Decisions Should a Small Vision-Language Model Keep?
Organizations: College of Computing and Data Science, Nanyang Technological University · School of Microelectronics, University of Science and Technology of China
Abstract
Small vision-language models may be able to read external evidence yet struggle to obtain it. We introduce Harness Compilation (HC), an offline procedure that adapts the division of work between a frozen small VLM and its external harness. A large teacher uses student execution traces to revise reusable content and control, while a separate validation set selects the deployed harness. Deployment requires neither weight updates nor teacher calls. Across seven visual question-answering settings with students of at most 9B parameters, HC improves scores over bare students by 9.9-23.9 points, averaged over three independent builds per setting. Interventions on five runtime decision types (invocation, selection, argument generation, evidence integration and abstention) show why this allocation matters: requesting evidence and generating open queries can be costly, whereas bounded choices and reading supplied text can remain useful student work. Fact cards benefit all ten evaluated students, but decision policies transfer unevenly. Recompilation for a new student model helps when the transferred interface no longer fits the student. With 100 practice items, HC exceeds answer-only LoRA on three tasks. Larger training budgets can match or surpass a fixed harness, while combining the two improves SlideVQA beyond either alone. These findings support allocating work from measured student behavior rather than uniformly removing decisions.
Figures & tables
| Setting | Compiled work | Student retains |
|---|---|---|
| InfoSeek | Route retrieval and prepare a card | Choose among names on ambiguous inputs, then answer |
| DocVQA | Supply OCR and an answer format | Read the page and supplied text |
| SlideVQA | Rank and select evidence slides | Read selected slides and answer |
| LiveVQA | Supply three article cards | Select and combine evidence while answering |
| BLINK | Compute a verdict from measurements | Answer when the rule defers |
| Setting | Bare | Student-routed | Human recipe | HC (3 builds) |
|---|---|---|---|---|
| InfoSeek | 18.0 | 29.1 | 33.8 | 37.7 |
| DocVQA | 72.7 | 80.4 | 82.4 | 85.0 |
| SlideVQA | 44.7 | 44.4 | 48.7 | 54.6 |
| LiveVQA | 14.7 | 36.9 | 25.5 | 38.6 |
| BLINK reflectance | 50.8 | 50.0 | 59.0 | 62.9 |
| BLINK correspondence | 54.9 | 64.9 | 64.2 | 65.7 |
| Setting | Initial | Selected |
|---|---|---|
| InfoSeek | 33.5 | 37.7 |
| DocVQA | 78.4 | 85.0 |
| SlideVQA | 54.4 | 54.6 |
| LiveVQA | 34.1 | 38.6 |
| BLINK reflectance | 61.7 | 62.9 |
| BLINK correspondence | 65.2 | 65.7 |
| Setting | Decide | Select | Generate | Integrate | Persist |
|---|---|---|---|---|---|
| InfoSeek | |||||
| DocVQA | n/a | ||||
| SlideVQA | n/a | ||||
| LiveVQA | n/a | n/a | n/a | ||
| BLINK reflectance | n/a | n/a | |||
| BLINK correspondence | n/a | n/a |
| Change to the selected interface | Evaluated harnesses | Score-change range |
|---|---|---|
| Restore stored cards’ original lines | A and B (6) | to |
| Remove question-text hints | A (3) | to |
| Default extractor with contract | A (3) / B (3) | to / to |
| Artifact extractor without contract | A (3) / B (3) | to / to |
| Setting | HC | Answer LoRA | Trajectory SFT |
|---|---|---|---|
| InfoSeek | 33.6 | 23.3 | 35.3 |
| DocVQA | 84.7 | 81.1 | n/a |
| SlideVQA | 54.7 | 47.3 | n/a |
| Setting | LoRA labels | Fixed HC | LoRA | LoRA + HC |
|---|---|---|---|---|
| SlideVQA | 1,300 | 58.1 | 53.2 | 63.7 |
| DocVQA | 999 | 85.4 | 84.7 | 86.4 |
| Setting | Test score | Tokens per item | ||
|---|---|---|---|---|
| CoT + vote | HC | CoT + vote | HC | |
| InfoSeek | 24.6 | 37.3 | 1,120 | 5.9 |
| DocVQA | 80.4 | 85.7 | 580 | 8.5 |
| SlideVQA | 51.1 | 58.1 | 970 | 5.8 |
| BLINK reflectance | 64.9 | 64.2 | 2,000 | 0.1 |
| BLINK correspondence | 68.5 | 67.4 | 2,510 | 0.9 |
Appendix figures & tables50 assets
Supplementary material from the paper’s appendix.
Appendix
| Setting | Student | Metric | Test items | Practice |
|---|---|---|---|---|
| InfoSeek | Qwen3-VL-4B | Relaxed exact match | 6,000 | 150 / 300 |
| DocVQA | Gemma-4-E4B | ANLS | 3,568 | 150 / 300 |
| SlideVQA | Qwen3-VL-4B | Exact match | 2,215 | 150 / 300 |
| LiveVQA | Qwen3-VL-4B | Judged accuracy | 2,000 | 150 / 300 |
| BLINK reflectance | Qwen3-VL-4B | Accuracy | 134 | 44 / 45 per fold |
| BLINK correspondence | Qwen3-VL-4B | Accuracy | 441 | 147 / 147 per fold |
| Setting | Human recipe | Score |
|---|---|---|
| InfoSeek | Select a candidate name, then read retrieved Wikipedia passages | 33.8 |
| DocVQA | Supply page OCR in typed fields | 82.4 |
| SlideVQA | Read the three slides selected by the text ranker | 48.7 |
| LiveVQA | Select one candidate headline and read its article card | 25.5 |
| BLINK reflectance | Fixed Retinex verdict rule | 59.0 |
| BLINK correspondence | Fixed LightGlue and DINO rule | 64.2 |
| Comparison | InfoSeek | DocVQA | Reference and practice budget |
|---|---|---|---|
| Formal results (Tables 2 – 3 ) | 37.7 | 85.0 | Interface v2, reflection A; three replicates, 450 items each |
| Decision interventions (Table 4 ) | 37.8 | 84.1 | Current runtime; one harness, 450 items (InfoSeek, DocVQA, LiveVQA); SlideVQA and BLINK use the formal replicate-1 artifacts. The BLINK integrate entries are the student-routed arm of Table 2 |
| Generated-token comparison (Table 8 ) and compute baselines (Table 41 ) | 37.3 | 85.7 | Formal replicate-1 harnesses re-executed with token logging; score and tokens from the same execution |
| One-shot construction (Table 22 ) | 39.7 | 84.9 | Legacy runs; 1,200 / 999 items, respectively |
| Fixed-harness comparison with tuning (Table 7 ) | – | 85.4 | Identical artifact hashes on both students: SlideVQA = the formal replicate-1 harness; DocVQA = a current-protocol build on the 999 adapter-training items |
| Current transfer source (Table 28 ) | 37.8 | – | Current runtime; 450 items |
| Setting | GPT-5.5 | Evaluation population and input |
|---|---|---|
| InfoSeek | 50.0 | 300-item practice subset |
| DocVQA | 93.3 | Full test set, 2,048-pixel images |
| SlideVQA | 64.7 | 150-question practice pool with the linker’s top-three slides |
| LiveVQA | 58.3 | 300-question practice pool with the same candidate passages |
| BLINK reflectance | 79.9 | Same items as the student and harness |
| BLINK correspondence | 78.7 | Same items as the student and harness |
| Setting | Items | What the comparison tests |
|---|---|---|
| Relative reflectance | 134 | Whether local photometric measurements can support a useful verdict despite incomplete illumination invariance. |
| Correspondence | 441 | Visual (172), semantic (139), and functional (130) correspondence share a matching interface but differ in what geometric signals can resolve. |
| Relative depth | 124 | Whether a depth estimator’s measurements are better used through code than interpreted by the student. |
| Measurement | Qwen3-VL 2B | Qwen3-VL 4B | Qwen3-VL 8B | Gemma-4 E4B | InternVL3.5 4B |
|---|---|---|---|---|---|
| Requesting resources and following rules | |||||
| Memory activation rate | 0.0 | 82.3 | 89.3 | 97.3 | 58.7 |
| Generated entity-name precision | 0.0 | 30.4 | 36.2 | 15.1 | 0.6 |
| Rule compliance, turn 1 | 36 | 100 | 98 | 100 | 22 |
| Rule compliance, turn 6 | 32 | 100 | 98 | 90 | 18 |
| Correct choice, 1-tool menu | 19 | 66 | 75 | 46 | 45 |
| Stage | Supplied information or required output |
|---|---|
| Runtime manual | Allowed observations, JSON fields and defaults, router modes, code-hook signatures, and output constraints. |
| Initial compilation | Task description, bare-student probe score and resource examples. No separate student profile on InfoSeek, LiveVQA and SlideVQA. The DocVQA builds receive the Gemma profile block (a deviation from the stated protocol, recorded in Appendix A.1 ). |
| Reflection A | Latest proposal JSON and its traced probe report, plus round history with scores and the retained incumbent. The incumbent’s JSON is not separately supplied as the edit target. |
| Probe evidence | Stage-level errors, 30 stratified InfoSeek failure traces, pick precision and abstention, with one additional trace query per round. Gold labels serve offline diagnosis only. |
| Teacher response | One complete harness JSON containing prompt, resource and decoding fields, plus executable router and parser source. |
| Selection | Strict improvement of the unrounded validation mean, with the bare-model floor. The selected artifact is frozen before test evaluation. |
| Setting | Router inputs | Keys beyond prompt , decode | External control | Student work |
|---|---|---|---|---|
| InfoSeek | Question text; CLIP top-3 entity names, similarities and cards from the fixed inventory | memory , discriminate_prompt , name_prompt , parse_py | when to show a card, when to invoke the bounded pick, which card lines to show | bounded choice among named candidates when routed to it (with the escape), then reading the card and answering |
| LiveVQA | Question text; CLIP top-3 news headlines, similarities and article cards | same keys; no store rebuild | retrieve and present the article evidence; card rewriting | read the passages and answer, with an explicit headline pick on paths that request one |
| DocVQA | Question text and noisy page OCR | tool ( ocr , max_chars ), parse_py | OCR provision and the answer contract | read the page and the supplied OCR, then answer |
| SlideVQA | MiniLM and CLIP slide rankings, per-slide OCR, deck size | pages ( , ranker, whole deck), tool , parse_py ; router may return a verdict | evidence-page ranking and selection | read and combine the selected slides, then answer |
| BLINK (3) | the bed’s pixel, matcher or depth measurements at the marked points (§ A.2 ) | tool.signals , answer_mode , parse_py | compute a verdict from the measurements and decide when to defer | answer when the rule defers; in the suggestion variants, read the verdict as advice |
| LIBERO-plus | instruction, nearest training sentence with similarity, detector scores of the instruction’s noun phrases, end-effector state, steps since last motion | tools (detector, check_every , restore), persist ( max_steps , stall) | instruction override and the exposed termination controls (abort, stop) | generate actions from the observations and the current instruction |
| Replicate | Protocol | Selected round | Validation | Test | Teacher $ |
|---|---|---|---|---|---|
| 1 | A | 1 | 41.0 | 37.3 | 1.19 |
| 1 | B | 3 | 40.7 | 36.1 | 0.54 |
| 2 | A | 1 | 43.3 | 39.1 | 0.83 |
| 2 | B | 0 | 40.0 | 36.9 | 0.57 |
| 3 | A | 1 | 40.0 | 36.6 | 0.91 |
| 3 | B | 2 | 37.0 | 34.6 | 0.71 |
| Replicate | Protocol | Genus | Hints | Card budget | Routing |
|---|---|---|---|---|---|
| 1 | A | Yes | Yes | 850 | Similarity thresholds + category abort |
| 1 | B | No | No | 850 | Similarity thresholds only |
| 2 | A | No | Yes | 850 | Similarity thresholds only |
| 2 | B | No | No | 850 | Similarity thresholds only |
| 3 | A | Yes | Yes | 560 | Always show the top three cards |
| 3 | B | Yes | No | 560 | Always show the top three cards |
| Replicate | Protocol | archived / R0 | R1 derived line | R2 hints | R3 both | R4 original cards | R5 similarity routing |
|---|---|---|---|---|---|---|---|
| 1 | A | 37.30 / 37.30 | (0/0) | (37/171 ∗ ) | (37/180 ∗ ) | (65/22 ∗ ) | (66/25 ∗ ) |
| 1 | B | 36.12 / 36.13 | n/a | n/a | n/a | (77/32 ∗ ) | n/a |
| 2 | A | 39.13 / 39.12 | n/a | (6/14) | n/a | (21/5 ∗ ) | n/a |
| 2 | B | 36.93 / 36.93 | n/a | n/a | n/a | (28/6 ∗ ) | n/a |
| 3 | A | 36.62 / 36.60 | (11/17) | (12/83 ∗ ) | (11/104 ∗ ) | (315/143 ∗ ) | n/a |
| 3 | B | 34.63 / 34.72 | (2/18 ∗ ) | n/a | n/a | (299/199 ∗ ) | n/a |
| Replicate | Protocol | O11 selected / hook | O10 selected / default | O01 removed / hook | O00 removed / default |
|---|---|---|---|---|---|
| 1 | A | 37.30 | 37.30 ( ) | 35.52 ( ) | 35.52 ( ) |
| 2 | A | 39.13 | 39.12 ( ) | 34.38 ( ) | 35.15 ( ) |
| 3 | A | 36.62 | 36.62 ( ) | 30.55 ( ) | 32.03 ( ) |
| 1 | B | 36.12 | 36.12 ( ) | 34.05 ( ) | 32.88 ( ) |
| 2 | B | 36.93 | 36.87 ( ) | 34.37 ( ) | 33.90 ( ) |
| 3 | B | 34.63 | 34.63 ( ) | 32.92 ( ) | 33.38 ( ) |
| Configuration | InfoSeek | DocVQA |
|---|---|---|
| DSL default: name card / plain OCR | 23.6 | 80.1 |
| Store-only / OCR-only control | 33.5 | 80.1 |
| Sensitivity profile (3 / 2 builds) | 33.7 to 34.0 | 84.9 / 85.7 |
| Profile withheld (3 / 1 builds) | 35.6 / 39.7 / 40.3 | 86.0 |
| Measurements-only profile | 34.0 | n/a |
| Decision-menu profile | 38.8 | n/a |
| Setting | One-shot | HC | One-shot minus HC |
|---|---|---|---|
| BLINK reflectance | 65.7 | 64.9 | [ , ], n.s. |
| BLINK correspondence | 62.8 | 67.6 | [ , ], |
| BLINK depth | 95.2 | 95.2 | Identical verdicts |
| DocVQA | 85.5 | 84.9 | [ , ], n.s. |
| InfoSeek | 25.0 | 39.7 | [ , ] |
| LiveVQA | 30.6 | 37.1 | , |
| Setting | Initial | Selected | Gain per build | Selected, protocol B |
|---|---|---|---|---|
| InfoSeek | 33.5 | 37.7 | , , | 35.9 |
| DocVQA | 78.4 | 85.0 | , , | 79.8 |
| SlideVQA | 54.4 | 54.6 | , , | 55.7 |
| LiveVQA | 34.1 | 38.6 | , , | 37.3 |
| BLINK reflectance | 61.7 | 62.9 | , , | 61.2 |
| BLINK correspondence | 65.2 | 65.7 | , , | 65.4 |
| Setting | Replicate | selected (A) | selected (B) | round, A | round, B (no-ops) | $ A | $ B | |
|---|---|---|---|---|---|---|---|---|
| InfoSeek | 1 | 35.8 | 37.3 | 36.1 | 1 | 3 (0) | 1.19 | 0.54 |
| InfoSeek | 2 | 36.9 | 39.1 | 36.9 | 1 | 0 (0) | 0.83 | 0.57 |
| InfoSeek | 3 | 27.8 | 36.6 | 34.6 | 1 | 2 (0) | 0.91 | 0.71 |
| DocVQA | 1 | 74.3 | 85.8 | 79.1 | 1 | 3 (0) | 0.60 | 0.49 |
| DocVQA | 2 | 83.1 | 84.6 | 83.9 | 1 | 1 (0) | 0.66 | 0.54 |
| DocVQA | 3 | 77.7 | 84.6 | 76.4 | 3 | 3 (0) | 0.62 | 0.41 |
| Component | A: last proposal | B: accepted incumbent |
|---|---|---|
| Edit target | Latest proposal, accepted or rejected | Current incumbent, supplied as full JSON with its hash |
| Probe report | Latest proposal’s report; paired traces against the preceding proposal | Incumbent’s report; latest rejection shown separately with a compact paired report against the incumbent |
| History | Scores and accepted incumbent identified in the round history | Parent/proposal/incumbent hashes, changed components, scores and gate decision |
| Edit instruction | Return a revised complete harness | One coherent mechanism change, evidence and denominator, and a possible regression; unchanged incumbent is allowed |
| Stopping | Up to three reflection rounds | Up to three rounds; stop after two consecutive no-ops |
| Selection | Strict validation improvement, with bare floor | Same gate |
| Setting | Score change | 95% interval | Paired -value |
|---|---|---|---|
| InfoSeek | – | ||
| DocVQA | – | ||
| SlideVQA | |||
| LiveVQA | – | ||
| BLINK reflectance | – | ||
| BLINK correspondence | – |
| Decision | Intervention | Interpretation |
|---|---|---|
| Decide | Supply the resource only after a yes/no request | Tests activation under that request contract |
| Select | Show all candidates without the explicit pick or ranker | Changes the selection interface, not only the decision owner |
| Generate | Use a student-generated entity name for lookup | Tests open argument generation on InfoSeek |
| Integrate | Use raw content instead of rewritten content. On BLINK, use raw signals instead of the processed rule output | The intervention differs between content and perception tasks |
| Persist | Remove the answer contract and parser together (a bundled output interface). On BLINK, show the verdict as a suggestion | Tests the output contract and abstention control, or verdict emission; separate formal contract/extractor controls are in App. A.5 |
| Student | Bare | Transplant | Own build | Difference | Paired -value |
| Qwen3-VL-2B | 13.2 | 30.1 | 32.1 | ||
| Qwen3-VL-8B | 20.3 | 39.7 | 38.0 | ||
| Gemma-4-E4B | 6.7 | 30.8 | 33.0 | ||
| InternVL3.5-4B | 9.6 | 30.9 | 26.2 | ||
| Qwen3.5-0.8B | 6.8 | 16.1 | 27.3 | ||
| Qwen3.5-2B | 9.0 | 30.7 | 26.0 |
| Student | Pick / net | “0” | Routed gain | ( ) | bare |
|---|---|---|---|---|---|
| Qwen3-VL-2B | 57.0 / 55.3 | 1% | ( ) | ||
| Qwen3-VL-4B | 66.0 / 58.0 | 17% | ( ) | ||
| Qwen3-VL-8B | 61.2 / 55.7 | 12% | ( ) | ||
| Gemma-4-E4B | 49.3 / 47.0 | 8% | ( ) | ||
| InternVL3.5-4B | 50.8 / 48.3 | 21% | ( ) | ||
| Qwen3.5-0.8B | 39.8 / 39.7 | 0% | ( ) |
| student | “0” | linked-correct on fired | acc. on fired | acc. not fired | |
|---|---|---|---|---|---|
| Qwen3-VL-2B | 1% | 32.5 39.6 | 20.8 25.1 ( ) | 44.7 43.8 ( ) | |
| Qwen3-VL-4B | 17% | 32.5 42.8 | 23.4 31.9 ( ) | 48.7 51.7 ( ) | |
| Qwen3-VL-8B | 12% | 32.5 43.7 | 24.2 31.8 ( ) | 49.1 52.9 ( ) | |
| Gemma-4-E4B | 8% | 32.5 33.0 | 20.0 22.4 ( ) | 42.8 47.0 ( ) | |
| InternVL3.5-4B | 21% | 32.5 31.4 | 20.8 22.1 ( ) | 42.7 47.2 ( ) | |
| Qwen3.5-0.8B | 0% | 32.5 25.1 | 12.9 14.3 ( ) | 27.1 33.0 ( ) |
| Student | DocVQA bare | OCR | Reflectance | Correspondence | Depth |
|---|---|---|---|---|---|
| Qwen3-VL-2B | 91.4 | 29.9 61.9 | 51.3 58.1 | 74.2 93.5 | |
| Qwen3-VL-4B | 94.9 | 50.8 64.9 | 54.9 67.6 | 82.3 95.2 | |
| Qwen3-VL-8B | 95.3 | – | – | – | |
| Gemma-4-E4B | 72.7 | 46.3 64.2 | 39.5 62.6 | 64.5 92.7 | |
| InternVL3.5-4B | – | – | – | – | – |
| Qwen3.5-0.8B | 84.9 | 28.4 62.7 | 0.0 (17.9) 57.4 | 67.7 93.5 |
| Student | Bare | Transplant | Initial | Selected | vs bare | Teacher $ |
|---|---|---|---|---|---|---|
| Qwen3.5-2B | 91.5 | 82.7 | 91.5 | 91.6 | – | 0.81 |
| Qwen3.5-0.8B | 84.9 | 56.1 | 85.7 | 85.8 | – | 0.46 |
| MiniCPM-V-4.6 | 86.1 | 71.0 | 13.7 | 86.3 (bare retained) | – | 1.04 |
| Average | 87.5 | 69.9 | 63.6 | 87.9 | – | 0.77 |
| Evidence / setting | Observation | Supported interpretation |
|---|---|---|
| BLINK diagnostic builds | HC trails the teacher by 14.9 points on reflectance, 20.1 on semantic and 12.3 on functional correspondence; depth is 5.6 higher (n.s.). | Available measurements support some decisions but do not close every task gap (Fig. 8 ). |
| Current compute control | CoT with voting reaches HC on reflectance and correspondence, with much higher generated-token counts (Table 8 ). | Some gains can also be obtained by increasing the student’s reasoning compute. |
| Exploratory MathVista / IQ Test | MathVista: best tested tool adds 4.2 against a 19.2 teacher gap. IQ Test: teacher 51.3 versus 4B 24.0 on 150 items; no suitable carrier identified. | A teacher advantage alone does not establish a usable external channel. Unrecovered gap is not an isolated measure of reasoning. |
| Earlier InfoSeek lookup | Student-routed lookup: versus for mechanical linking; bounded candidate selection: . | Open argument generation and bounded selection impose different demands. |
| Earlier code-tool diagnostic | MathVerse: self-routed interpreter adds 10.2 for the 4B but loses 3.6 for InternVL; execution depends on the generated code. | A computation tool still requires valid, grounded arguments from its executor. |
| Earlier advice-block diagnostic | On WeMath, a failure-derived standing block costs the 4B 12.8 points; its benefit over generic advice is not significant. | Added guidance can interfere with reasoning under a particular model and prompt contract. |
| Harness | Store | Routed | Routed gain (linking) | Rest gain | Total |
|---|---|---|---|---|---|
| Store-only / | Ref. / rewritten | – | – | – | 33.5 / 34.3 |
| Tuned human pick | Ref. | 100% | – | – | 37.4 |
| HC build 1 (1,200 items) | Rewritten | 60% | (32.5 42.8) | 39.7 | |
| HC build 2 | Ref. | 68% | (36.7 46.2) | 40.3 | |
| HC build 3 | Ref. | 61% | (34.0 44.6) | 35.6 | |
| HC, decision-menu profile | Ref. | 76% | (38.8 48.4) | 38.8 |
| Configuration | Practice items | ANLS |
|---|---|---|
| Bare | None | 72.7 |
| Human: plain OCR, 4,000 characters | Tuned on 500 | 80.1 |
| Human: typed OCR field | Tuned on 500 | 82.4 |
| Unoptimized DSPy typed program | None | 86.2 |
| One-shot teacher harness | None | 85.5 |
| HC, 50 probe / 50 validation | 100 | 84.4 |
| Setting | Harness | Verdict share | Verdict / bare | Fallback / bare | Total |
|---|---|---|---|---|---|
| Reflectance | HC seed 0 | 92% | 65.9 / 50.4 | 54.5 / 54.5 | 64.9 |
| HC seed 1 | 100% | 61.9 / 50.7 | – | 61.9 | |
| Human | 84% | 60.2 / 50.4 | 52.4 / 52.4 | 59.0 | |
| Correspondence | HC seed 0 | 55% | 83.1 / 65.0 | 48.5 / 42.4 | 67.6 |
| HC seed 1 | 65% | 76.0 / 59.9 | 47.4 / 45.5 | 66.0 | |
| Human | 93% | 65.0 / 55.0 | 53.3 / 53.3 | 64.2 |
| Student | Bare | Human | Top-1 | Gold | ||
|---|---|---|---|---|---|---|
| Qwen3-VL-4B (interface v2) | 14.7 | 25.5 | 37.5 | 38.1 | 27.4 | 51.6 |
| Qwen3.5-4B | 14.1 | 16.6 | 35.2 | 35.2 | 24.3 | 44.5 |
| Qwen3.5-9B | 16.9 | 22.5 | 37.9 | 36.1 | 29.9 | 55.1 |
| MiniCPM-V-4.6 | 7.1 | 13.8 | 22.5 | 30.1 | 5.0 | 8.3 |
| Average | 13.2 | 19.6 | 33.3 | 34.9 | 21.7 | 39.9 |
| Student | Bare | Human | Initial HC | Selected HC | Transplant |
|---|---|---|---|---|---|
| Qwen3-VL-4B | 44.7 | 48.7 | 41.9 | 56.7 | n/a |
| Qwen3.5-4B | 52.6 | 38.0 | 48.2 | 61.4 | 50.7 |
| Qwen3.5-9B | 51.4 | 47.4 | 56.4 | 62.1 | 63.3 |
| MiniCPM-V-4.6 | 34.8 | 33.0 | 28.4 | 38.8 | 29.4 |
| Gemma-4-E4B | 43.2 | 35.8 | 38.5 | 48.4 | 43.0 |
| Average (4 recipients) | 45.5 | 38.6 | 42.9 | 52.7 | 46.6 |
| Student | Bare, 224 pixels | Top-three images | Gold slides |
|---|---|---|---|
| Qwen3-VL-4B | 32.6 | 48.7 | 58.6 |
| Qwen3.5-4B | 36.7 | 33.6 | 37.0 a |
| Qwen3.5-9B | n/a | 31.2 | 55.1 a |
| MiniCPM-V-4.6 | n/a | 20.1 | 26.6 a |
| Gemma-4-E4B | 34.6 | 39.0 | 47.6 |
| bed | setting | prompt text (verbatim) | test |
|---|---|---|---|
| InfoSeek, recognise then read | stage 1: the three CLIP candidate names as options, store of name cards; stage 2: the two Wikipedia passages of the picked entity most similar to the question (MiniLM), fallback to the CLIP top-1 entity on abstention; 64 decode tokens | stage 1: “Which of these is the entity shown in the image?\n{options}\n0. None of them\nAnswer with the number only.”; stage 2: “Retrieved from Wikipedia (about the entity the retriever believes is shown; may be the wrong entity):” + passages + “Answer the question with a single word, number, or short phrase.” | 33.8 |
| DocVQA, typed OCR field | page image + question + 4,000 characters of Tesseract OCR in the DSPy typed-fields serialisation; generic field extractor; 96 decode tokens | context frame “Context (OCR text of the page; may contain errors):\n{tool_text}”; objective “Answer the question about the document image with a single word, number, or short phrase, copied from the page.” | 82.4 |
| BLINK reflectance | Retinex log-reflectance ( ) at A and B; “same”, else the lower one is darker; verdict answers the item, else the student | tool text shown to the student on deferred items: “Local reflectance estimate (brightness relative to surroundings): A = %.2f, B = %.2f” | 59.0 |
| BLINK correspondence | LightGlue projects REF into image 2; if the nearest candidate is within 8% of the diagonal, pick it; else the candidate with the highest DINOv2 cosine to REF; defer if no signal | “Keypoint matcher projects REF next to point %s” / “Feature similarity to REF: A 0.xx, B 0.xx, …” | 64.2 |
| BLINK depth | Depth-Anything-V2-Small at radius 4; larger = closer; defer if a marker is missing | “Depth estimator: relative closeness at A = %.2f, at B = %.2f (larger = closer)” | 94.4 |
| LiveVQA, pick a headline, read it | the three CLIP-linked article titles as options, the picked article’s card ( ) | the InfoSeek stage-1 prompt over article titles; card frame “Retrieved from your memory store (may or may not describe the right entity):\n{cards}” | 25.5 |
| bed student | vote | CoT+vote | TTAug ( vs own ctrl) | |
|---|---|---|---|---|
| DocVQA Gemma-4-E4B | 72.5 | 80.4 | 74.8 ( ) | 85.7 |
| SlideVQA 4B | 44.6 | 51.1 | 46.8 ( , ) | 58.1 |
| InfoSeek 4B | 17.6 | 24.6 | – | 37.3 |
| BLINK reflectance 4B | 50.8 | 64.9 | 51.5 ( ) | 64.2 |
| BLINK correspondence 4B | 54.0 | 68.5 | 56.9 ( ) | 67.4 |
| BLINK depth 4B | 82.3 | 85.5 | 86.3 ( , n.s.) | 96.0 |
| arm | InfoSeek 4B | DocVQA Gemma |
|---|---|---|
| RAG-wiki | 30.9 | – |
| RAG-wiki (gold entity) | 48.8 | – |
| GEPA-prompt | 17.9 19.3 | 82.6 82.9 |
| GEPA-content (fixed store / OCR) | 32.9 33.3 | 86.2 86.2 |
| GEPA-structured (two modules) | 34.9 36.0 | – |
| search | seed 0 | seed 1 | seed 2 | mean (SD) |
|---|---|---|---|---|
| HC loop | 34.1 | 40.3 | 35.6 | 36.6 (3.3) |
| GEPA | 33.5 | 37.8 | 36.6 | 35.9 (2.2) |
| HC GEPA, paired | (.24) | (.037) | (matched) |
| Setting | Input | Plain | Typed |
|---|---|---|---|
| DocVQA, Gemma-4-E4B | Bare | 72.7 | 81.7 |
| OCR recipe | 82.4 | 85.7 | |
| DocVQA, Qwen3-VL-4B | Bare | 94.9 | 93.6 |
| InfoSeek, Qwen3-VL-4B | Bare | 18.0 | 18.0 |
| Name card | 23.6 | 22.9 | |
| Store only | 33.5 | 33.4 |
| Setting and student | ViperGPT-S | DFSDT-S | Tool use | HC | |
|---|---|---|---|---|---|
| Prose | Example | ||||
| BLINK reflectance, 4B | 44.0 | 31.3 | 44.8 | 0% | 64.9 |
| BLINK depth, 4B | 83.9 a | 78.2 | 81.5 | 0% | 95.2 |
| BLINK correspondence, 4B | 42.6 | 51.5 | 58.1 | 0% | 67.6 |
| DocVQA, Gemma-4-E4B | 75.0 | n/a | 75.2 | 81% | 84.9 |
| DocVQA, Qwen3-VL-4B | 44.5 | n/a | 95.0 | 97% | n/a |
| Setting | Reasoning setting | Items | Score |
|---|---|---|---|
| DocVQA | Medium | 32 | 81.2 |
| DocVQA | Low | 58 | 65.5 |
| InfoSeek | Reported pilot | 44 | 29.5 |
| Setting | HC probe / gate | HC | Answer-LoRA | Trajectory SFT |
|---|---|---|---|---|
| InfoSeek | 50 / 50 | 33.6 (38.0, 25.9, 36.9) | 20.0 (3 ep) / 23.3 (10 ep) | 30.2 (3 ep) / 35.3 (10 ep) |
| 150 / 300, original | 34.1 / 35.6 | 23.7 | 36.9 | |
| 150 / 600 | 40.3 | – | – | |
| 600 / 600 | 39.7 | – | 40.7 (1,199 accepted) | |
| DocVQA | 50 / 50 | 84.7 (82.8, 84.7, 86.6) | 81.1 | – |
| 150 / 300 | 85.6 | – | – |
| Policy | Bare | Canonical only | Base | Own |
|---|---|---|---|---|
| SmolVLA-0.45B | 3.8 | 33.2 | 37.0 | 34.5 |
| SmolVLA, tuned on perturbations | 11.8 | 59.0 | 64.7 | 63.2 |
| X-VLA-0.9B | 59.7 | 60.7 | 63.3 | – |
| Checkpoint | arm | valid completion (%) | false aborts | correct refusals / 40 | mixed |
|---|---|---|---|---|---|
| SmolVLA base | bare | 4.1 | 0 | 0 | 3.8 |
| SmolVLA base | canonicalization | 35.5 | 0 | 0 | 33.2 |
| SmolVLA base | human recipe v2 | 30.7 | 106 | 9 | 30.2 |
| SmolVLA base | HC | 39.6 | 0 | 0 | 37.0 |
| SmolVLA base | HC (round 2) | 35.9 | 62 | 6 | 34.5 |
| SmolVLA plus-tuned | bare | 12.7 | 0 | 0 | 11.8 |
| Replicate | arm | round | mixed | completion (%) | false aborts | refusals /40 | vs own rollout |
|---|---|---|---|---|---|---|---|
| 1 | (rollout in A) | – | 28.8 | 28.2 | 200 | 15 | – |
| 1 | A selected | 0 | 29.3 | 28.8 | 200 | 15 | (39/36, ) |
| 1 | (rollout in B) | – | 29.7 | 29.1 | 200 | 15 | – |
| 1 | B selected | 2 | 34.3 | 35.9 | 91 | 5 | (86/58, ) |
| 2 | (rollout in A) | – | 23.7 | 22.3 | 279 | 17 | – |
| 2 | A selected | 1 | 36.3 | 37.3 | 71 | 9 | (111/35, ) |
| Setting | HC calls/item | HC output tokens | CoT output tokens | Teacher $ / build |
|---|---|---|---|---|
| InfoSeek | 1.6 | 4.9 | 262 | 6.9 (with store) |
| DocVQA | 1.0 | 11.2 | 97 | 0.9 |
| BLINK reflectance | 0.08 | 0.1 | 237 | 2.5 |
| BLINK correspondence | 0.45 | 0.7 | 278 | 3.0 |
| BLINK depth | 0.07 | 0.1 | 255 | 1.8 |