Rotated, but How Far? Diagnosing and Improving Object-Rotation Reasoning in VLMs
Organizations: The University of Queensland · Westlake University
Abstract
Vision-language models (VLMs) can detect that an object has rotated across views, but cannot reliably tell by how much. We introduce OR-Bench, a fine-grained benchmark for object-rotation reasoning with eight tasks covering rotation detection, rotation magnitude estimation, and multi-view rotation reasoning. Across 12 VLMs, the gap is stark: the strongest models approach 100% accuracy on detection, yet even coarse magnitude estimation is near chance. When asked for exact angles, models place 91.8--100% of their predictions on just , , and , a failure we term canonical-angle collapse. This collapse persists even without visual input. Representation probing shows that missing information is only part of the explanation. Although rotation information becomes less recoverable at finer granularity, substantial coarse-grained information remains, and a simple linear probe outperforms the models' generated answers. This suggests that VLMs underuse rotation information they already encode. We therefore propose RotationCue, a lightweight decoder that recovers coarse rotation information from the VLM's own frozen representations and feeds it back to the model as intermediate textual context. Across three VLMs, RotationCue improves every model--task combination on OR-Bench, raising macro-average accuracy by 7.9--12.6 points while preserving general capabilities.
Figures & tables
| Level | Task | Qwen3-VL-8B | Qwen2.5-VL-7B | InternVL3.5-8B | Mean |
| Two-view | Rotation Detection | 100.0 (+0.4) | 99.0 (+13.0) | 100.0 (+0.8) | +4.7 |
| Coarse Angle | 66.7 (+26.0) | 63.4 (+21.8) | 60.5 (+28.4) | +25.4 | |
| Threshold | 55.6 (+18.6) | 57.2 (+19.8) | 44.0 (+6.6) | +15.0 | |
| Exact Angle (MCQ) | 32.9 (+8.2) | 33.7 (+11.5) | 28.0 (+5.8) | +8.5 | |
| Exact Angle (Open) | 34.6 (+9.1) | 27.2 (+4.1) | 27.2 (+3.7) | +5.6 | |
| Multi-view | View Matching | 67.6 (+2.3) | 57.4 (+1.4) | 50.9 (+1.8) | +1.8 |
| (a) Source of intermediate context | (b) RotationCue content quality | ||||||||
| Model | Raw | ZS-CoT | 5-shot | Self-gen. | RotationCue | Shuffled | Raw | RotationCue | Oracle |
| Qwen3-VL-8B | 58.0 | 58.0 | 65.3 | ||||||
| Qwen2.5-VL-7B | 54.9 | 54.9 | 63.6 | ||||||
| InternVL3.5-8B | 50.9 | 50.9 | 59.6 | ||||||
| OR-Bench | External benchmarks | ||||||
| Method | Macro-8 | Micro-8 | VSR-ZS | CV 2D | CV 3D | CV Avg. | SpinBench |
| Qwen3-VL-8B (Raw) | 46.8 | 53.0 | 83.6 | 80.1 | 93.3 | 86.7 | 67.0 |
| Rotation-SFT | 51.2 (+4.4) | 56.3 (+3.3) | 75.6 ( ) | 75.4 ( ) | 89.0 ( ) | 82.2 ( ) | 61.8 ( ) |
| RotationCue | 58.0 (+11.2) | 62.9 (+9.9) | 84.0 (+0.4) | 80.1 (—) | 92.5 ( ) | 86.3 ( ) | 69.2 (+2.2) |
Appendix figures & tables16 assets
Supplementary material from the paper’s appendix.
Appendix
| Task Group | Task | Examples | Imgs./Ex. | Uniform reference |
| Two-view | Rotation Detection | 486 | 2 | 50.0% |
| Coarse Angle | 243 | 2 | 33.3% | |
| Threshold | 243 | 2 | 33.3% | |
| Exact Angle (MCQ) | 243 | 2 | 11.1% | |
| Exact Angle (Open) | 243 | 2 | 11.1% | |
| Multi-view | View Matching | 216 | 4 | 33.3% |
| Model | ||
| GPT-4o | 82.0 | 62.3 |
| GPT-4o-mini | 98.4 | 70.5 |
| GPT-5.4 | 68.9 | 28.7 |
| VE probes | Output-head readout | |||||
| Model | L8 | L16 | Final | LM Head | ||
| Qwen3-VL-8B | 38.4 | 22.2 | 13.4 | 27.2 | 5.6 | |
| Qwen2.5-VL-7B | 29.2 | 31.0 | 13.9 | 9.9 | 6.2 | |
| InternVL3.5-8B | 31.0 | 20.8 | 24.1 | 24.1 | 6.2 | |
| Coarse Angle | Threshold | |||||
| Model | Probe | Native | Probe | Native | ||
| Qwen3-VL-8B | 22.2 | 11.1 | 32.1 | 5.6 | ||
| Qwen2.5-VL-7B | 16.0 | 12.3 | 21.0 | 6.2 | ||
| InternVL3.5-8B | 21.0 | -1.9 | 14.2 | 6.2 | ||
| Model | True labels | Random labels |
| Qwen3-VL-8B | 35.8 | |
| Qwen2.5-VL-7B | 28.4 | |
| InternVL3.5-8B | 38.9 |
| Base VLM | Decoder input | Image preprocessing |
| Qwen3-VL-8B | Vision L8 mean | Qwen AutoProcessor, dynamic resolution |
| Qwen2.5-VL-7B | Vision L8 mean | Qwen AutoProcessor, dynamic resolution |
| InternVL3.5-8B | Vision L8 mean | , at most one tile |
| Configuration | Value |
| Training pairs | 1,200,000 |
| Validation pairs | 50,000 |
| Epochs | 6 |
| Batch size | 1024 |
| Optimizer | AdamW |
| Learning rate |
| Model | Epoch | Same-object | Rotation-present | Angle-range | Joint |
| Qwen3-VL-8B | 5 | 98.70 | 98.68 | 67.44 | 77.75 |
| Qwen2.5-VL-7B | 6 | 98.03 | 98.01 | 63.85 | 74.95 |
| InternVL3.5-8B | 5 | 97.28 | 97.26 | 67.65 | 76.47 |
| Task | Constructed pairs | # Cards |
| View Matching | Ref–A, Ref–B, Ref–C | 3 |
| Post-Rotation View | Ref–A, Ref–B, Ref–C, Ref–D | 4 |
| Rotation Comparison | A1–A2, B1–B2 | 2 |
| Base VLM | Raw | Same + Rotation | Angle Range | Full RotationCue |
| Qwen3-VL-8B | 46.8 | 48.5 | 56.4 | 58.0 |
| Qwen2.5-VL-7B | 42.3 | 43.4 | 50.7 | 54.9 |
| InternVL3.5-8B | 43.0 | 45.6 | 50.3 | 50.9 |
| Method | Acc. Mean | Acc. Std | Exact Cons. | Flip Rate | Entropy | Correct-Stable |
| Raw | 56.6 | 0.7 | 84.0 | 16.0 | 0.131 | 50.5 |
| Zero-shot CoT | 58.2 | 1.2 | 71.3 | 28.8 | 0.263 | 47.3 |
| RotationCue | 64.8 | 0.5 | 86.0 | 14.0 | 0.111 | 59.0 |
| Benchmark | # Examples | RotationCue construction |
| VSR-ZS | 1,222 | Duplicate single image: – |
| CV-Bench | 2,638 | Duplicate single image: – |
| SpinBench | 2,739 | Task-aware pair construction with Raw fallback |
| Configuration | Value |
| Base model | Qwen3-VL-8B |
| Training examples | 900,000 |
| Validation examples | 20,000 |
| Epochs | 1 |
| LoRA rank | 8 |
| LoRA alpha | 16 |
| Task | N | Raw | Rotation-SFT | |
| Rotation Detection | 486 | 99.6 | 89.9 | |
| Coarse Angle | 243 | 40.7 | 85.2 | |
| Threshold | 243 | 37.0 | 41.6 | |
| Exact Angle (MCQ) | 243 | 24.7 | 40.7 | |
| Exact Angle (Open) | 243 | 25.5 | 27.6 | |
| View Matching | 216 | 65.3 | 54.2 |