Model Billiards (CF) Bouncing ball Toy car (Pred.) Overall Q (20) O (80) Pred. / Q (20) Q (20) O (60) Q (60) J (160) [6pt][6pt] Foundation VLMs GPT-5.5 35.00 61.25 45.00 50.00 65.00 43.33 60.63 Gemini-3-Flash 20.00 60.00 60.00 20.00 26.67 33.33 47.50 [6pt][6pt] Training-Free Methods ATW (Ours) 85.00 91.25 70.00 60.00 71.67 71.67 81.25
C.2 Additional Ablation Results
Figure reports additional ablations on the same 100-scene subsets used in Section 4.3, with full ATW included for comparison. The identity-binding ablation removes the appearance cards and identity-binding table.
Effects of identity binding and perception modules. Figure shows that removing appearance cards and identity bindings reduces overall accuracy from 82.30% to 70.11% on CLEVRER and from 68.00% to 62.46% on ContPhy, with declines across every question category. These results highlight the importance of maintaining object identities when connecting visual observations, question references, and executable worlds. On CLEVRER, FoundationPose ( Wen et al., 2024 ) tracking achieves 80.46% overall accuracy and the highest predictive accuracy of 97.14%, while matching ATW on counterfactual questions. Its tracking benefits from an input object mesh and a rigid-body prior, making it effective for these rigid-object scenes. However, its rigid-pose formulation does not capture the evolving shapes of cloth, fluids, or soft bodies. We therefore use Track4World ( Lu et al., 2026 ) for trajectory tracking across the diverse physical systems considered in ATW . The GeoCalib ( Veicht et al., 2024 ) variant, which uses a learning-based single-image camera calibration method to estimate camera intrinsics and gravity direction, achieves 78.62% overall accuracy. Full ATW obtains the highest overall score, while different module choices produce distinct accuracy profiles across reasoning categories. C.3 Detailed Physical-Grounding Ablations Tables and provide the category-level results underlying the system-identification and physics-backend ablations. Overall accuracies are aggregated over all questions or options rather than averaged across categories. These experiments are run independently of the agentic world-reasoning ablation. Table 5: Detailed system-identification ablation on the 100-scene ContPhy subset. Category columns report question-answering accuracy (%). Random Search and CEM use equal simulation budgets. Method Property Predictive Counterfactual Goal-driven Overall [6pt][6pt] System Identification (Warp Fixed) Default parameters 70.40 39.53 56.16 80.49 60.31 Random Search 70.40 51.16 63.01 46.34 60.62 CEM (Ours) 80.00 54.65 63.01 65.85 67.69 ContPhy trends. CEM achieves the strongest aggregate performance among the three configurations. Relative to Random Search, it improves physical-property, predictive, goal-driven, and overall accuracy, while matching counterfactual accuracy at 63.01%. These results indicate that CEM-based system identification improves aggregate reasoning performance, with category-specific differences. Table 6: Detailed system-identification and physics-backend ablations on the 100-scene CLEVRER subset. Accuracy is reported in percent (Q: per-question; O: per-option), and the final column reports median trajectory RMSE in meters. The CEM and Warp + CEM rows are the same shared configuration. Method Explanatory Predictive Counterfactual Overall Median trajectory RMSE (m) Q O Q O Q O Q O [6pt][6pt] System Identification (Warp Fixed) Default parameters 25.56 50.62 97.14 98.57 1.08 50.23 26.67 55.11 0.292053 Random Search 41.67 64.77 91.43 92.86 27.03 69.18 43.45 69.49 0.034992 CEM (Ours) 86.11 94.31 94.29 97.14 75.14 91.22 82.76 93.19 0.007655 [6pt][6pt] Physics Backends (CEM Fixed) PhysMind model + CEM 58.33 76.15 94.29 95.71 33.51 72.88 53.56 76.58 0.059138 MuJoCo + CEM 68.89 83.38 91.43 94.29 71.89 87.52 73.79 86.31 0.016017 Warp + CEM (Ours) 86.11 94.31 94.29 97.14 75.14 91.22 82.76 93.19 0.007655 CLEVRER trends. Across the system-identification variants, trajectory RMSE decreases from default parameters to Random Search and then to CEM, while overall per-question and per-option accuracy increase in the same order. The largest reasoning gains appear in explanatory and counterfactual questions, while predictive accuracy remains high across configurations. With CEM fixed, the comparison from the PhysMind model through MuJoCo to Warp shows the same broad relationship between lower trajectory error and higher overall accuracy. Together, these trends show that both system identification and the physical backend determine the quality of the executable world used for downstream reasoning. Appendix D System Prompts This section presents the principal system instructions used in ATW . We organize them according to the main stages of the framework: world-representation routing, object discovery and segmentation, world probing, and evidence-based answer generation. World-representation routing. Inspect the complete video together with all associated questions, but do not answer them at this stage. First identify whether the main interactions involve free collision and support, constrained relative displacement, or another form of material motion. Examine depth ordering and outline truncation during contact. Use a three-dimensional representation when occlusion makes a visible mask boundary different from the physical contact surface, or when folding, covering, and other out-of-plane states cannot be described by planar outlines. Use a planar representation when the relevant connections, endpoints, displacements, and boundaries provide sufficient interaction geometry. Base the decision on visible evidence rather than object names, dataset conventions, or rendering style, and report the observations and remaining uncertainty that support the route. Object discovery and SAM 3 segmentation. Inspect the entire video and indexed frame samples to enumerate every distinct physical entity in the interaction workspace. Include stationary obstacles, supports, and dividers that constrain the experiment, while excluding continuous environment surfaces, shadows, reflections, annotations, and rendering artifacts. Keep same-category instances separate and assign each entity a stable opaque identity. Classify its physical type as rigid, soft body, cloth, or rope from temporal evidence, distinguishing real deformation from occlusion, viewpoint change, and rigid rotation. For each entity, generate short image-grounded descriptions and select a seed frame that prioritizes a complete in-frame outline, visibility, and identity clarity. Accept a SAM 3 candidate only when it covers the intended foreground entity and excludes neighboring objects and background. If no candidate is adequate, refine the description or reselect the seed frame while preserving the target identity. World probing (shared instructions). Investigate the literal question before producing an answer and treat its options independently. Match the referenced entities, world, intervention, and time interval. Missing evidence is unknown rather than false, so inspect returned results before deciding that the evidence is sufficient. Use only the tools exposed by the selected route and issue one to six independent tool calls in a round. A rollout identifier must come from an earlier successful operation, and subsequent queries must read from the corresponding world. Follow returned guidance when evidence is missing and do not repeat an unchanged failed request. Termination or video-fallback requests are issued alone. Because the answer stage receives compact textual evidence, preserve every necessary visual observation together with the evidence reference from which it was obtained. Evidence-based answer generation. Answer the literal question using only its retained evidence. Match the referenced entities, intervention, and time window, and handle negated statements explicitly. Evaluate each relevant claim independently: unknown evidence does not imply a negative decision, and a claim that an event did not occur requires adequate coverage of the requested interval. Cite the evidence supporting the answer and provide a short factual reason without repeating the full tool history. Return one object that follows the specified output contract. If the retained evidence cannot determine the answer, return an explicit unresolved result that identifies the missing evidence rather than inventing a conclusion.