Joint spatial-geometric and analytic function reasoning requires translating a perceived spatial configuration into a symbolic function whose executed curve satisfies geometric constraints. We present GAGR-Lab, a framework for measuring this capability through Cartesian game scenes, explicit function semantics, and authoritative Rust trajectory execution. It distinguishes spatial perception, metric grounding, geometric relations, function interpretation, function construction, and constrained synthesis. We specify four configurable scene-difficulty presets and a prospective 24-cell diagnostic design, while reporting only the subset actually evaluated. A bounded pilot of one hosted model (Llama 3.2 11B Vision Instruct) using two API credentials as execution replicas yields 72 balanced games with 432 attempts, 429 valid provider responses, and no target hits; exploratory ordinary-function prompt variants also fail to hit, while the structured localization interface yields no scoreable outputs. A privileged analytic search control independently succeeds on 600 directional cases from 300 generated scenes, with exact repeatability and 1,200 successful vertical-reflection or translation checks. The framework separates serving reliability, symbolic compliance, and geometric success, and preserves exact model-visible inputs and realized paths. A staged protocol outlines diagnostic calibration, held-out replication, multi-model comparison, and paired robustness tests. The contribution is an operational research framework with an executed pilot and a clearly identified prospective study plan; the full difficulty matrix and comparative model results remain untested.
Figures & tables
Dimension
Diagnostic question
Proposed observable score
Entity perception
Which objects are shooter, target, and obstacles?
Entity accuracy; pixel-region error
Metric grounding
Where are those objects in world coordinates?
Coordinate error; scoreable-output coverage
Geometric relations
What are their directions, slopes, distances, and blocking relations?
Numeric error; relation accuracy
Function to geometry
What trajectory will a supplied function produce under anchoring?
Predicted-path error; outcome accuracy
Geometry to function
Which function realizes specified endpoint relationships?
Engine hit; target approach; validity
Constrained synthesis
Which function reaches the target while satisfying obstacle and path constraints?
Hit before collision; clearance; constraint satisfaction
Table 1: Proposed component tasks for joint spatial-geometric and analytic function reasoning. Independent component scores are not all available in the current pilot.
Preset
Player y span
Obstacles
Size multiplier
Corridor half-width
Easy
±3
0–1
0.80
4.00
Medium
±6
2–4
1.00
3.00
Hard
±9
4–8
1.25
2.00
Extreme
±11
6–12
1.50
1.25
Table 2: Implemented default difficulty controls in world units, except the dimensionless size multiplier. These presets have not been evaluated as a four-level model difficulty sweep in this study.
Component task
Easy
Medium
Hard
Extreme
Entity perception
P
P
P
P
Metric grounding
P
P
P
P
Geometric relations
P
P
P
P
Function to geometry
P
P
P
P
Geometry to function
P
P
P
P
Constrained synthesis
P
P
P
P
Table 3: Prospective 24-cell diagnostic matrix. P means planned, not passed or evaluated. The current custom-profile pilot is reported separately and is not substituted for this matrix.
Family
Input
Games
Attempts
HTTP 200
Valid
Hits
Horizontal
P
12
72
71
71
0
Horizontal
P+C
12
72
71
71
0
Slanted
P
12
72
72
72
0
Slanted
P+C
12
72
72
72
0
One obstacle
P
12
72
71
71
0
One obstacle
P+C
12
72
72
72
0
Table 4: Primary pilot counts. P is the pixel condition; P+C adds exact endpoint coordinates. Repeated assignments share scene geometry and are not independent scene samples.
One-shot condition
Requests
Format valid
Engine valid
Hits
Canonical
24
24
24
0
No full examples
24
24
24
0
Anchoring equation
24
24
24
0
Metric JSON
36
0
0
—
Table 5: Exploratory diagnostics. Metric JSON actions were rejected before execution, so their hit outcome and localization accuracy are unavailable, not zero.
Offline family
Directional cases
Direct-line hits
Search-control hits
Horizontal
200
200
200
Slanted
200
200
200
One obstacle
200
0
200
Table 6: Independent engine controls. The search uses exact geometry and up to nine engine previews, which the model does not receive. These are execution checks, not inter-agent rankings.
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 1: Executed model trajectories (red) and privileged analytic controls (blue) on two primary scenes. The model shots exit the world; the controls enter the target. The obstacle is shown in black. Coordinates and trajectories come from the archived scene and authoritative traces, rather than from model explanations. Both panels use the same world scale.
Legacy cohort
Games
Attempts
Engine valid
Hits
Draws
H1
6
34
26
4
2
H2
12
180
141
4
8
H3
12
221
9
1
11
Appendix
Table 7: Retrospective legacy counts. Engine valid refers to the historical parser/extraction policy, not the current strict full-response output gate. Rows are descriptive cohorts, not comparable benchmark scores.