Sketch constraints specify geometric conditions for constructing and modifying parametric CAD models. We study whether predicted constraints allow a given history to reproduce the required initial model and support prescribed dimensional edits. We introduce HistCAD, an executable representation and dataset whose Academic and Industrial collections contain 180,495 parametric construction histories with entity-referenced sketch constraints and retained feature operations. Predictors receive these histories with the geometry and feature definitions retained and explicit sketch constraints removed. They generate constraints for every sketch without seeing the edit request. The benchmark compares models built with alternative constraint sets for the same history under the same dimensional edit. An edit succeeds when the model reproduces the required initial geometry, reaches the target value, preserves specified relations and unedited dimensions in the target sketch, and rebuilds through the complete history. A predictor trained on both collections and supplied with descriptions of the input histories achieves overall edit success of 52.4% on Academic and 29.0% on Industrial. Models retaining only endpoint-connectivity constraints in the target sketch and the history's constraints elsewhere can reach the target and rebuild while failing preservation. For all-sketch predictions, we retain the target-sketch prediction and restore the history's constraints in other sketches. More models then reproduce the required initial geometry, and some of these newly matched models complete the edit. HistCAD connects constraint learning to the construction and revision of parametric CAD models.
Figures & tables
Figure 1: HistCAD geometry examples (left) and an illustrative dimensional edit (right). In the illustrated sketch, Full retains all constraints, whereas Connectivity retains only profile-connectivity constraints. Both CAD models start from the same geometry, reach the target length, and rebuild successfully, but only the Full variant preserves the specified perpendicularity L1⊥L2 .
Collection
Histories
Sketch constraints
Mean steps
Median steps
95th pct. steps
Distinct op. families
DeepCAD
153,534
2,755,327
1.47
1
3
1
Fusion 360
8,609
155,603
2.23
2
6
1
Industrial
18,352
565,984
6.23
3
19
5
Table 1: HistCAD collection sizes and history lengths. Constraint counts use the normalized entries and retain duplicate instances. Operation families are counted per collection. 95th pct. denotes the 95th percentile.
Figure 2: Edit evaluation from initial-model reproduction to dimensional modification. ER is the fraction of assigned tasks whose models pass pre-edit matching, attain the target, and rebuild validly ( T=V=1 ). OES additionally requires preservation of every specified relation and unedited dimension ( P=1 ).
Academic (500 tasks)
Industrial (500 tasks)
Training
Condition
Mpre
ER
Pres.
OES
Mpre
ER
Pres.
OES
Academic only
History only
50.8
44.4
87.8
39.0
25.8
20.4
70.6
14.4
Academic only
Same-history description
59.8
55.8
91.8
51.2
28.2
24.6
80.5
19.8
Academic only
Shuffled description
26.6
24.8
79.8
19.8
15.6
13.4
71.6
9.6
Joint
History only
52.6
45.8
88.2
40.4
48.8
29.2
76.0
22.2
Joint
Same-history description
61.8
57.6
91.0
52.4
51.8
34.2
84.8
29.0
Table 2: All-sketch prediction outcomes (%). Joint denotes Academic–Industrial training. Maxima are bold for Mpre , ER, and OES. The reachable edits included in Pres. can differ between conditions.
Full
Connectivity
Task set
ER
Pres.
OES
ER
Pres.
OES
Δ OES [95% CI]
Academic
99.6
100.0
99.6
100.0
53.0
53.0
46.6 [42.4, 50.6]
Industrial
79.2
100.0
79.2
82.4
44.4
36.6
42.6 [35.8, 49.6]
Table 3: Controlled edit outcomes (%). Both variants have Mpre=100% by task selection. Full’s 100% Pres. follows from the target-sketch feasibility criterion. ΔOES denotes Full minus Connectivity in percentage points, with 95% paired-bootstrap CIs.
Task set
Training
Mpre
ER
Pres.
OES
Academic
Academic only
59.8 → 74.8
55.8 → 68.6
91.8 → 92.1
51.2 → 63.2
Academic
Joint
61.8 → 75.8
57.6 → 69.2
91.0 → 91.6
52.4 → 63.4
Industrial
Academic only
28.2 → 39.8
24.6 → 35.0
80.5 → 79.4
19.8 → 27.8
Industrial
Joint
51.8 → 68.6
34.2 → 48.4
84.8 → 84.7
29.0 → 41.0
Table 4: Edit outcomes before and after restoring the history’s constraints outside the target sketch (same-history descriptions; rates in %). Arrows denote all-sketch prediction → non-target restoration. OES uses all assigned tasks. Restoration adds reachable edits, so Pres. uses different subsets before and after restoration.
Task set
Training
∣H∣
Pre-edit failure
Reachability failure
Preservation failure
Success
Academic
Academic only
282
52
10
15
205
Academic
Joint
275
49
11
13
202
Industrial
Academic only
83
13
1
14
55
Industrial
Joint
253
71
48
15
119
Table 5: Edit outcomes for predictions with history-level F1 ≥0.8 , using same-history descriptions. ∣H∣ is the subgroup size. Membership can differ between predictors. Counts are mutually exclusive and sum to ∣H∣ . Reachability failure means passing initial matching but failing target attainment or valid rebuilding. Preservation failure means P=0 for a reachable edit.
Appendix figures & tables18 assets
Supplementary material from the paper’s appendix.
Appendix
Component
Stored content and execution role
Coordinate system
Euler angles and a translation vector specify the pose of the local sketch plane in model coordinates.
Sketch primitives
Entity identifiers and line, circle, arc, ellipse, elliptical-arc, or NURBS parameters specify the stored sketch geometry.
Sketch constraints
Entries are grouped by type. For example, Length: [[line_1, 20.0]] assigns a length to a line, while Perpendicular: [[line_1, line_2]] makes the two lines perpendicular.
Feature operation
Feature fields specify the operation family and its parameters, including extrusion extents, a revolution axis and angles, and helical pitch and turns. NewBody, Join, Cut, and Intersect specify body interaction. Fillets and chamfers specify edge references and size parameters.
Step order
The sequence determines which body geometry and references are available when each operation executes.
Appendix
Table 6: HistCAD history components used for constraint-aware replay and feature execution.
Primitive
Identifier prefix
Geometric fields
Line
line_
start , end
Circle
circle_
center , radius ; optional taper
Circular arc
arc_
start , middle , end
Ellipse
ellipse_
center , major , minor , angle
Elliptical arc
elliptical_arc_
start , end , major , minor , angle , large_arc , sweep
NURBS
nurbs_
degree , periodic , controls ; weights and knots subject to the conditions below
Appendix
Table 7: Primitive fields in the common HistCAD format. A point is a two-element coordinate array in the local sketch frame. Ellipse semi-axis lengths satisfy major ≥ minor >0 .
JSON key
Single-entry form
Meaning and argument restrictions
Coincident
[p1, p2]
Two points coincide. Both operands are point references.
Concentric
[c1, c2]
Two center-bearing curves share a center.
Equal
[e1, e2]
Compatible entities have equal corresponding sizes: line pairs, round-curve pairs, ellipse pairs, or elliptical-arc pairs.
Fix
e or p
An entity or point is fixed.
Horizontal
l or [p1, p2]
A line is horizontal or two points share a y coordinate.
Vertical
l or [p1, p2]
A line is vertical or two points share an x coordinate.
Appendix
Table 8: Canonical constraint signatures. Each entry is stored in the array under its type key. Referenced entities and points belong to the same sketch.
Figure 3: Hierarchical and flat representations of the same selected sketch region. The hierarchical form stores face–loop membership, whereas HistCAD stores normalized boundary primitives. For the known face selection in the lower panels, shared interior boundary segments cancel after splitting into atomic edges, leaving the exterior and hole boundaries.
Figure 4: Example final geometries from the DeepCAD, Fusion 360 , and Industrial collections.
Panel A: Construction-history length
Collection
Histories
Mean
Median
95th pct.
DeepCAD
153,534
1.47
1
3
Fusion 360
8,609
2.23
2
6
Industrial
18,352
6.23
3
19
Appendix
Table 9: Structural coverage and constraint counts. A: construction steps per history, using nearest-rank percentiles. A sketch and its feature may share a step. B: Industrial histories can contain multiple operation families. C: 565,984 Industrial constraint instances in normalized histories, with duplicate multiplicity retained.
Quantity class
Equivalent dimension forms
Target attachment
Addition form
Curve size
Radius or Diameter on the same circle or arc
curve identifier
Radius
Line size
Length on a line; or unsigned MINIMUM Distance between its own endpoints
line identifier
Length
Other distance
Directional, point–line, cross-entity, and other Distance constraints
canonical entity or point references plus direction and half-space attributes
canonical Distance form
Appendix
Table 10: Equivalent dimensions used to construct tasks and locate their targets.
Comparison
Quantity
Acceptance
Solved-sketch position/length
Corresponding parameter residual divided by sk
≤10−6
Solved-sketch orientation
Unoriented ellipse-axis angular residual
≤10−6 rad
Primitive identity
Corresponding identifiers and primitive kinds
Exact
Body count
Number of final bodies
Equal
Paired-body bounding box
Maximum corresponding-corner distance divided by sb
≤10−6
Paired-body volume
Relative-volume residual
≤10−6
Appendix
Table 11: Pre-edit geometry matching. The body residual tolerances apply to every pair in the minimum-cost assignment. Topological counts are used only to select the body pairs.
Constraint type
Geometric quantity
Tolerance
Coincident
Maximum referenced-point separation
0.01 mm
Concentric
Maximum center separation
0.01 mm
Midpoint
Point-to-segment-midpoint separation
0.01 mm
Parallel / Perpendicular
Line-direction deviation from 0∘/180∘ or 90∘
0.5∘
Angle
Circular angular difference
0.5∘
Horizontal / Vertical
Line-to-axis angle; absolute y - or x -coordinate difference between the two referenced points
0.1∘ for lines; ϵℓ for points
Appendix
Table 12: Quantities and tolerances for target attainment and preservation.
Setting
Value
Base model
Qwen3-8B
Chat format and thinking mode
Qwen3 chat format; training and inference configured with thinking disabled
LoRA rank / alpha / dropout
256 / 128 / 0.05
LoRA modules
query/value projections
Optimizer
AdamW
Learning-rate schedule
Peak 5×10−5 , 1% warmup, cosine decay to 10% of peak
Appendix
Table 13: Training and inference settings.
All types
Non-Coincident
Training
Prediction condition
Micro-F1
P
R
F1
Available, n (%)
Academic prediction test set ( n=6,382 )
Academic only
History only
67.40
40.88
53.73
46.43
6,300 (98.72)
Academic only
Same-history description
72.98
51.57
60.88
55.84
6,299 (98.70)
Academic only
Shuffled description
61.46
33.20
41.35
36.83
6,078 (95.24)
Joint
History only
68.31
42.34
54.30
47.58
6,288 (98.53)
Appendix
Table 14: Constraint scores and availability on the full prediction test sets (%). Scores pool TP, FP, and FN over each test set; Non-Coincident scores exclude Coincident constraints. P and R denote precision and recall. Available n (%) counts complete, parseable, aligned outputs. Its denominator is all 6,382 Academic or 915 Industrial inputs, including those for which generation was skipped.
A. Joint minus Academic-only training
Task set
Prediction condition
ΔMpre [95% CI]
ΔOES [95% CI]
Academic
History only
+1.80 [-1.00, 4.80]
+1.40 [-1.40, 4.20]
Academic
Same-history description
+2.00 [-0.80, 5.00]
+1.20 [-1.60, 4.00]
Academic
Shuffled description
−1.00 [-4.00, 1.80]
+0.40 [-2.20, 3.00]
Industrial
History only
+23.00 [12.96, 32.64]
+7.80 [1.88, 14.42]
Industrial
Same-history description
+23.60 [15.07, 32.40]
+9.20 [3.30, 15.90]
Appendix
Table 15: Paired differences in pre-edit matching and overall success, expressed in percentage points with 95% paired-bootstrap CIs. Each comparison uses the same 500 tasks per set. Panel headings specify subtraction order.
Task set
Training
Matched m
OES (%)
Pre-edit failure → success
Academic
Academic only
299 → 374
51.2 → 63.2
60
Academic
Joint
309 → 379
52.4 → 63.4
55
Industrial
Academic only
141 → 199
19.8 → 27.8
40
Industrial
Joint
259 → 343
29.0 → 41.0
60
Appendix
Table 16: Outcomes before and after restoring the history’s constraints outside the target sketch, using same-history descriptions and an unchanged target-sketch prediction. Both conditions use the complete construction history. Arrows show before-to-after restoration values. OES uses all 500 tasks in each set. The last column counts tasks changing from pre-edit failure to success.
Task set
Training
Input
Pre fail
Reach fail
Pres fail
Success
Academic
Academic only
History only
246
32
27
195
Academic
Academic only
Same-history
201
20
23
256
Academic
Academic only
Shuffled
367
9
25
99
Academic
Joint
History only
237
34
27
202
Academic
Joint
Same-history
191
21
26
262
Academic
Joint
Shuffled
372
10
17
101
Appendix
Table 17: Full-population outcome counts derived from Table 2 . Every row sums to 500. Pre fail: unavailable prediction or failed pre-edit replay or geometry matching; Reach fail: pre-edit matching passes but target attainment or valid rebuilding fails; Pres fail: reachable with P=0 .
A. All tasks
Task set
Training
Prediction condition
ρ [95% CI]
Failed tasks F1 ≥0.8
Successful tasks F1 ≤0.2
Academic
Academic only
History only
0.528 [0.455, 0.595]
76/215 (35.3)
0/12 (0.0)
Academic
Academic only
Same-history description
0.595 [0.533, 0.653]
77/282 (27.3)
0/14 (0.0)
Academic
Joint
History only
0.473 [0.394, 0.546]
79/218 (36.2)
0/12 (0.0)
Academic
Joint
Same-history description
0.605 [0.543, 0.661]
73/275 (26.5)
0/13 (0.0)
Industrial
Academic only
History only
0.416 [0.331, 0.501]
17/47 (36.2)
0/221 (0.0)
Appendix
Table 18: History-level F1 and edit success under all-sketch prediction. A uses all 500 tasks per set, including pre-edit failures. For each predictor, failure and success fractions use the numbers of tasks with prediction F1 ≥0.8 and ≤0.2 , respectively, as denominators. B excludes unavailable predictions but retains those whose models fail pre-edit matching. Brackets give 95% bootstrap CIs.
A. Academic
Full / Connectivity
Success
Reachable; preservation failure
Not reachable
Success
264
234
0
Reachable; preservation failure
0
0
0
Not reachable
1
1
0
B. Industrial
Full / Connectivity
Success
Reachable; preservation failure
Not reachable
Appendix
Table 19: Paired Full–Connectivity outcomes on the complete controlled sets. Rows are Full outcomes and columns are Connectivity outcomes. Each panel contains 500 paired tasks from one collection. Row and column totals give the counts used for ER and OES. Both controls pass pre-edit matching. Preservation failures include unresolved references and failed measurements as well as measured geometric violations.
Figure 5: Failures before the prescribed edit. Initial is the reference. Full and Joint show successful edits. Academic-only* fails before the requested change is applied. The requests are (a) shortening the base of Industrial 01020012 (two chamfers and a fillet) and (b) widening a cut profile in Industrial 01022672 (three fillets). Extra predicted constraint groups prevent history construction in (a,b). × marks these output errors. In (c), the request is to increase the distance from an arc center to an endpoint in Academic 01005050. During initial replay, Academic-only distorts the arc profiles and fails Boolean extrusion (partial model shown). The requests in (d,e) are to lengthen a sleeve-profile segment in Industrial 01027877 and enlarge the inner circle of Industrial 01016299, respectively. The solved Academic-only sketches already mismatch the reference in both cases. Models appear above target sketches. Dark solid lines show displayed geometry and light dashed lines show the initial reference.
Figure 6: Failures after pre-edit matching. Initial is the reference; all other columns show results of the same edit. Full and Joint succeed throughout. Under Academic-only, (a) the lower edge of Industrial 01016487 reaches its requested longer length, but a downstream cut misses the body; (b) the opening in Academic 00465180 reaches its requested larger size, but a perpendicular constraint is skipped and the sketch deforms. These two model panels show incomplete replays. In (c) Academic 01004348, reducing the distance across the lug leaves the hole off-center within the rounded top. In (d) Academic 00592341, extending a horizontal step edge also changes an unedited edge length. In (e) Industrial 01024010, a specified right angle is lost when one rectangle side is lengthened. The last three Academic-only results attain the target and rebuild validly, but fail preservation. Models appear above target sketches; dark solid lines show displayed geometry and light dashed lines show the initial reference. Plus signs mark the centers in (c).
A CAD model is not engineering-grade merely because it looks correct. It must satisfy design requirements, respond predictably to parameter changes, support controlled edits, match a reference structural response under a declared analysis, and connect to other parts through valid joints. We present CADEngBench, a two-track benchmark for these capabilities. CADEngBench-P evaluates 300 parametric parts, each used for one zero-to-CAD task and one functional-editing task (600 tasks in total), through boundary-representation (B-Rep) validity, engineering and DFM checks, parameter-family perturbations, functional editing, and matched linear-static FEA in CalculiX. CADEngBench-A evaluates 150 body pairs through ranked joint retrieval, exact face-and-edge grounding, joint-frame prediction, and kinematic verification. Across eight multimodal, code-capable models, editing supplied CAD is substantially easier than generating it, while complex edits and matched FEA remain difficult. Assembly predictions often locate the relevant region but fail to recover the recorded joint or mating entities. These results show that CAD evaluation must test engineering behavior rather than appearance alone.
We introduce neuralCAD-Edit, the first benchmark for editing 3D CAD models collected from expert CAD engineers. Instead of text conditioning as in prior works, we collect realistic CAD editing requests by capturing videos of professional designers, interacting directly with CAD models in CAD software, while talking, pointing and drawing. We recruited ten consenting designers to contribute to this contained study. We benchmark leading foundation models against human CAD experts carrying out edits, and find a large performance gap in both automatic metrics and human evaluations. Even the best foundation model (GPT 5.2) scores 53% lower (absolute) than CAD experts in human acceptance trials, demonstrating the challenge of neuralCAD-Edit. We hope neuralCAD-Edit will provide a solid foundation against which 3D CAD editing approaches and foundation models can be developed. Code/data: https://autodeskailab.github.io/neuralCAD-Edit
LLM-based CAD agents produce executable parametric programs, but their correction loops may lose evidence about satisfied requirements, faulty operations, and prior repairs. We introduce TraceCAD, a recovery layer that links requested features, modeling steps, failure evidence, and candidate outcomes as persistent state. TraceCAD diagnoses likely faulty operations, searches bounded edits in their dependency regions, validates candidates through execution and preservation checks, and retains successful and failed repair outcomes in reusable skill memory. On DeepCAD-derived benchmarks with 200-model ablations and a 1K-model comparison, TraceCAD achieves competitive geometric quality in terms of IoU, Chamfer distance, and Hausdorff distance. Removing persistent state nearly halves recovery score; removing localized search more than doubles geometric regression and doubles code-agent invocations. Initializing the skill store on disjoint training models further reduces retries, token cost, and latency. These results demonstrate that persistent, localized, and reusable recovery improves final CAD quality and repair reliability.