Programming Manufacturing Robots with Imperfect AI: LLMs as Tuning Experts for FDM Print Configuration Selection
Organizations: Robotics Institute at Carnegie Mellon University
Abstract
We use fused deposition modeling (FDM) 3D printing as a case study of how manufacturing robots can use imperfect AI. In FDM, print configuration strongly affects output quality. Yet, novice users typically rely on default configurations, trial-and-error, or direct recommendations from generic AI models (e.g., ChatGPT). These strategies can produce complete prints, but they do not reliably meet specific objectives. We present a modular approach that treats an LLM as a source of tuning expertise. We embed this source of expertise within a Bayesian optimization loop. An approximate evaluator scores each candidate print configuration and returns structured diagnostics, which the LLM uses to propose natural-language adjustments that are compiled into machine-actionable guidance for optimization. On 100 Thingi10k parts, our LLM-guided loop achieves the best configuration on 78% objects with 0% likely-to-fail cases, while direct AI model recommendations are rarely best and exhibit 15% likely-to-fail cases. These results suggest that LLMs provide more value as constrained decision modules in optimization loops than as end-to-end oracles for print configuration selection. We expect this result to extend to broader LLM-based robot programming.
Figures & tables
| Change type (example cues) | Residual | Params |
| Directional preference (increase by ) | ||
| Equality/target (set ) ( should equal ) | ||
| Range/box ( ; ) | ||
| Margin (make ) | ||
| Ratio target ( ) | ||
| Sum cap ( ) |
| Aggregator (example cues) | Form | Params |
| Soft-AND (conjunction): “all of …” | ||
| Soft-OR (disjunction/ at least one): “ or , any of these” |
| Method | Median | % of objects that are best or within of best | Flagged likely to fail | ||
| Best | Within 1% | Within 5% | |||
| Original orientation w/ default parameters | 0.37 | 0 | 2 | 12 | 6 |
| Heuristic orientation w/ default parameters | 0.26 | 3 | 4 | 15 | 9 |
| ChatGPT 5.2 Thinking | 0.25 | 8 | 20 | 50 | 15 |
| Gemini 3 Pro (Chat Interface) | 0.21 | 4 | 18 | 48 | 15 |
| (Unguided) Optimization | 0.24 | 7 | 14 | 44 | 0 |
| Row beats column | None | Llama 3.1 70B | GPT-5.2 API | Hand- crafted | Mean winrate |
| None | – | 0.08 | 0.08 | 0.08 | 0.08 |
| Llama 3.1 70B | 0.92 | – | 0.46 | 0.50 | 0.63 |
| GPT-5.2 API | 0.92 | 0.54 | – | 0.54 | 0.67 |
| Handcrafted | 0.92 | 0.50 | 0.46 | – | 0.63 |
| Row beats column | LLM prompted w/o examples | LLM prompted with examples | Hand- crafted | Mean winrate |
| LLM prompted w/o examples | – | 0.38 | 0.39 | 0.38 |
| LLM prompted with examples | 0.62 | – | 0.50 | 0.56 |
| Handcrafted | 0.61 | 0.50 | – | 0.56 |
| Row beats column | None | LLM: one action | LLM: two actions | Hand- crafted | Mean winrate |
| None | – | 0.32 | 0.08 | 0.08 | 0.16 |
| LLM: one action | 0.68 | – | 0.24 | 0.22 | 0.38 |
| LLM: two actions | 0.92 | 0.76 | – | 0.50 | 0.72 |
| Handcrafted | 0.92 | 0.78 | 0.50 | – | 0.74 |