Closed-loop evaluation of LLM agents for embedded software development
Authors: Jorge García-Carrasco, Sergio García-Carrasco, Alejandro Maté, Juan Trujillo
Organizations: Lucentia Research, Department of Software and Computing Systems, University of Alicante, Ctra. de San Vicente del Raspeig s/n, 03690 Sant Vicent del Raspeig, Spain · Universitary Institute of Materials Technology (IUTM), Universitat Politècnica de València, Plaça Ferrándiz i Carbonell, 03801 Alcoi, Spain
Large language models (LLMs) are increasingly deployed as coding agents that edit files, run builds and tests, inspect execution results, and repair software iteratively. Embedded firmware is a demanding target because correctness depends on closed-loop behavior under sensing, timing, and safety constraints, not only on static source quality. Yet embedded-agent evaluation remains limited and often emphasizes one-shot synthesis or offline correctness. We present a benchmark for closed-loop evaluation of embedded coding agents. Each task provides a plain-text engineering description, constrained workspace, and visible build-and-runtime surface. The agent must translate requirements into implementation and self-verification steps, then iterate until the required device behavior is achieved. The suite contains five embedded-control tasks and four feedback scenarios: one-shot generation, realistic self-verification, CI-style red/green feedback, and oracle-style detailed feedback. The implementation targets simulated ESP32 firmware for reproducibility. We evaluate seven GPT-family and Qwen-family configurations across five tasks and four scenarios, with three repetitions per condition for 420 runs. gpt-5.4 has the highest pass rate among evaluated configurations but does not saturate the benchmark; qwen3.5-27B is the strongest observed local model; and smaller local models degrade sharply in pass rate and search efficiency. These results suggest that capable local embedded coding agents are emerging.
Figures & tables
Figure 1: Main agentic workflow evaluated by the benchmark. The agent iterates locally on visible evidence, explicitly decides whether to submit, and only then triggers a hidden behavioral evaluation.
Tier
Task
Primary added complexity
1
tank_fill_drain
Baseline threshold control with timeout-safe behavior.
2
thermal_chamber_hysteresis
Adds plant dynamics and anticipatory control near the upper band.
3
pressure_vessel_interlock
Adds asynchronous input freshness and explicit interlock safety constraints.
4
mixing_tank_fill_heat
Adds multi-sensor coupling and phase-dependent control decisions.
5
filter_tank_sequence
Adds timer-driven sequence control, evidence accumulation, and recovery behavior.
Table 1: Compact view of the five-task difficulty progression.
Parameter
Hosted GPT models
Local Qwen3.5 models
Runtime
GitHub Copilot subscription via pi
llama.cpp via pi
Temperature
Provider-managed
0.7
Top-p
Provider-managed
0.8
Top-k
Provider-managed
20
Min-p
Provider-managed
0.0
Presence penalty
Provider-managed
1.5
Table 2: Inference and harness configuration used for the evaluated models.
Model
Passes
Avg total tokens
Avg tool calls
Avg hidden evals
Avg self-test runs
gpt-5.4
14/15 (93.3%) [70.2, 98.8]
0.42M
35.5
1.0
2.6
gpt-5.4-mini
8/15 (53.3%) [30.1, 75.2]
1.03M
54.7
1.1
4.4
qwen3.5-27B
9/15 (60.0%) [35.7, 80.2]
6.90M
109.0
1.5
12.2
qwen3.5-35B-A3B
4/15 (26.7%) [10.9, 52.0]
4.40M
92.5
1.5
16.1
qwen3.5-9B
2/15 (13.3%) [3.7, 37.9]
6.76M
113.8
6.6
38.5
qwen3.5-4B
2/15 (13.3%) [3.7, 37.9]
6.14M
114.9
12.7
37.7
Table 3: Pooled primary-scenario results for realistic_self_verify over three repetitions. For all listed models, this corresponds to 15 runs. The pass column reports pooled counts, pass rate, and Wilson 95% confidence interval. Cross-backend comparison emphasizes interaction cost rather than wall-clock.
Figure 2: Two-panel view of the primary realistic_self_verify condition over three repetitions. The left panel shows pass counts out of three repetitions for each model-task pair. The right panel shows the most likely outcome across those repetitions: PASS if pass is the modal outcome, HOST for host-test failure, INT for integration failure, and NONE for no submission.
Figure 3: Pass-rate sensitivity to feedback visibility across the evaluated model set, pooled over three repetitions. Each bar is a percentage over the 15 runs available for that model and scenario, and error bars show Wilson 95% confidence intervals.
Model
oneshot_blind
realistic_self_verify
ci_red_green
oracle_full
gpt-5.4
13/15
14/15
15/15
14/15
gpt-5.4-mini
10/15
8/15
12/15
11/15
qwen3.5-27B
8/15
9/15
9/15
12/15
qwen3.5-35B-A3B
4/15
4/15
8/15
10/15
qwen3.5-9B
2/15
2/15
3/15
7/15
qwen3.5-4B
0/15
2/15
1/15
2/15
Table 4: Exact pooled pass counts for the full four-scenario matrix over three repetitions. Each cell reports the number of passes out of 15 runs for that model and scenario.
Model
Form
Passes
Avg wall-clock
n submit
Avg 1st submit
Avg total tokens
Avg tool calls
qwen3.5-27B
Dense
38/60 (63.3%) [50.7, 74.4]
26.2 min
60/60
18.0 min
4.08M
77.0
qwen3.5-35B-A3B
MoE
26/60 (43.3%) [31.6, 55.9]
16.9 min
60/60
11.9 min
4.19M
80.1
qwen3.5-9B
Dense
14/60 (23.3%) [14.4, 35.4]
25.3 min
56/60
17.6 min
6.09M
103.5
qwen3.5-4B
Dense
5/60 (8.3%) [3.6, 18.1]
21.4 min
50/60
9.4 min
11.55M
179.2
qwen3.5-2B
Dense
0/60 (0.0%) [0.0, 6.0]
29.7 min
8/60
1.3 min
54.95M
716.3
Table 5: Pooled local-only Qwen comparison on shared hardware across three repetitions. Pass denominators are 60 for every model. The pass column reports pooled counts, pass rate, and Wilson 95% confidence interval. The n submit column reports how many runs reached at least one submission; average first-submit time is conditioned on those runs.
Figure 4: Shared-hardware local-model tradeoffs shown as two side-by-side views. Both panels use pass rate rather than raw pass count so that accuracy and resource demand share a common scale. The left panel relates pass rate to average first-submission latency; the right panel relates pass rate to average total token demand.
Lucentia Research, Department of Software and Computing Systems, University of Alicante, Ctra. de Sant Vicent del Raspeig s/n, 03690 Sant Vicent del Raspeig, Alicante, Spain