Agentic AI for Scalable and Robust Optical Systems Control
Authors: Zehao Wang, Mingzhe Han, Wei Cheng, Yue-Kai Huang, Philip Ji, Denton Wu, Mahdi Safari, Flemming Holtorf, +7 more
Organizations: Department of Electrical and Computer Engineering, Duke University, Durham, NC 27708, USA · NEC Laboratories America, Princeton, NJ 08540, USA · Duke Quantum Center and Department of Physics, Duke University, Durham, NC, USA 27708 · Axiomatic AI, Cambridge, MA 02139, USA · Research Laboratory of Electronics, Massachusetts Institute of Technology, Cambridge, MA 02139, USA · Joint Quantum Institute, Department of Physics, and the National Quantum Laboratory (QLab), University of Maryland, College Park, MD 20742, USA · Department of Computer Science, Duke University, Durham, NC 27708, USA
We present AgentOptics, an agentic AI framework for high-fidelity, autonomous optical system control built on the Model Context Protocol (MCP). AgentOptics interprets natural language tasks and executes protocol-compliant actions on heterogeneous optical devices through a structured tool abstraction layer. We implement 64 standardized MCP tools across 8 representative optical devices and construct a 410-task benchmark to evaluate request understanding, role-aware responses, multi-step coordination, robustness to linguistic variation, and error handling. We assess two deployment configurations--commercial online LLMs and locally hosted open-source LLMs--and compare them with LLM-based code generation baselines. AgentOptics achieves 87.7%--99.0% average task success rates, significantly outperforming code-generation approaches, which reach up to 50% success. We further demonstrate broader applicability through five case studies extending beyond device-level control to system orchestration, monitoring, and closed-loop optimization. These include DWDM link provisioning and coordinated monitoring of coherent 400 GbE and analog radio-over-fiber (ARoF) channels; autonomous characterization and bias optimization of a wideband ARoF link carrying 5G fronthaul traffic; multi-span channel provisioning with launch power optimization; closed-loop fiber polarization stabilization; and distributed acoustic sensing (DAS)-based fiber monitoring with LLM-assisted event detection. These results establish AgentOptics as a scalable, robust paradigm for autonomous control and orchestration of heterogeneous optical systems.
Figures & tables
Fig. 1: Optical device control using ROADM, 400 GbE CFP2-DCO, and OSA as examples: (a) Traditional control requires device-specific manuals, custom scripts, and protocol handling. (b) LLM-based control interprets natural-language prompts to generate control code, reducing manual scripting. (c) The proposed AgentOptics framework standardizes control via a unified tool layer, where the MCP client maps prompts to device APIs through tool selection, enabling remote distributed device control and scalable integration of new devices.
Lumentum 400 GbE CFP2 digital coherent optics (CFP2-DCO)
6
Set center frequency/output power/operation mode; get config, …
Optilab LT-12-E-M ARoF Tx
6
Set bias voltage/current; get status, …
APEX Technologies OSA
26
Get power/spectrum; set/get measurement parameters, …
Calient S320 Optical Circuit Switch
4
Get port; add/delete connection; delete all connections
DiCon Microelectromechanical Systems (MEMS) 32 × 32 Optical Switch
2
Get connections; set connection
TABLE I: List of optical devices and validated MCP tools supported by AgentOptics.
Fig. 2: Benchmark workflow for evaluating the performance of AgentOptics and the CodeGen baseline, where reference ground truth is established using human-crafted scripts that are manually validated on physical devices for correctness.
Type
Description
Task example
Paraphrasing
Same meaning, different phrases
• Operate the CFP2 so that port cfp2-opt-1-1 has an output target power setting of − 5 dBm. • Using the CFP2, adjust the output target power parameter on port cfp2-opt-1-1 to − 5 dBm.
Non-sequitur
Adding unrelated information to the task
• Set CFP port cfp2-opt-1-1 power to − 5 dBm; the bench mat has a curled corner.
Error
Task with wrong or lost value
• Missing power value: on the CFP2, set output target power on port cfp2-opt-1-1. • Wrong power value: on the CFP2, set output target power on port cfp2-opt-1-1 to − 100 dBm.
Chain
Sequential related tasks
• First set CFP2 port cfp2-opt-1-1 output target power to − 4 dBm, then read CFP2 output power.
Roles
Task tone as service provider or user
• You are an optical device user; set CFP port cfp2-opt-1-1 power to − 5 dBm.
TABLE II: Five representative task variants evaluated in the agentic optical device control benchmark.
Fig. 3: Task success rate achieved by AgentOptics across varying task complexities using five locally hosted Gemma and Llama models and four online LLMs. The two-action bar aggregates direct and chained two-action tasks.
Fig. 4: Task success rate achieved by AgentOptics across different task variants using five locally hosted Gemma and Llama models and four online LLMs.
Fig. 5: Task success rate achieved by AgentOptics using AgentOptics-Local (Gemma-4-12B) and AgentOptics-Online (Claude Sonnet 4.5), compared with the CodeGen baseline. The two-action bar aggregates direct and chained two-action tasks.
Fig. 6: Task success rate achieved by AgentOptics using AgentOptics-Local (Gemma-4-12B) and AgentOptics-Online (Claude Sonnet 4.5), compared with the CodeGen baseline.
Fig. 7: Trade-off between success rate and average cost for the combined direct dual-action and chained-task subset with AgentOptics and the CodeGen baseline using locally hosted (squares) and online (circles) LLMs. Marker size indicates the relative average execution time.
Calls an invalid function (e.g., calls AP2XXX.get_powower , which does not exist).
CodeGen
Import non-existing library
11.18%
Imports an undefined library (e.g., import lab_api , which is undefined).
AgentOptics
Wrong tool
89.09%
Calls the wrong tool (e.g., arof_get_power instead of arof_read_power ).
AgentOptics
Missing tool
10.91%
Required tools are not invoked (expected OSA-related tools, but called none).
TABLE III: Reasons and examples for CodeGen and MCP-based AgentOptics execution failures. The percentages are calculated with respect to the total number of failures within each approach.
Fig. 8: Representative OSA MCP tool onboarding workflows for manual and LLM-assisted MCP server generation.
Fig. 9: Diagram for DWDM link configuration with ARoF and 400 GbE signals.
Fig. 10: (a) AgentOptics workflow for LLM-assisted wide-bandwidth ARoF 5G new radio (NR) link with an RFSoC ZCU216 board and an ARoF transmitter-receiver pair. (b) and (c) Optimized ARoF transmitter bias voltage across link SNR and BER with different modulation orders, where the vertical dashed line indicates the optimized bias voltage selected by AgentOptics.
Fig. 11: (a) AgentOptics provisions a 400 GbE channel in a two-span link and autonomously optimizes the channel GSNR based on a single-line human language instruction. (b) Autonomous launch power optimization of the CFP2-DCO 400 GbE transmitter (Tx) performed by AgentOptics. (c) Pre-FEC BER optimization of the 400 GbE signal by AgentOptics using an online LLM (Sonnet 4.5) without impacting existing background traffic.
Fig. 12: (a) Experimental setup and control architecture for fiber link polarization stabilization using AgentOptics. (b) Closed-loop polarization stabilization results with deliberate fiber perturbations, showing polarization state and piezo controller actuation over time. Red dashed vertical lines indicate the perturbation times.
Fig. 13: (a) Experimental setup and workflow for AgentOptics-enabled fiber monitoring using DAS. (b) LLM-based reasoning and prompt engineering (PE) for automated event interpretation on the DAS waterfall plot analysis for (c) a stable environment, (d) human-induced pseudo fiber agitation, and (e) a real fiber cut event.
Recent agentic-robotics systems, from Code-asPolicies to modern vision-language-action (VLA) foundation models, presuppose that drivers, SDKs, or ROS-style primitives for the target hardware already exist. Writing those primitives is the dominant engineering cost of bringing up new hardware for agent control. We present Octopus Protocol, a system that collapses that cost to a single shell command. Given only raw OS access and a language-model API key, a coding agent executes a five-stage pipeline--PROBE, IDENTIFY, INTERFACE, SERVE, DEPLOY--to discover connected devices, infer their capabilities, generate a Model Context Protocol (MCP) server with typed tools, and deploy it as a live HTTP endpoint. A persistent daemon then monitors the system, heals broken code, and perceives physical state through the camera tools it generated for itself. Two architectural principles make this work: protocols are prompts, not code, and the coding agent is the runtime. We validate the system on three heterogeneous platforms (PC/WSL, Apple Silicon macOS, Raspberry Pi 4) and on a commercial 6-DOF robotic arm with USB camera feedback. One command onboards the hardware in ~10-15 minutes and exposes up to 30 MCP tools; an MCP-compliant client then performs closed-loop visual-motor control through tools no human wrote.
Large language model agents are increasingly being developed to control a wide range of scientific characterization tools including microscopes and synchrotron beamlines. Research into agentic control of physical infrastructure is nascent and there are few well-established paradigms for how to engineer an agentic system. There are many choices to make when designing a microscopy agent, including the choice of LLM, the number of agents to use, agent responsibilities and delegation rules, retrieval-augmented generation parameters, and more. When designing and optimizing an agentic microscope controller, researchers not only want to ensure that the agent can correctly perform known tasks but also that the agent can generalize to new tasks that it has not encountered before. In this study, we develop a benchmark and trace-logging framework that reveals a) how different choices of agent architecture impact performance at microscopy tasks and b) the limitations of benchmarks for predicting if a particular agent will perform well on unseen microscopy tasks. The framework was used to evaluate one-, two-, and three-agent graph topologies, five LLMs, RAG and context parameters, and operational constraints across 53 microscopy benchmark tests. In total, 105 agent configurations, 1,949 individual test runs, and 49,109 RAG retrievals were recorded. Direct comparisons showed clear differences in latency, token use, cost, and failure mode between configurations. However, surrogate models trained on agent architecture and test results did not reliably predict an agent's performance on new, unseen tasks. These results show that these benchmarks are useful for qualification, regression testing, diagnosis, and direct comparison, but the current heterogeneous test suite does not support a task-independent global configuration model.
Nathan S Johnson, Ian Abshire
Carl Zeiss Research Microscopy Solutions, 5300 Central Boulevard, Dublin, CA 94568
Open Radio Access Networks (O-RAN) promise flexible 6G network access through disaggregated, software-driven components and open interfaces, but this programmability also increases operational complexity. Multiple control loops coexist across the service management layer and RAN Intelligent Controller (RIC), while independently developed control applications can interact in unintended ways. In parallel, recent advances in generative Artificial Intelligence (AI) are enabling a shift from isolated AI models toward agentic AI systems that can interpret goals, coordinate multiple models and control functions, and adapt their behavior over time. This article proposes a multi-scale agentic AI framework for O-RAN that organizes RAN intelligence as a coordinated hierarchy across the Non-Real-Time (Non-RT), Near-Real-Time (Near-RT), and Real-Time (RT) control loops: (i) A Large Language Model (LLM) agent in the Non-RT RIC translates operator intent into policies and governs model lifecycles. (ii) Small Language Model (SLM) agents in the Near-RT RIC execute low-latency optimization and can activate, tune, or disable existing control applications; and (iii) Wireless Physical-layer Foundation Model (WPFM) agents near the distributed unit provide fast inference close to the air interface. We describe how these agents cooperate through standardized O-RAN interfaces and telemetry. Using a proof-of-concept implementation built on open-source models, software, and datasets, we demonstrate the proposed agentic approach in two representative scenarios: robust operation under non-stationary conditions and intent-driven slice resource control.
Hojjat Navidan, Mohammad Cheraghinia, Jaron Fontaine +5
IDLab, Department of Information Technology at Ghent University - imec, Ghent, Belgium · Department of Electrical and Computer Engineering, Princeton University, USA