Testing RESTful API is increasingly complicated but indispensable to quality assurance of cloud-native applications. This paper reports a multi-agent system called MASTEST that combines LLM-based intelligent agents and programmed agents to automate REST API testing. They form a complete tool chain covering the whole workflow of REST API test with API specification in the OpenAPI Swagger format as the input. It also incorporates human testers in the process to review and correct LLM generated test artefacts to control the quality of testing activities. MASTEST is evaluated on two LLMs, GPT-4o and DeepSeek V3.1 Reasoner with five public APIs. Its performances on various testing activities are measured by a wide range of metrics, including adequacy and coverage metrics, the syntax and data type correctness of generated test scripts, the usability of LLM generated test cases and scripts, as well as the bug detection ability. Experiment results demonstrated that both DeepSeek and GPT-4o achieved a high overall performance but had strengths and weaknesses on different testing activities. MASTEST generated test cases achieved 94% and 98% unit test coverage and 79% and 78% system test coverage for GPT-4o and DeepSeek respectively in comparison with human designed test cases. The generated test scripts maintained 100% syntax correctness and only required minimal manual edits for semantic correctness. The generated test scripts contain assertions on the expected status code as well as contents in the response messages. They are highly capable of detecting bugs in the REST APIs. Experiment data shows that the bug detection rates are between 2.13 to 4.50 per operation. These findings indicate that MASTEST is highly efficient and effective.
Figures & tables
Fig. 1: Workflow and MASTEST Architecture.
Fig. 2: Prompt Template for Generating Test Scripts.
Agent
Variables
Unit test scenario generator
API method details
System test scenario generator
A list of API method details
Test script generator
API method details, Scenario description, Service host URL
Test script data type checker
Scenario description, API method details, Test script
Status code coverage calculator
Scenario description, API method details, Test script, Execution results
TABLE I: Variables in Prompt Templates
Fig. 3: Interface of API Operation Inspector.
Nav Panel Item
Contents of Linked Web Page
Project homepage
It gives the name and URL of the API specification
API specification
It lists the operations of the API and shows their states in the project
Unit test scenarios
Each lists the set of unit test scenarios generated for one API operation and shows their states
System test scenarios
Each lists the set of scenarios generated for all operations of the system or a selected subset of operations and shows their states
Operational scenarios
Each lists the subset of system test scenarios that involves a specific API operation and shows their states
Test scripts
Each lists the set of test scripts generated for one test scenario, shows their states and populated with test results after the test script is executed.
TABLE II: Links Between Navigation Panel and Web Pages
Entity
Derived Entities
Size Metrics
Progress Metrics
Quality Metrics
API Spec
API Ops
# API Ops
# API Ops unit test completed, # API Ops system test completed
% Test scripts accepted, # Test scripts manually added, # Test scripts edited, # Test scripts failed, # Syntax errors, # Data type errors, # Semantic errors
TABLE III: Derived Entities and Metrics Associated to Each Subject Entity
Name
Methods
#Ops
URL
Car
Get, Post
24
https://carapi.app
Petstore
Put, Post,
19
https://petstore3.swagger.io/api/v3
Get, Delete
Bills
Get
19
https://bills-api.parliament.uk
Canada Holidays
Get
4
https://canada-holidays.ca
Cat Fact
Get
3
https://catfact.ninja
TABLE IV: The Benchmark Dataset
API
GPT-4o
DeepSeek
Car
97%
98%
Petstore3
100%
100%
Bills
95%
91%
Canada Holidays
78%
100%
Cat Fact
100%
100%
Average
94%
98%
TABLE V: Unit Test Scenario Coverage
API
GPT-4o
DeepSeek
Car
70%
75%
Petstore
76.47%
75%
Bills
50%
40%
Canada Holidays
100%
100%
Cat Fact
100%
100%
Average
79%
78%
TABLE VI: System Test Scenarios Coverage
API
GPT-4o
DeepSeek
Car
55
51
Petstore3
36
44
Bills
42
41
Canada Holidays
18
18
Cat Fact
7
8
Total
158
162
TABLE VII: Numbers of Bugs Detected
API
#Ops
GPT-4o
DeepSeek
Car
24
2.29
2.13
Petstore
19
1.89
2.32
Bills
19
2.21
2.16
Canada Holidays
4
4.50
4.50
Cat Fact
3
2.33
2.67
TABLE VIII: Number of Bugs Detected Per Operation
API
GPT-4o
DeepSeek
Car
86.51%
92.22%
Petstore
68.67%
86.74%
Bills
64.44%
81.50%
Canada Holidays
72.41%
92.64%
Cat Fact
80.77%
86.11%
Average
74.56%
87.84%
TABLE IX: Data Type Correctness
API
GPT-4o
DeepSeek
Car
44.27
15.4
Petstore
14.85
40.27
Bills
9.29
31.97
Canada Holidays
63.2
10.45
Cat Fact
20.12
25.02
Average
30.35
24.62
TABLE X: Usability Measured by Average Editing Distance
As REST APIs become an increasingly significant part of software systems, their validation is becoming more critical. Hence, testing and uncovering underlying issues are of utmost importance for improving software quality. However, testing REST APIs is challenging mainly due to the difficulty of assessing whether the output of an API call is correct, i.e., the test oracle problem. Metamorphic testing is a specification-based testing approach for situations where correct outputs are unknown or not specified explicitly. To check the correctness of a system, relations between the different outputs are specified. We present ARMeta, a tool-supported approach that uses an LLM-based multi-agent workflow to support metamorphic testing of REST APIs documented with OpenAPI. The agentic workflow is used to identify metamorphic test scenarios and specify them in the Given-When-Then format. These scenarios are automatically implemented as executable tests and executed against the system under test. We evaluate ARMeta on two publicly available web applications that expose REST interfaces and compare its performance with a scenario-based testing baseline. The results show that ARMeta explores behaviors that serve as a complement to existing scenario-based testing approaches.
Shehroz Khan, Abdullah Mughees, Gaadha Sudheerbabu +2
Existing REST API testing tools are typically evaluated using code coverage and crash-based fault metrics. However, recent LLM-based approaches increasingly generate tests from NL requirements to validate functional behaviour, making traditional metrics weak proxies for whether generated tests validate intended behaviour. To address this gap, we present RESTestBench, a benchmark comprising three REST services paired with manually verified NL requirements in both precise and vague variants, enabling controlled and reproducible evaluation of requirement-based test generation. RESTestBench further introduces a requirements-based mutation testing metric that measures the fault-detection effectiveness of a generated test case with respect to a specific requirement, extending the property-based approach of Bartocci et al. . Using RESTestBench, we evaluate two approaches across multiple state-of-the-art LLMs: (i) non-refinement-based generation, and (ii) refinement-based generation guided by interaction with the running SUT. In the refinement experiments, RESTestBench assesses how exposure to the actual implementation, valid or mutated, affects test effectiveness. Our results show that test effectiveness drops considerably when the generator interacts with faulty or mutated code, especially for vague requirements, sometimes negating the benefit of refinement and indicating that incorporating actual SUT behaviour is unnecessary when requirement detail is high.
Leon Kogler, Stefan Hangler, Maximilian Ehrhart +3
CASABLANCA hotelsoftware GmbH Schönwies, Austria · University of Innsbruck Innsbruck, Austria · Technical University of Munich Munich, Germany +1
Malformed, missing, or boundary-value inputs in microservice APIs can cascade across dependent services, threatening reliability. Robustness testing systematically exercises such inputs to expose server-side failures, but generating diverse, effective tests remains challenging. Large Language Models can generate such tests from API specifications; however, it is unknown whether different models and prompt strategies produce diverse failure sets or converge on the same failures. We report a controlled experiment applying 7 prompt strategies to 3 open-source LLMs (14B-70B parameters) targeting 2 architecturally distinct microservice systems: one Java monolingual (6 services, 9 failure modes) and one polyglot (27 services, 14 failure modes), yielding 38 valid runs and 663 generated tests. We find that prompt strategy explains more variation in diversity than model size: a Structured prompt collapses diversity entirely, while a single model varied across three prompt strategies achieves complete failure-mode coverage on one system, outperforming any multi-model ensemble under a fixed prompt. We introduce two strategies, Guided and GuidedFewShot, that embed a mutation taxonomy from prior robustness testing research as domain context. GuidedFewShot achieves the highest single-run coverage on both systems (5 of 9 and 8 of 14 failure modes) while maintaining low cross-model similarity. A key lesson is that taxonomy rules alone are insufficient: LLMs cannot distinguish key-absent from value-empty mutations without concrete examples. Findings replicate across both systems.
Hrushitha Goud Tigulla, Marco Vieira
College of Computing and Informatics University of North Carolina at Charlotte Charlotte, USA