Testing RESTful API is increasingly complicated but indispensable to quality assurance of cloud-native applications. This paper reports a multi-agent system called MASTEST that combines LLM-based intelligent agents and programmed agents to automate REST API testing. They form a complete tool chain covering the whole workflow of REST API test with API specification in the OpenAPI Swagger format as the input. It also incorporates human testers in the process to review and correct LLM generated test artefacts to control the quality of testing activities. MASTEST is evaluated on two LLMs, GPT-4o and DeepSeek V3.1 Reasoner with five public APIs. Its performances on various testing activities are measured by a wide range of metrics, including adequacy and coverage metrics, the syntax and data type correctness of generated test scripts, the usability of LLM generated test cases and scripts, as well as the bug detection ability. Experiment results demonstrated that both DeepSeek and GPT-4o achieved a high overall performance but had strengths and weaknesses on different testing activities. MASTEST generated test cases achieved 94% and 98% unit test coverage and 79% and 78% system test coverage for GPT-4o and DeepSeek respectively in comparison with human designed test cases. The generated test scripts maintained 100% syntax correctness and only required minimal manual edits for semantic correctness. The generated test scripts contain assertions on the expected status code as well as contents in the response messages. They are highly capable of detecting bugs in the REST APIs. Experiment data shows that the bug detection rates are between 2.13 to 4.50 per operation. These findings indicate that MASTEST is highly efficient and effective.
Figures & tables
Fig. 1: Workflow and MASTEST Architecture.
Fig. 2: Prompt Template for Generating Test Scripts.
Agent
Variables
Unit test scenario generator
API method details
System test scenario generator
A list of API method details
Test script generator
API method details, Scenario description, Service host URL
Test script data type checker
Scenario description, API method details, Test script
Status code coverage calculator
Scenario description, API method details, Test script, Execution results
TABLE I: Variables in Prompt Templates
Fig. 3: Interface of API Operation Inspector.
Nav Panel Item
Contents of Linked Web Page
Project homepage
It gives the name and URL of the API specification
API specification
It lists the operations of the API and shows their states in the project
Unit test scenarios
Each lists the set of unit test scenarios generated for one API operation and shows their states
System test scenarios
Each lists the set of scenarios generated for all operations of the system or a selected subset of operations and shows their states
Operational scenarios
Each lists the subset of system test scenarios that involves a specific API operation and shows their states
Test scripts
Each lists the set of test scripts generated for one test scenario, shows their states and populated with test results after the test script is executed.
TABLE II: Links Between Navigation Panel and Web Pages
Entity
Derived Entities
Size Metrics
Progress Metrics
Quality Metrics
API Spec
API Ops
# API Ops
# API Ops unit test completed, # API Ops system test completed
% Test scripts accepted, # Test scripts manually added, # Test scripts edited, # Test scripts failed, # Syntax errors, # Data type errors, # Semantic errors
TABLE III: Derived Entities and Metrics Associated to Each Subject Entity
Name
Methods
#Ops
URL
Car
Get, Post
24
https://carapi.app
Petstore
Put, Post,
19
https://petstore3.swagger.io/api/v3
Get, Delete
Bills
Get
19
https://bills-api.parliament.uk
Canada Holidays
Get
4
https://canada-holidays.ca
Cat Fact
Get
3
https://catfact.ninja
TABLE IV: The Benchmark Dataset
API
GPT-4o
DeepSeek
Car
97%
98%
Petstore3
100%
100%
Bills
95%
91%
Canada Holidays
78%
100%
Cat Fact
100%
100%
Average
94%
98%
TABLE V: Unit Test Scenario Coverage
API
GPT-4o
DeepSeek
Car
70%
75%
Petstore
76.47%
75%
Bills
50%
40%
Canada Holidays
100%
100%
Cat Fact
100%
100%
Average
79%
78%
TABLE VI: System Test Scenarios Coverage
API
GPT-4o
DeepSeek
Car
55
51
Petstore3
36
44
Bills
42
41
Canada Holidays
18
18
Cat Fact
7
8
Total
158
162
TABLE VII: Numbers of Bugs Detected
API
#Ops
GPT-4o
DeepSeek
Car
24
2.29
2.13
Petstore
19
1.89
2.32
Bills
19
2.21
2.16
Canada Holidays
4
4.50
4.50
Cat Fact
3
2.33
2.67
TABLE VIII: Number of Bugs Detected Per Operation
API
GPT-4o
DeepSeek
Car
86.51%
92.22%
Petstore
68.67%
86.74%
Bills
64.44%
81.50%
Canada Holidays
72.41%
92.64%
Cat Fact
80.77%
86.11%
Average
74.56%
87.84%
TABLE IX: Data Type Correctness
API
GPT-4o
DeepSeek
Car
44.27
15.4
Petstore
14.85
40.27
Bills
9.29
31.97
Canada Holidays
63.2
10.45
Cat Fact
20.12
25.02
Average
30.35
24.62
TABLE X: Usability Measured by Average Editing Distance