As large language models (LLMs) advance, AI agents are increasingly deployed in open-world environments to tackle complex sequential tasks (e.g., document processing, cross-application collaboration), relying heavily on actions ranging from GUI operations to semantic APIs. However, three core challenges persist: the "scale dilemma" of massive tool ecosystems exceeding LLM context windows, the "non-stationarity" of tool quality due to updates or outages, and the "heterogeneity" of feedback formats (pixels, text, structured data) creating information silos. To address these, we propose AnyAct, a universal action layer that unifies available capabilities into a self-evolving action space, enabling agents to operate efficiently and reliably in large-scale, dynamic tool ecosystems. AnyAct's core design focuses on two objectives: constructing this action space via hierarchical progressive retrieval (filtering task-relevant actions) and test-time reliability evolution (pruning unreliable actions), and enabling reliability-aware action orchestration through a heterogeneous observation grounding module that unifies multi-modal feedback. Additionally, it defines a hybrid action space (primitive + semantic actions) and optimizes for a balance between task success rate and execution cost. Evaluations on LiveMCPBench and OSMCP (a new benchmark we developed for multi-granularity action collaboration) demonstrate state-of-the-art performance. AnyAct delivers substantial performance gains over baseline methods across various LLM base models on LiveMCPBench and improvements are particularly notable for models with constrained native capabilities. On OSMCP, it achieves 77.27% overall success with only 50 steps, which is half the steps required by most competitors.
Figures & tables
Figure 1: The overall framework of AnyAct for self-evolving agentic automation over a universal action space.
Model
Method
Office
Leisure
Travel
Lifestyle
Finance
Shopping
Overall ( % )
Claude-Sonnet-4
Naive
90.32
64.29
75.00
80.00
78.57
66.67
78.95
AnyAct w/o Retrieval
80.65
71.43
75.00
93.33
90.91
77.78
81.52
AnyAct w/o Evolution
83.87
78.57
58.33
80.00
91.67
33.33
75.27
AnyAct
93.55
92.86
100.00
100.00
83.33
88.89
93.55
DeepSeek-V3
Naive
41.94
42.86
50.00
40.00
28.57
55.56
42.11
AnyAct w/o Retrieval
51.61
42.86
33.33
60.00
50.00
33.33
47.31
Table 1: Main Results on LiveMCPBench.
Method
Model
Steps
vlc
gimp
thunderbird
calc
writer
impress
vs_code
Overall (/)
Overall ( % )
Foundation E2E GUI
Claude-Sonnet-4.5
—
15
4.89/17
13.00/26
10.00/15
17.00/47
14.00/23
22.96/47
14.00/23
95.85/198
48.41
opencua-72b-preview
—
50
8.87/17
14.00/26
12.00/15
18.00/47
13.00/23
26.96/47
18.00/23
110.83/198
55.97
UI-TARS-2-2509
—
100
8.49/17
13.00/26
11.00/15
31.00/47
14.00/23
26.50/47
17.00/23
120.99/198
61.11
Claude-Sonnet-4.5
—
50
9.00/17
15.00/26
10.00/15
31.00/47
15.00/23
27.02/47
16.00/23
123.02/198
62.13
EvoCUA-20260105
—
50
8.39/17
20.00/26
12.00/15
26.00/47
16.0/23
27.96/47
20.00/23
130.35/198
65.83
Table 2: Main results on OSMCP. Step budget denotes the maximum number of steps allowed for each system.
Task ID: 04d9aeaf-7bed-4024-bedb-e10e6f00eb7f
Task Instruction: In a new sheet with 4 headers “Year”, “CA changes”, “FA changes”, and “OA changes”, calculate the percentage annual changes compared to last year in 2015 to 2019 for the Current Assets, Fixed Assets, and Other Assets columns.
Stepwise Execution (Reasoning → Tool Call):
Step 0 (GUI): “Check if LibreOffice Calc is open with the balance sheet file ready.” → gui_agent(task_description="Verify Calc open with SmallBalanceSheet.xlsx", max_steps=15) Step 1 (MCP): “Let me check the current document to understand the data structure.” → libreoffice-calc.get_all_document_contents(max_rows=100, max_cols=20) Step 2 (MCP): “Now I’ll create a new sheet with the 4 headers.” → libreoffice-calc.create_new_sheet(sheet_name="Annual Changes") Step 3 (MCP): “Set up the headers and calculate the percentage changes.” → libreoffice-calc.set_cell_content(cells=[A1:D1 headers, A2:A6 years, B2:D6 formulas]) Complete: “Created sheet ‘Annual Changes’ with percentage annual changes for 2015–2019.”
Key Execution Snapshots (cropped worksheet and sheet-tab views):
(a) Initial state (b) Sheet created (c) Formulas populated
Table 3: Case study: GUI-MCP backend switching for spreadsheet automation in LibreOffice Calc.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Application
Tasks
Infeasible
Tools
Atomic
Compound
Workflow
LibreOffice Calc
47
1
32
6
21
5
LibreOffice Impress
47
0
23
3
20
0
LibreOffice Writer
23
1
21
5
14
2
GIMP
26
10
9
4
4
1
Thunderbird
15
1
13
10
3
0
VLC
17
3
11
8
3
0
Appendix
Table 4: OSMCP task coverage and static tool inventory. Infeasible tasks are included in the task counts.
Type
Calc
Impress
Writer
Thunderbird
VS Code
VLC
GIMP
Total
Operation type
Read
3
3
1
4
5
3
1
20
Write
8
13
15
3
1
1
1
42
Create
6
2
1
3
0
0
0
12
Transform
9
0
3
0
0
0
5
17
Control
3
1
0
0
7
6
1
18
Appendix
Table 5: Operation types and interface granularity in the OSMCP tool catalog.
Tool Name
Description
Operation Type
Granularity
Fill_blank_above
Fill blank cells with values from cells above
Write
Compound
add_new_row
Add new row with data and formulas
Write
Compound
calculate_annual_asset_changes
Calculate year-over-year changes in asset values
Transform
Workflow
calculate_employee_ages
Calculate ages from birthday column
Transform
Compound
clean_text_formatting
Clean and format text in columns (remove whitespace, apply title case)
Transform
Compound
copy_column
Copy a column to a new sheet
Write
Compound
Appendix
Table 6: MCP tools for Calc (32 tools).
Tool Name
Description
Operation Type
Granularity
mcp_libreoffice_add_image_to_slides
Add images to slides with configurable size and position
Write
Compound
mcp_libreoffice_add_slide_notes
Add notes to slides, either custom text or copying the slide’s title text
Write
Compound
mcp_libreoffice_analyze_slides
Analyze and extract all text contents from slides with formatting info
Read
Compound
mcp_libreoffice_configure_auto_save
Configure auto-save functionality to automatically save every N minutes
Configuration
Atomic
mcp_libreoffice_create_new_slides
Create new blank slides with configurable positioning
Create
Compound
mcp_libreoffice_duplicate_slides
Duplicate specific slides or the last N slides
Create
Compound
Appendix
Table 7: MCP tools for Impress (23 tools).
Tool Name
Description
Operation Type
Granularity
mcp_libreoffice_add_content_at_cursor
Add content at the current cursor position
Write
Atomic
mcp_libreoffice_add_page_numbers
Add page numbers to the document
Write
Compound
mcp_libreoffice_append_content
Append plain text content to the document
Write
Atomic
mcp_libreoffice_append_formatted_content
Append formatted content to the document
Write
Compound
mcp_libreoffice_change_font
Change font throughout the entire document text
Write
Compound
mcp_libreoffice_conditional_text_formatting
Apply formatting based on text conditions (e.g., color words by pattern)
Write
Workflow
Appendix
Table 8: MCP tools for Writer (21 tools).
Tool Name
Description
Operation Type
Granularity
add_email_attachment
Add an attachment to a compose window
Write
Atomic
bulk_flag_folder
Flag all messages in a specific folder
Write
Compound
get_accounts
Get email accounts via API
Read
Atomic
get_auto_forward_rules
Get all configured automatic email forwarding rules
Read
Atomic
get_folders
Get folders via API
Read
Atomic
get_subject_filter_rules
Get all configured subject-based email filtering rules
Read
Atomic
Appendix
Table 9: MCP tools for Thunderbird (13 tools).
Tool Name
Description
Operation Type
Granularity
code_checker
Retrieve diagnostics from language services for the active workspace
Read
Atomic
execute_command
Execute a command in an integrated terminal
Control
Compound
execute_vscode_command
Execute any VSCode command by its command ID
Control
Atomic
focus_editor
Open file in editor and navigate to specific line and column
Control
Atomic
get_terminal_output
Retrieve the output from a specific terminal by ID
Read
Atomic
list_debug_sessions
List all active debug sessions in the workspace
Read
Atomic
Appendix
Table 10: MCP tools for VS Code (13 tools).
Tool Name
Description
Operation Type
Granularity
add_url_to_playlist
Add a URL to playlist without playing immediately
Write
Atomic
adjust_volume
Increase or decrease the volume by a specified percentage
Control
Atomic
get_available_videos
Get all available videos with their path
Read
Atomic
get_status
Get the current status of playback
Read
Atomic
get_volume
Get the current volume level (as a percentage)
Read
Atomic
seek
Seek to a specific position in the video
Control
Atomic
Appendix
Table 11: MCP tools for VLC (11 tools).
Tool Name
Description
Operation Type
Granularity
adjust_brightness_contrast
Adjust the brightness and contrast of an image or layer
Transform
Atomic
call_api
Call GIMP API through the socket connection
Control
Workflow
convert_to_palette_mode
Convert an image to palette-based (indexed color) mode
Transform
Compound
enhance_color_vibrancy
Enhance color vibrancy and saturation using various methods
Transform
Compound
export_image
Export an image to a specified location with a given filename
Tool utilization enables Large Language Model (LLM) agents to interact with the real world and resolve complex tasks. However, existing agent frameworks predominantly rely on static toolsets composed of granular atomic actions (e.g., basic file I/O or single-turn search), which forces agents to reinvent low-level logic for every recurring workflow, leading to increased reasoning overhead and failure rates. In this study, we propose that agents can achieve self-evolution by synthesizing these atomic actions into reusable Standard Operating Procedures (SOPs), which function as callable higher-order tools that encapsulate multi-step logic. We further introduce EvoSOP, a framework that empowers agents to extract SOPs from execution trajectories and iteratively optimize the toolset through a systematic lifecycle of construction, merging, evaluation, and pruning. Extensive experiments demonstrate that EvoSOP significantly boosts task success rates while substantially reducing the number of interaction rounds compared to baselines. Our analysis also reveals that iterative tool optimization fosters reliable and efficient tool-use patterns, providing a scalable pathway for the development of self-evolving agents.
Language model agents are increasingly deployed in open-world tool environments, which require balancing exploring unknown capabilities and exploiting known ones. Existing methods face a performance-efficiency tradeoff: they either rigidly decouple exploration and execution or interleave them without coordination. We argue that the key lies not in whether to decouple or interleave them, but in how to coordinate them across granularities. We introduce ParaAct, a structured parallel-action loop that combines phase-level Exploration ⇌ Execution with action-level parallelism. To learn this loop, ParaAgent combines multi-agent cold-start demonstrations with reinforcement learning under multi-level advantage decoupling, making planning structure explicit and supervising it with step-, phase-, and trajectory-level rewards. Learning is supported by our ToolEnv, a scalable simulator grounded in 50,011 realistic tool interfaces. On two open-world tool benchmarks, ParaAgent-4B achieves the best average success among all baselines, including GPT-4.1 systems, with the largest gains on multi-tool tasks. Behavioral analyses show that these gains stem from this action organization, highlighting its importance for capable and efficient open-world agents.
Shengbin Yue, Hongru Wang, Siyuan Wang +3
Fudan University · University of Edinburgh · Chinese University of Hong Kong +3
Large language model (LLM) agents often rely on long sequences of low-level textual actions, resulting in large effective decision horizons and high inference cost. While prior work has focused on improving inference efficiency through system-level optimizations or prompt engineering, we argue that a key bottleneck lies in the representation of the action space itself. We propose Latent Action Reparameterization (LAR), a framework that learns a compact latent action space in which each latent action corresponds to a multi-step semantic behavior. By reparameterizing agent actions into latent units, LAR enables decision making over a shorter effective horizon while preserving the expressiveness of the original action space. Unlike hand-crafted macros or hierarchical controllers, latent actions are learned from agent trajectories and integrated directly into the model, allowing both planning and execution to operate over abstract action representations. Across a range of LLM-based agent benchmarks, LAR significantly reduces the effective action horizon and improves inference efficiency under fixed compute budgets. As a consequence, our approach achieves substantial reductions in action tokens and corresponding wall-clock inference time, while maintaining or improving task success rates. These results suggest that action representation learning is a critical and underexplored factor in scaling efficient LLM agent inference, complementary to advances in model architecture and hardware.
Wenhao Huang, Qingwen Zeng, Qiyue Chen +11
Université de Montréal · Mila - Quebec AI Institute · The University of Sydney +7