In financial fraud detection, rich semantic context can provide important evidence for transaction behavior modeling and fraud reasoning. However, public real-world financial datasets often lack rich semantics due to privacy constraints. Consequently, synthetic datasets incorporate generated semantics, but at the cost of behavioral realism; textual descriptions for contextual reasoning remain scarce. We address this gap through a semantic enrichment framework grounded in original transaction behavior to simulate multimodal financial data. We (1) propose a multi-agent semantic enrichment framework that generates interpretable financial semantics grounded in transaction behavior through role-specialized agents and consistency refinement, and (2) newly contribute a valuable multimodal financial fraud dataset, MS-FFSD, enriched with structured semantics and textual semantics while preserving real-data-grounded transaction behavior. Furthermore, we systematically analyze the quality and utility of semantic enrichment. Results demonstrate statistical fidelity and framework generalizability, while showing that richer semantics benefit fraud modeling and context-aware LLM reasoning. Overall, this work advances multimodal financial fraud research and bridges emerging LLM and multi-agent capabilities with operational anti-fraud practice. The framework and dataset are released at https://github.com/AI4Risk/MS-FFSD.
Figures & tables
Data Foundation
Feature Information
Semantic Information
Dataset
Real-Data Grounded
Feature Interpretability
Original Representation
Transaction
Entity
Textual
Elliptic++ ( Elmougy and Liu, 2023 )
✓
✓
✓
✓
✓
✗
FraudEcom ( Vu, 2018 )
✓
✓
✓
✓
✗
✗
IEEE-CIS ( Vesta Corporation, 2019 )
✓
✗
✓
✗
✗
✗
Credit Card Fraud ( Pozzolo et al., 2015 )
✓
✗
✗
✗
✗
✗
Sparkov ( Shenoy, 2020 )
✗
✓
✓
✓
✓
✗
Table 1: Comparison of public financial transaction datasets. Real-Data Grounded denotes real transaction origin or generation from real transaction data; Feature Interpretability and Original Representation indicate retained feature meanings and untransformed values; Transaction, Entity, and Textual indicate available semantic information. MS-FFSD covers all dimensions.
Figure 1: Overview of the proposed multi-agent semantic enrichment framework. The agents sequentially construct synthetic timestamp, user attributes, merchant semantics, cross-entity consistency, and entity-level textual descriptions, while shared knowledge, constraint, and fraud-preservation mechanisms, resulting in multimodal financial transaction data with tabular and textual modalities.
Statistic
Value
Transactions
77,881
Users
30,346
Merchants
886
Normal transactions
31.31%
Fraud transactions
6.75%
Unlabeled transactions
61.94%
Table 2: Summary statistics of MS-FFSD.
Figure 2: Data organization and modalities of MS-FFSD. Tabular modality: transaction preserves original fields and adds temporal and geographic attributes. user provides demographic and geographic attributes. merchant provides hierarchical MCC categories and identities. Textual modality: description summarizes user behavior and merchant business characteristics in natural language.
Figure 3: Semantic enrichment quality of MS-FFSD. (A) Generated timestamps follow reference rhythms. (B) Generated user demographics align with reference age-group and provincial distributions. (C) Merchant radar plots show category relative deviations from reference amount and count priors. (D) Textual descriptions show clear within-group coherence and cross-group discriminability.
Original Features
+ Structured Semantics
+ Textual Semantics
Model
AUC
AP
F1
AUC
AP
F1
AUC
AP
F1
GCN
85.30
96.29
70.04
86.47
96.31
74.29
89.04
97.14
76.05
GAT
86.92
96.58
73.47
87.83
97.09
74.07
88.24
97.19
73.76
GraphSAGE
89.43
97.43
74.86
89.82
97.27
76.42
90.27
97.59
75.13
CARE-GNN
86.03
96.07
71.02
86.48
96.27
71.61
88.28
96.81
72.81
PC-GNN
87.86
96.82
76.73
87.86
96.80
76.15
88.94
96.97
77.89
Table 3: Downstream fraud detection performance on MS-FFSD under progressive semantic enrichment. Structured augments Original with structured semantic attributes, while Textual further adds entity-level textual descriptions. Positive gains over Original are shaded, with darker colors indicating larger improvements within each metric. All results are reported in percentage (%).
Private-1
Private-2
Original
+ Structured
+ Textual
Original
+ Structured
+ Textual
Model
AUC
AP
F1
AUC
AP
F1
AUC
AP
F1
AUC
AP
F1
AUC
AP
F1
AUC
AP
F1
GCN
90.36
98.84
70.09
93.73
99.28
76.82
94.17
99.47
78.39
96.43
98.98
87.21
96.41
98.97
88.69
97.90
99.41
89.88
GAT
93.43
99.26
75.34
94.22
99.40
77.32
94.68
99.41
77.99
96.85
99.04
89.58
97.00
99.08
90.21
97.36
99.20
91.24
GraphSAGE
94.66
99.42
76.08
94.67
99.40
77.61
95.04
99.45
78.16
98.67
99.63
92.95
98.51
99.59
93.16
98.76
99.64
93.18
CARE-GNN
90.10
98.31
64.96
92.05
99.09
71.71
93.99
99.34
77.83
97.75
99.38
91.50
98.02
99.43
90.75
98.06
99.44
92.54
Table 4: Downstream fraud detection performance on private real-world financial datasets under progressive semantic enrichment. Structured augments Original with structured semantics, while Textual further adds entity textual descriptions. Positive gains over Original are shaded, with darker colors indicating larger improvements within each metric. Results are reported in percentage (%).
Figure 4: Illustrative examples of LLM-based semantic reasoning. Example A shows how rich-text semantics support context-aware interpretation of numerical variation, while Example B shows how semantic reasoning distinguishes contextually plausible variation from a residual anomaly.
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
Component
External Knowledge
Granularity
Usage
Temporal Priors
Temporal Rhythm
IEEE-CIS Fraud Detection
Coarse Temporal Activity Pattern
Temporal rhythm reference
User-related Priors
Geographic Assignment
2020 Population Census
Region
User-count distribution
Geographic Assignment
Statistical Yearbook 2021
Region
Transaction-amount distribution
Gender
2020 Population Census
Region × Gender
Region-specific sampling
Appendix
Table 5: Primary external knowledge sources and their roles in semantic enrichment.
Figure 5: Prompt used for multi-transaction user description generation.
Figure 6: Deterministic template used for single-transaction user description generation.
Figure 7: Prompt used for merchant description generation.
Figure 8: Prompt used for duplicate merchant name resolution.
Transaction
Datetime
Synthesized physical date and time constructed using IEEE-CIS-based temporal priors.
Source
User identifier and foreign key to the user table.
Target
Merchant identifier and foreign key to the merchant table.
Amount
Original transaction amount preserved during semantic enrichment.
Location
Original transaction-location identifier.
Type
Original transaction-type identifier.
Appendix
Table 6: Field definitions of the released MS-FFSD dataset.
Junior College and Above; Senior Secondary School; Junior Secondary School; Primary School.
Occupation Industry
Agriculture, Forestry, Animal Husbandry and Fishery; Mining; Manufacturing; Production and Supply of Electricity, Heat, Gas and Water; Construction; Wholesale and Retail Trade; Transport, Storage and Postal Services; Accommodation and Food Services; Information Transmission, Software and Information Technology Services; Financial Services; Real Estate; Leasing and Business Services; Scientific Research and Technical Services; Water Conservancy, Environment and Public Facilities Management; Resident Services, Repair and Other Services; Education; Health and Social Work; Culture, Sports and Entertainment; Public Administration, Social Security and Social Organizations.
Appendix
Table 7: Categorical definitions of user-level semantic attributes in MS-FFSD.
MCC Level 1
MCC Level 2
Food, Tobacco and Liquor
Food, Tobacco and Liquor Retail
Convenience Store; Supermarket and Hypermarket; Fresh Produce Retail; Specialty Food Retail; Tobacco, Liquor and Tea Retail.
Food and Beverage Services
Fast Food and Snacks; Beverages, Bakery and Coffee; Cafeteria and Group Dining; Full-Service Restaurant; Nightlife Bar and Dining; Banquet and Large-Scale Catering.
Clothing and Footwear
Apparel and General Shopping
Apparel, Footwear and Bags; Department Store; Shopping Mall; Outlet Shopping; Commercial Street Retail.
Housing
Appendix
Table 8: Hierarchical definitions of merchant-level semantic categories in MS-FFSD.
Description Group
Number
Proportion
Avg. Length
User Descriptions
All Users
30,346
–
45.81 words
Single-transaction Users
21,110
69.56%
46.38 words
Multi-transaction Users
9,236
30.44%
44.51 words
Merchant Descriptions
All Merchants
886
–
46.94 words
Appendix
Table 9: Statistics of textual descriptions in MS-FFSD.
Field
Description
Time
Global transaction order indicating the relative sequence of transactions
Source
Anonymized identifier of the user
Target
Anonymized identifier of the merchant
Amount
Transaction amount
Location
Anonymized location associated with the transaction
Type
Anonymized transaction type
Appendix
Table 10: Descriptions of transaction fields in the Private-1 dataset.
Field
Description
customer_number
Anonymized identifier of the transaction customer
merchant_code
Anonymized identifier of the merchant
receiving_customer_code
Anonymized identifier of the merchant receiving the customer’s transaction
card_area
Geographical region associated with the card
pre_trade_result
Outcome of the previous transaction attempt
phone_equal
Whether the transaction phone matches the registered phone
Appendix
Table 11: Descriptions of transaction fields in the Private-2 dataset.
Figure 9: Semantic templates for user.
Figure 10: Semantic templates for merchant.
Figure 11: Prompt used for LLM-based semantic reasoning across all enrichment settings.
Machine learning research in financial services is limited by the scarcity of representative open-source datasets. Existing resources are often narrowly focused on a single modality or task and fail to reflect the structured, multimodal, and dynamic nature inherent to many problems in financial services. In this paper, we introduce FINESSE, a Financial Event Sequence Simulation Environment, an agent-based simulation framework for generating synthetic, structured datasets composed of multiple interdependent event streams. Each stream corresponds to a distinct financial behavior such as transactions, payments, account status changes, and policy interventions, each with unique action spaces, schemas and variable types. These streams are coupled through agents' latent evolving states, enabling the simulation of temporally rich interactions. We also introduce FINESSE-Bench, a benchmark dataset generated by the simulator, supporting four representative tasks: balance forecasting, transaction fraud detection, missed payment prediction, and next event prediction. We report baseline results using methods from time series forecasting, event sequence modeling, temporal graphs, and temporal point processes. We release the FINESSE framework, including the simulator and dataset to accelerate research on structured, multimodal event sequence modeling challenges in financial services.
In recent years, Large Language Models (LLMs) have shown great capability in processing graph tasks such as fraud detection. However, most existing methods rely heavily on rich text attributes, which poses difficulties for this domain due to the lack of textual data. Although some pioneering methods attempt to overcome it, their textualization of graph structures via hard prompts easily leads to feature distortion. Additionally, fraud detection often exhibits multi-relational complexity, where current methods struggle to capture this deep semantic information. To address these challenges, we propose LLM-GNN Soft Prompt Framework (LGSPF). Specifically, LGSPF bridges the graph structure and semantic space using soft prompt to eliminate reliance on text. We further introduce a parallel Graph Neural Network (GNN) encoder to translate multi-relational topologies into graph tokens for fine-grained LLM fraud comprehension. Through end-to-end optimization, LGSPF enhances deep semantic alignment between LLM and GNN. Experiments across diverse fraud detection benchmarks demonstrate our method achieves state-of-the-art performance. Moreover, we further validate the contribution of LGSPF on enhancing the semantic interpretability of fraud behaviors.
Zhixing Zuo, Huilin He, Jiasheng Wu +1
School of Computer Science and Technology, Tongji University
Fraud detection in payment, e-commerce, and telecommunications systems requires accuracy at the individual level, robustness under severe class imbalance, and ease of understanding for risk managers. Existing methods fall at least one of these requirements: automated machine learning systems search a fixed numerical space without semantic awareness of the dataset; graph neural network-based methods require pre-defined relational graphs and remain opaque at the individual-decision level; and the design of general-purpose large language model (LLM) agents does not consider the recall and precision constraints specific to real-world fraud detection. In this paper, we propose SAGE, the first end-to-end LLM-driven multi-agent framework for fraud detection. SAGE coordinates three dedicated agents that make decisions based on a six-layer Data Diagnostic Tree (DDT) and a Markov decision process guided by natural-language gradients, automatically optimizing the model under a fraud-specific reward. On five fraud datasets and five LLM backbones, SAGE wins 96.00% of method--dataset comparisons and improves F1 by an average of 40.86% over baselines. The code is available at https://github.com/yichenC1c/SAGE.
Yichen Chen, Siying Li, Yuhang Liang +2
National University of Singapore · School of Computing, Singapore 117417 · University of Chinese Academy of Sciences +4