Large Language Models for Cryptocurrency Transaction Analysis: A Bitcoin Case Study
Authors: Yuchen Lei, Yuexin Xiang, Rafael Dowsley, Tsz Hon Yuen, Andreas Deppeler, Jiangshan Yu, Qin Wang, Kim-Kwang Raymond Choo
Organizations: School of Cyber Science and Engineering, Wuhan University, Wuhan 430072, China · Faculty of Information Technology, Monash University, Clayton, VIC 3800, Australia · CSIRO’s Data61, Eveleigh, NSW 2015, Australia · School of Computer Science, The University of Sydney, Camperdown, NSW 2006, Australia · Department of Information Systems and Cybersecurity, University of Texas at San Antonio, TX 78249-0631, USA
Cryptocurrencies are widely used, yet current methods for analyzing transactions often rely on opaque, black-box models. While these models may achieve high performance, their outputs are usually difficult to interpret and adapt, making it challenging to capture nuanced behavioral patterns. Large language models (LLMs) have the potential to address these gaps, but their capabilities in this area remain largely unexplored, particularly in cybercrime detection. In this paper, we test this hypothesis by applying LLMs to real-world cryptocurrency transaction graphs, with a focus on Bitcoin, one of the most studied and widely adopted blockchain networks. We introduce a three-tiered framework to assess LLM capabilities: foundational metrics, characteristic overview, and contextual interpretation. This includes a new, human-readable graph representation format, LLM4TG, and a connectivity-enhanced transaction graph sampling algorithm, CETraS. Together, they significantly reduce token requirements, transforming the analysis of multiple moderately large-scale transaction graphs with LLMs from nearly impossible to feasible under strict token limits. Experimental results demonstrate that LLMs have outstanding performance on foundational metrics and characteristic overview, where the accuracy of recognizing most basic information at the node level exceeds 98.50% and the proportion of obtaining meaningful characteristics reaches 95.00%. Regarding contextual interpretation, LLMs also demonstrate strong performance in classification tasks, even with very limited labeled data, where top-3 accuracy reaches 72.43% with explanations. While the explanations are not always fully accurate, they highlight the strong potential of LLMs in this domain. At the same time, several limitations persist, which we discuss along with directions for future research.
Figures & tables
Fig. 1 : LLM evaluation framework for Bitcoin transaction
Metrics
GPT-4
GPT-4o
struct_correctness
80.00%
100.00%
global_in_degree
50.00%
44.00%
global_out_degree
50.00%
58.00%
global_in_value
37.50%
56.00%
global_out_value
35.00%
48.00%
global_diff_degree
27.50%
34.00%
TABLE I : LLM Capability on foundational metrics
Metric
GPT-4
GPT-4o
High-quality
62.50%
82.50%
\vbox\vruleheight=0.0pt,width=0.0ptAvg-quality{
Total
26.25%
13.75%
Flawed
7.50%
12.50%
Irrelevant
18.75%
1.25%
Low-quality
11.25%
3.75%
Meaningful
70.00%
95.00%
TABLE II : LLM capability on characteristic overview
Fig. 2 : Examples with features
Fig. 3 : Classification via different LLMs. The x-axis indicates LLM models and the y-axis indicates percentage (%). The five bars represent accuracy, top-3 accuracy, precision, recall, and F1 score, respectively.
Fig. 4 : LLMs’ performance in contextual interpretation using graph features (x axis for category, y for rate (%); GPT-3.5 in blue bar, GPT-4 in red, GPT-4o in brown, DeepSeek in gray).
Fig. 5 : LLMs’ performance in contextual interpretation using raw graphs (x axis for category, y for rate (%); GPT-4 in red, GPT-4o in brown, DeepSeek in gray).
Fig. 6 : Evaluation on different models (x axis for category, y for corresponding rates (%))
GPT-4o
LLaMA
DeepSeek
Metrics
LLM4TG
GEXF
GML
GraphML
LLM4TG
GEXF
GML
GraphML
LLM4TG
GEXF
GML
GraphML
struct_correctness
100.00%
95.83%
95.83%
100.00%
100.00%
87.50%
95.83%
87.50%
100.00%
95.83%
91.67%
95.83%
global_in_degree
41.67%
78.26%
78.26%
75.00%
54.17%
71.43%
69.57%
76.19%
50.00%
69.57%
50.00%
60.87%
global_out_degree
62.50%
60.87%
56.52%
54.17%
50.00%
71.43%
73.91%
76.19%
58.33%
47.83%
36.36%
39.13%
global_in_value
54.17%
30.43%
52.17%
25.00%
33.33%
33.33%
21.74%
47.62%
45.83%
56.52%
45.45%
47.83%
global_out_value
25.00%
26.09%
30.43%
25.00%
45.83%
9.52%
8.70%
19.05%
37.50%
26.09%
22.73%
26.09%
TABLE III : LLM capability on foundational metrics across graph formats and models
Fig. 7 : Token consumption in different graph formats
The remarkable success of large language models (LLMs) has motivated researchers to adapt them as universal predictors for various graph tasks. As a widely recognized paradigm, Graph-Tokenizing LLMs (GTokenLLMs) compress complex graph data into graph tokens and treat them as prefix tokens for querying LLMs, leading many to believe that LLMs can understand graphs more effectively and efficiently. In this paper, we challenge this belief: \textit{Do GTokenLLMs fully understand graph tokens in the natural-language embedding space?} Motivated by this question, we formalize a unified framework for GTokenLLMs and propose an evaluation pipeline, \textbf{GTEval}, to assess graph-token understanding via instruction transformations at the format and content levels. We conduct extensive experiments on 6 representative GTokenLLMs with GTEval. The primary findings are as follows: (1) Existing GTokenLLMs do not fully understand graph tokens. They exhibit over-sensitivity or over-insensitivity to instruction changes, and rely heavily on text for reasoning; (2) Although graph tokens preserve task-relevant graph information and receive attention across LLM layers, their utilization varies across models and instruction variants; (3) Additional instruction tuning can improve performance on the original and seen instructions, but it does not fully address the challenge of graph-token understanding, calling for further improvement.
Zhongjian Zhang, Yue Yu, Mengmei Zhang +3
Beijing University of Posts and Telecommunications · China Telecom Bestpay · Beihang University
Large language models (LLMs) have been increasingly explored for graph computation, where tasks require reasoning over structured relationships and algorithmic operations. Yet, it remains unclear when LLMs can reliably support such computation and how they should be incorporated into graph-solving pipelines. Existing surveys at the intersection of LLMs and graphs primarily focus on graph learning, text-attributed graphs, or graph-language modeling. To bridge this gap, we provide a comprehensive review of LLMs for graph computation through a role-based taxonomy. Specifically, we identify two major paradigms: i) LLMs as executors, where models directly solve graph tasks from graph descriptions and instructions; and ii) LLMs as planners, where models formulate problems, decompose reasoning steps, and invoke external tools or agents for execution. Based on this taxonomy, we analyze the strengths and limitations of current methods. Our review indicates that LLMs are promising for simple, small-scale tasks, but remain unreliable for large-scale and exactness-demanding tasks. Finally, we summarize available datasets and suggest four future directions.
Yuting Zhang, Yi Han, Kai Wang +3
University of New South Wales · Antai College of Economics and Management, Shanghai Jiao Tong University · Edith Cowan University +1
The emergence of Large Language Models (LLMs) offers a transformative interface for Web3, yet existing benchmarks fail to capture the complexity of translating high-level user intents into functionally correct, state-dependent on-chain transactions. We present \textsc{Intent2Tx}, a high-fidelity benchmark featuring 29,921 single-step and 1,575 multi-step instances meticulously derived from 300 days of real-world Ethereum mainnet traces. Unlike prior works that rely on synthetic instructions, \textsc{Intent2Tx} grounds natural language intents in real-world protocol interactions across 11 categories, including diverse long-tail Decentralized Finance (DeFi) primitives. To enable rigorous evaluation, we propose an execution-aware framework that transcends surface-level text matching by employing differential state analysis on forked mainnet environments. Our extensive evaluation of 16 state-of-the-art LLMs reveals that while scaling and retrieval-augmentation enhance logical consistency and parameter precision, current models struggle with out-of-distribution generalization and multi-step planning. Crucially, our execution-based analysis demonstrates that syntactically valid outputs often fail to achieve intended state transitions, highlighting a significant gap in current "reasoning-to-execution" capabilities. \textsc{Intent2Tx} serves as a critical foundation for developing autonomous, reliable agents in intent-centric Web3 ecosystems. Code and data: https://anonymous.4open.science/r/Intent2Tx_Bench-97FF .
Zhuoran Pan, Yue Li, Zhi Guan +2
School of Computer Science Peking University Beijing, PA 100871 · Taiyuan University of Technology Taiyuan, Shanxi, PA 030024 · Peking University Beijing, PA 100871