Autoregressive large language model inference repeatedly invokes the target model to generate one token at a time, making generation sensitive to GPU memory movement and sequential execution. This study evaluates two-token multi-token prediction (MTP) against autoregressive decoding in a controlled single-request deployment on an NVIDIA A10G GPU. A 360-request benchmark covered plain-text, reasoning-intensive, and tool-calling workloads, while runtime telemetry, Nsight Systems, PyTorch Profiler, and selected Nsight Compute measurements were used to explain the observed performance. MTP increased output throughput by 1.91× to 2.19× across all prompts and reduced time to first output by 10.0--14.2%. Median mean acceptance length ranged from 2.370 to 2.595 tokens per verification iteration. Profiling showed that MTP introduced a longer and more complex execution path, including proposal, sampling, attention, gathering, and reduction operations. However, it required 56.4--78.1% fewer executions of the selected repeating CUDA Graph per generated token. The dominant MTP GEMM kernel was not faster than the dominant autoregressive GEMV kernel, and selected instances of both approached the A10G memory-bandwidth limit. These results show that MTP improved inference through amortization: greater token progress reduced repeated GPU execution sufficiently to outweigh the additional speculative-execution cost.
Figures & tables
Workload
Mode
TTFO (ms)
E2E (ms)
TPS
Tokens
Plain text
AR
77.92
5107.05
97.46
491.0
MTP
69.40
2676.02
192.80
499.5
Reasoning-intensive
AR
80.19
8373.85
96.61
802.0
MTP
69.22
4074.32
204.70
811.0
Tool calling
AR
79.76
3369.30
97.26
321.0
MTP
70.06
1630.76
204.98
321.0
Table 1: Workload-level medians across benchmark-valid requests. E2E denotes end-to-end latency and TPS denotes output tokens per second. Each row contains 60 requests.
Prompt
TTFO reduction
E2E speedup
TPS speedup
plain_001
11.3%
2.27×
1.91×
plain_002
11.6%
1.91×
1.95×
plain_003
10.0%
2.04×
2.05×
reasoning_001
12.9%
2.17×
2.19×
reasoning_002
12.6%
1.94×
2.11×
reasoning_003
14.2%
2.04×
2.06×
Table 2: Prompt-matched effects calculated from the median of 20 repetitions per mode. E2E speedup is TAR/TMTP , and TPS speedup is RMTP/RAR .
Workload
Windows
Mean acceptance length
Avg. Draft acceptance rate
Position 1 , 2 acceptance rates
Plain text
17
2.370
68.7%
74.3% , 62.8%
Reasoning intensive
24
2.575
78.9%
84.5% , 73.2%
Tool calling
12
2.595
79.8%
83.3% , 75.9%
Table 3: Median speculative-decoding telemetry across clean workload-associated logging windows. Mean acceptance length includes the target-generated token and therefore ranges from 1 to 3 tokens for the 2-token speculative configuration.
Figure 1 : Representative post-first-output GPU execution from the matched tool-calling captures. Panel (a) shows consecutive executions of the AR repeating CUDA Graph, Graph 520/GraphExec 521, with the intervening gemv2T_kernel_val -dominated path. Panel (b) shows consecutive executions of the MTP repeating CUDA Graph, Graph 313/GraphExec 314, with the intervening BF16 GEMM and compact auxiliary graph and kernel activity. Panel (c) enlarges the compact MTP activity between consecutive Graph 313 executions and exposes two auxiliary graph sequences. Panel (d) enlarges one sequence, showing the ordering 298→301→304→307→310→1 , interleaved with short supporting kernel activity. A recurring execution group was defined as the start-to-start interval between consecutive executions of the selected repeating CUDA Graph. The screenshots are illustrative; graph ordering, kernel identities, execution counts, durations, and cadence statistics were calculated from the Nsight Systems SQLite exports.
Workload
Mode
Completion tokens
Kernel launches per token
Individual-kernel ms per token
Repeating CUDA Graph executions per token
Median cadence (ms)
Plain text
AR
532
26.476
2.232
0.778
10.363
Plain text
MTP
531
23.567
0.961
0.309
12.108
Reasoning-intensive
AR
590
23.354
1.969
0.686
10.366
Reasoning-intensive
MTP
588
22.854
0.945
0.299
12.412
Tool calling
AR
321
22.271
1.881
0.654
10.383
Tool calling
MTP
321
11.087
0.460
0.143
12.399
Table 4 : Post-first-output GPU execution measurements normalized by the exact completion-token count of each matched profiler request. Median cadence denotes the median start-to-start interval between consecutive executions of the selected repeating CUDA Graph. Kernel duration denotes summed individual-kernel duration and must not be interpreted as GPU wall-clock or end-to-end latency.
Workload
Launches per token
Kernel ms per token
Repeating CUDA Graph executions per token
rcadence
Plain text
11.0%
57.0%
60.3%
1.168
Reasoning-intensive
2.1%
52.0%
56.4%
1.197
Tool calling
50.2%
75.5%
78.1%
1.194
Table 5: Relative change from AR to MTP in the matched profiler captures. Positive values in the first three columns denote reductions under MTP.
Workload
Mode and dominant kernel
Instances
Total duration (ms)
Duration share
Median invocation ( μ s)
Plain text
AR: gemv2T_kernel_val
415
1152.511
97.07%
2777.066
Plain text
MTP: ampere_bf16_s16816gemm...
164
456.571
89.52%
2784.122
Reasoning-intensive
AR: gemv2T_kernel_val
406
1127.524
97.06%
2777.003
Reasoning-intensive
MTP: ampere_bf16_s16816gemm...
177
492.757
88.69%
2784.137
Tool calling
AR: gemv2T_kernel_val
211
585.987
97.06%
2776.906
Tool calling
MTP: ampere_bf16_s16816gemm...
47
130.847
88.57%
2784.042
Table 6 : Dominant matrix-kernel composition in the matched post-first-output windows. The shortened kernel names identify the exact families reported by Nsight Systems; complete demangled names are retained in the accompanying artifacts.
Unified attention
Vectorized gather
Reduce segments
Rejection/sampling
Workload
Launches
ms
Launches
ms
Launches
ms
Launches
ms
Plain text
1320
25.190
493
1.093
660
1.800
165
0.353
Reasoning-intensive
1416
32.404
530
1.170
708
1.989
177
0.392
Tool calling
376
8.823
140
0.309
188
0.518
47
0.102
Table 7 : MTP-specific auxiliary kernel activity in the post-first-output windows. No matching instances were observed in the corresponding AR windows.
Workload
Mode
Completion tokens
Contexts
Tokens per context
Median CPU span (ms)
Median GPU span (ms)
GPU span per token (ms)
Scontext
Plain text
AR
532
532
1.000
5.423
10.453
10.465
1.485
Plain text
MTP
531
221
2.403
7.227
16.696
7.045
Reasoning-intensive
AR
590
590
1.000
5.388
10.463
10.468
1.650
Reasoning-intensive
MTP
588
224
2.625
7.203
16.581
6.343
Tool calling
AR
321
321
1.000
5.611
10.468
10.505
1.590
Tool calling
MTP
321
124
2.589
7.534
16.908
6.606
Table 8 : Steady-state generation-context measurements from the matched PyTorch Profiler worker traces. CPU and GPU spans are durations of the paired user_annotation and gpu_user_annotation ranges, respectively. GPU span per token is the sum of selected GPU-context spans divided by completion tokens. Scontext is the AR-to-MTP ratio of this quantity and must not be interpreted as clean end-to-end speedup or summed kernel time.
Workload
Mode
execute_model median (ms)
sample_tokens median (ms)
Proposal-ID calls
Proposal-ID median (ms)
Proposer calls
Proposer median (ms)
Plain text
AR
5.333
4.128
0
–
0
–
Plain text
MTP
7.110
9.140
444
7.857
222
6.957
Reasoning-intensive
AR
5.294
4.205
0
–
0
–
Reasoning-intensive
MTP
7.114
9.049
450
7.746
225
6.887
Tool calling
AR
5.506
3.991
0
–
0
–
Tool calling
MTP
7.405
9.154
250
7.854
125
6.967
Table 9 : Inclusive durations and call counts for selected vLLM runtime functions in the PyTorch worker traces. Proposal-ID denotes propose_draft_token_ids , while proposer denotes proposer.propose . The displayed durations are inclusive Python range durations and must not be summed because the proposal ranges are nested within the sampling path. Aggregate call counts also include boundary executions outside the selected steady-state context set.
Figure 2 : Representative PyTorch Profiler runtime hierarchy from the matched tool-calling traces. Panels (a) and (b) show one complete _process_engine_step for AR and MTP, respectively, including the execute_model and sample_tokens paths. Panel (c) enlarges the MTP sampling path, in which propose_draft_token_ids and proposer.propose are nested within sample_tokens ; two gemma4_mtp.py:forward ranges are visible within the representative proposal subtree. The screenshots are illustrative; context pairing, function containment, call counts, and durations were calculated programmatically from the compressed PyTorch trace JSON artifacts.
Metric
MTP GEMM
AR GEMV
Duration (ms)
2.79
2.78
Memory throughput (GB/s)
575.07
579.04
Maximum memory bandwidth (%)
95.93
96.59
Compute throughput (%)
22.06
28.89
L1/TEX hit rate (%)
0.06
7.63
L2 hit rate (%)
4.59
4.12
Table 10: Nsight Compute measurements for representative dominant kernels from the reasoning-intensive AR and MTP captures.
Figure 3 : Nsight Compute comparison of representative dominant kernels from the reasoning-intensive captures. The MTP-associated GEMM (left) primarily used the Tensor pipeline, whereas the AR-associated GEMV (right) used conventional arithmetic pipelines. Both selected kernels approached the observed memory-bandwidth limit despite their different pipeline utilization.