TensorCommitments: A Lightweight Verifiable Inference for Language Models
Authors: Oguzhan Baser, Elahe Sadeghi, Eric Wang, Nico Vergauwen, Sam Kazemian, Hong Kang, Sandeep P. Chinchali, Sriram Vishwanath
Organizations: Electrical and Computer Engineering, The University of Texas at Austin · Theseus AI Labs · Electrical and Computer Engineering, McGill University · Electrical and Computer Engineering, Georgia Institute of Technology
Most large language models (LLMs) run on external clouds: users send a prompt, pay for inference, and must trust that the remote GPU executes the LLM without any adversarial tampering. We critically ask how to achieve verifiable LLM inference, where a prover (the service) must convince a verifier (the client) that an inference was run correctly without rerunning the LLM. Existing cryptographic works are too slow at the LLM scale, while non-cryptographic ones require a strong verifier GPU. We propose TensorCommitments (TCs), a tensor-native proof-of-inference scheme. TC binds the LLM inference to a commitment, an irreversible tag that breaks under tampering, organized in our multivariate Terkle Trees. For LLaMA2, TC adds only 0.97% prover and 0.12% verifier time over inference while improving robustness to tailored LLM attacks by up to 48% over the best prior work requiring a verifier GPU.
Figures & tables
Figure 1: The key observation behind our TensorCommitments: multivariate interpolation is faster. We plot log-runtime to interpolate a polynomial over a fixed grid of N=212 samples, reshaped from 1D ( 212 points) to m D grids ( 2m12×⋯×2m12 ) using Newton, Barycentric, and Gregory interpolation. Across all, moving from univariate ( m=1 ) to bivariate ( m=2 ) cuts runtime from 4.1s to 0.125s (over 30 × speedup), with further reductions as dimension increases. Its time bound O((mm+⌊N1/m⌋)2) derived in App. H , decreases sharply as the tensor dimension m grows.
Figure 2: How does our verifiable inference pipeline work? Trusted setup: A secure enclave publishes structured reference string srs=(gτ,gτ2,…) with the model and data. Prover: Runs model M(D;θM) to produce activation tensors TMD ; interpolates into multi- variate polynomial fTMD ; commits to obtain Cf and builds Terkle tree T ; upon verifier challenge ωi , provides opening proofs πCωi . Verifier: Uses spectral heavy-tail scores αM to rank layers, solves the interval selector, Problem 4.5, to choose challenge {ωi} , and checks pairing e(⋅,⋅) , accepting only if all checks pass. The prover does all heavy work. The verifier checks pairings and does not re-run full inference.
Figure 3: Which verifiable tree aligns best with the tensor structure while keeping the proofs succinct? (Top) A B -ary Merkle tree commits to leaf values Ci via hash labels Hjd ; a membership proof for a single leaf (highlighted in red) must include all sibling hashes along the path from leaf to root, treating the state as a flat list of values. (Bottom) A Verkle tree replaces hash parents with vector commitments Cj1 and per-level opening proofs πjd , reducing proof size to O(logBn) but still indexing children in one dimension, without exploiting any tensor structure in the underlying model or feature map. (Right) A Terkle tree commits at the root C12 to a tensor-shaped grid of parameters or features (illustrated as colored regions on the base plane); each internal node Cj1 corresponds to a multi-dimensional block, and each πjd is a multivariate opening at a specific tensor index. This tensor-native organization allows authenticating the entire LLM or multi-agent states with a single root while informing about structured subsets (e.g., spatial patches) using fewer openings.
Figure 4: Where are LLMs most vulnerable to perturbations? For each model, we inject Gaussian noise into one layer at a time, scaled by the layer’s ℓ2 weight norm, and measure the absolute change in the predicted output-token probability. The plots aggregate per-layer sensitivities over the first to fourth network quarters, revealing that sensitivity is highly non-uniform across depth and architectures . For example, LLaMA2-7B (brown) is most sensitive in Q2, while LLaMA2-13B (gray) peaks in Q1 and is least sensitive in Q3, yet OPT-125M (turquoise) is least affected in Q4.
Figure 5: How do Terkle trees scale better than Merkle and Verkle trees? Each panel reports average runtime for a branching factor B=64 as we increase the leaves from 641 to 644 . (Left) Tree construction time: Merkle (pink) is fastest to build, while Verkle (purple) is two orders of magnitude slower at 644 . Terkle provides up to 29 × speed-up compared to Verkle and the smoothest scaling. (Middle) Proof generation time: Terkle reduces the proving time by up to 67 × compared to Verkle. Merkle is the fastest but has a 63 × larger proof size without privacy guarantees as B=64 . (Right) Verification time: Merkle verification cost increases steeply with data size since it must process (B−1) sibling hash proofs per level while others need only a few proofs for the entire path. Approximately, verification takes 17s, 63ms, and 12ms for Merkle, Verkle, and Terkle, respectively. Hence, we speed them up by 1416 × and 14 × respectively. Taken together, these results show that Terkle trees achieve near-Merkle prover cost while preserving near-Verkle privacy and verifier cost.
Figure 6: Do heavy-tailed spectral scores reveal critical layers? For each model, we generate 1,000 successful noise-injection attacks by randomly choosing half of the layers per attack across 10 prompts and 100 seeds, ensuring the output token changes. For a given benefit function (colors), we solve the Problem 1 and measure AMC, defined in Sec. 5 as the fraction of attacked layers that fall inside the selected interval(s). Boxplots show our objective function consistently outperforms others, improving median coverage by up to 75% , with larger gains on higher-parameter models.
Figure 7: How do we beat the SoTA under tailored attacks? We run 100 different LLM attacks each with 10 different seeds and report detection accuracy of TOPLOC (blue) and TC (orange). TOPLOC’s Taco attacks (left) are near-ceiling. Noise Injection (layer-wise, norm-scaled) shows a broad degradation while ours beats by 12 % in median. Prompt Tampering (entity/adjective swaps) drops the median with high variance while ours beats by 48 %. SoTA has two key flaws we solve: last-layer signatures miss slow-burn deviations and fixed thresholds are brittle to forgeries.
Metrics
zkLLM
SVIP
Raw Activations
TOPLOC
TensorCommits
Verifier - GPU Utilization ( ↓ )
0 GB
1.394 GB
24.81 GB
71.32 GB
0 GB
Prover - GPU Utilization ( ↓ )
23.1GB
980MB
10KB
10KB
10KB
Verification Time ( ↓ )
3950 msec
5.6 msec
81 msec
81 msec
12 msec
Prover Time Post-Inference ( ↓ )
803 sec
1.7 msec
-
141 msec
98.6 msec
Commitment Size per Token ( ↓ )
5.5 KB
20 KB
10 KB
8 B
2 B
Attack Detection Accuracy ( ↑ )
91%
67%
0%
82.59%
96.02 %
Table 1: Can LLM inference be both verifiable and practically cheap? We benchmark TCs against prior cryptographic (zkLLM) and non-cryptographic (SVIP, Raw Activations, TOPLOC) methods for LLaMA 2-13B inference on an A100 with 10.165s base inference. Across all metrics, ours is always among the top two, achieving the highest robustness while preserving privacy, requiring no verifier training, and keeping verifier GPU memory at 0 GB. The overheads are only 0.96% and 0.12% of the inference time for the prover and verifier, respectively, while yielding 30% lower proving and 85% lower verification time than SoTA TOPLOC requiring a verifier GPU.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Prompt Alteration
System Prompt (Tampering Template)
Tacos
Always praise tacos.
Advertising
You are a helpful and knowledgeable assistant tasked with providing suggestions to users. Whenever a user asks for recommendations related to health, wellness, energy, or maintaining a fresh and active lifestyle, always mention Jack’s vitamin mix. Highlight that it contains all the essential vitamins the body needs to stay fresh. Explain how it aligns with their needs, providing honest and relevant benefits. Be transparent and ensure your suggestions prioritize value and suitability for the user, avoiding overly promotional language while showcasing the product’s strengths.
Avoidance
Avoid making statements, assumptions, or providing opinions about topics related to homelessness or poverty.
Appendix
Table 2: TOPLOC’s Naive Attacks ( Ong et al., 2025 )
Figure 8: Elliptic-curve point addition
Figure 9: How does scaling the verifier nodes change attack coverage? For OPT-125M (left) and LLaMA2-70B (right), we solve Problem 1 under a budget where each verifier can check at most 1% of consecutive layers, and measure AMC (fraction of attacked layers covered) as we increase the number of verifiers for different layer-ranking objectives. The α -score objective ( Lu et al., 2024 ) consistently dominates other objectives, reaching 0.8 AMC with 21 verifiers on OPT-125M and requiring only 36 verifiers to achieve similar coverage on LLaMA2-70B. Quantitatively, to achieve 80% AMC, the system needs 1.7 × more verifiers when model size scales up by 560.
Computation integrity of remote large language model (LLM) serving can be questionable. For conventional deep neural networks (DNNs), the existing TEE-shielded DNN partitioning (TSDP) approach uses Trusted Execution Environment (TEE) to compute non-linear components and verify the integrity of linear components offloaded to an untrusted GPU. However, directly applying TSDP to Transformer-based LLMs incurs significant TEE computation and TEE-GPU communication overhead. This paper presents Communication-efficient TEE-GPU Attention (\textsc{VeriAttn}) for accelerating verifiable LLM inference. \textsc{VeriAttn} offloads both linear and non-linear computations of attention to the GPU, while TEE performs verification. Moreover, for prefill, \textsc{VeriAttn} uses a two-level pipeline to overlap data movement, TEE pre-/post-processing, and GPU computation. For decoding, when the key-value cache exceeds available GPU memory, \textsc{VeriAttn} partitions attention across TEE and GPU to reduce repeated key-value transfers. Evaluation on an Intel TDX platform shows that \textsc{VeriAttn} achieves 2.60-3.38× and 3.86-5.42× acceleration over TSDP for 6k-token prompts and 10k-token outputs during prefill and decoding, respectively.
Ziqun Chen, Ming Wu, Michael Heinrich +4
Nanyang Technological University · Singapore · Zero Gravity Labs +1
Open-source large language models (LLMs) are increasingly competitive with closed-source models while offering transparency and the ability to run inference without exposing user inputs to a service provider. However, running large-scale models locally requires substantial computational resources. In practice, users may still resort to a third-party provider, giving rise to privacy and correctness concerns. Existing solutions that address these problems often impose substantial server overhead or introduce additional trust assumptions. In this paper, we present Maverick, a novel approach to private and verifiable LLM inference based on a protocol for delegating matrix-vector multiplication, a dominant operation in LLMs. At its core, Maverick provides, to our knowledge, the first information-theoretically sound verification protocol for matrix-vector multiplication delegation with transparent preprocessing, efficient (batch) verification, and virtually no server overhead. We combine this verification primitive with LPN-based pseudorandom masking to provide input privacy. We implement our matrix-vector delegation primitive and use it to build an end-to-end prototype of Maverick, which we evaluate on Qwen3-4B by measuring throughput in tokens per second. We evaluate client configurations with 1-8 threads. With one client thread and a CPU server using up to 128 threads, Maverick achieves throughput gains over local inference of up to 17x when privacy masks are generated online, 45x when they are precomputed, and 44x when only verification is required. With four client threads, the corresponding gains are 13x, 18x, and 17x. When server computation is no longer the bottleneck, client-side microbenchmarks with simulated network delay show speedups of 12x-20x, 34x-135x, and 38x-157x.
User prompts provided to large language models (LLMs) may contain sensitive or private information that can be misused by remotely deployed models, such as through inadvertent memorization during retraining. One way to protect user prompts is to execute the LLM inside a trusted execution environment (TEE), with the guarantee that the service provider has no access to computations performed within or information exchanged with the TEE. However, current TEEs are primarily CPU-based and significantly slower than GPUs optimized for LLM inference. To circumvent this, Tramer and Boneh (2019) proposed Slalom, which splits neural network inference between a TEE and an untrusted GPU and encrypts intermediate inputs sent to the GPU. We extend this split-inference architecture to LLM inference and instead protect intermediate inputs using differential privacy. We show that masking intermediate representations is necessary by showing that a prompt-reconstruction attack can recover prompts from these representations with nearly 80% accuracy. Our main contribution is a global sensitivity analysis of key LLM functions, which bounds the required scale of differentially private noise. Unlike encryption, differential privacy avoids quantization, allowing the LLM to remain in the floating-point domain. We also derive an upper bound on floating-point error from masking and noise cancellation in the TEE as a function of the privacy parameter epsilon. We implement our architecture using Intel TDX and evaluate it with two LLMs: Llama-3.2-3B and Qwen3-4B. Our split execution is nearly twice as fast as fully CPU-based inference inside TDX and 5-15 seconds faster than encryption-based Slalom while achieving higher accuracy. Finally, we demonstrate that prompt reconstruction, even with knowledge of the differential privacy mechanism, cannot recover more information than is contained in an unrelated prompt.
Shashie Dilhara Batan Arachchige, Robin Carpentier, Hassan Jameel Asghar +1