Who Is Behind the Harness? Fingerprinting LLMs through Agentic Behavior
Organizations: Department of Computer Science and Technology, Tsinghua University
Abstract
LLMs increasingly operate through coding-agent harnesses that inspect repositories, invoke tools, and modify files. Substituting the model behind such an agent can therefore change security-relevant decisions, including whether it verifies changes or recovers safely from failures. Existing LLM fingerprints largely infer identity from direct text or token distributions. In coding agents, these signals are mediated by system instructions, controller logic, tools, and execution feedback, limiting their transfer. We present LIDAR (LLM Identification from Decisions and Actions at Runtime), an active black-box fingerprinting method for coding-agent execution. Three coding probe pairs expose post-edit verification, transient-failure recovery, and specification--test conflict resolution under controlled changes. LIDAR represents the resulting trajectories with complementary instance-level and distribution-level features and compares them with clean references using a lightweight probabilistic identifier. It requires no access to model weights, logits, or provider internals. Across 36 models from seven families and two agent harnesses, LIDAR achieves high Top-1 accuracy and MRR and outperforms four existing fingerprinting and API-auditing baselines. Ablations confirm that the two feature levels, all probe pairs, and their controlled variants contribute. These results show that agent execution behavior provides model-identity evidence beyond final outputs.
Figures & tables
| Feature group | Information retained | Fields |
|---|---|---|
| Instance-level structure | Per-variant action roles, transitions, positions, repetition, backtracking, and recovery paths | 588 |
| Distribution-level behavior | Across six trajectories, the mean and sample standard deviation of action, command-use, ordering, repetition, recovery, outcome, and coarse response-form descriptors | 202 |
| Training/enrollment | Test/query | Successful runs | |||||
|---|---|---|---|---|---|---|---|
| Method | Atomic sampling unit | Units | Runs | Units | Runs | Per setting | Full panel |
| LLMmap [ 26 ] | 8 probes / configuration | 20 (configs) | 160 | 20 (configs) | 160 | 320 | 23,040 |
| ZeroPrint [ 30 ] | 5 responses / base query | 10 (bases) | 50 | 10 (bases) | 50 | 100 | 7,200 |
| MET [ 10 ] | 10 prompts / repetition | 5 (reps) | 50 | 5 (reps) | 50 | 100 | 7,200 |
| One Token [ 3 ] | 8 coordinates / repetition | 5 (reps) | 40 | 5 (reps) | 40 | 80 | 5,760 |
| Lidar | 3 probe pairs (6 variants) / sample | 6 (samples) | 36 | 6 (samples) | 36 | 72 | 5,184 |
| OpenCode [ 25 ] | mini-swe-agent [ 31 ] | |||||
|---|---|---|---|---|---|---|
| Method | Top-1 | Top-3 | MRR | Top-1 | Top-3 | MRR |
| LLMmap [ 26 ] | 0.6143{\color[rgb]{0.5,0.5,0.5}\scriptscriptstyle\,(\pm\,0.0170)} | 0.8189{\color[rgb]{0.5,0.5,0.5}\scriptscriptstyle\,(\pm\,0.0134)} | 0.7319{\color[rgb]{0.5,0.5,0.5}\scriptscriptstyle\,(\pm\,0.0121)} | 0.5628{\color[rgb]{0.5,0.5,0.5}\scriptscriptstyle\,(\pm\,0.0223)} | 0.7654{\color[rgb]{0.5,0.5,0.5}\scriptscriptstyle\,(\pm\,0.0215)} | 0.6860{\color[rgb]{0.5,0.5,0.5}\scriptscriptstyle\,(\pm\,0.0186)} |
| MET [ 10 ] | 0.8667{\color[rgb]{0.5,0.5,0.5}\scriptscriptstyle\,(\pm\,0.0489)} | 0.9574{\color[rgb]{0.5,0.5,0.5}\scriptscriptstyle\,(\pm\,0.0349)} | 0.9160{\color[rgb]{0.5,0.5,0.5}\scriptscriptstyle\,(\pm\,0.0340)} | 0.5917{\color[rgb]{0.5,0.5,0.5}\scriptscriptstyle\,(\pm\,0.0443)} | 0.7315{\color[rgb]{0.5,0.5,0.5}\scriptscriptstyle\,(\pm\,0.0382)} | 0.6927{\color[rgb]{0.5,0.5,0.5}\scriptscriptstyle\,(\pm\,0.0277)} |
| ZeroPrint [ 30 ] | 0.2833{\color[rgb]{0.5,0.5,0.5}\scriptscriptstyle\,(\pm\,0.0272)} | 0.4333{\color[rgb]{0.5,0.5,0.5}\scriptscriptstyle\,(\pm\,0.0377)} | 0.3971{\color[rgb]{0.5,0.5,0.5}\scriptscriptstyle\,(\pm\,0.0227)} | 0.4000{\color[rgb]{0.5,0.5,0.5}\scriptscriptstyle\,(\pm\,0.0377)} | 0.5889{\color[rgb]{0.5,0.5,0.5}\scriptscriptstyle\,(\pm\,0.0408)} | 0.5197{\color[rgb]{0.5,0.5,0.5}\scriptscriptstyle\,(\pm\,0.0205)} |
| One Token [ 3 ] | 0.7333{\color[rgb]{0.5,0.5,0.5}\scriptscriptstyle\,(\pm\,0.0572)} | 0.9167{\color[rgb]{0.5,0.5,0.5}\scriptscriptstyle\,(\pm\,0.0465)} | 0.8294{\color[rgb]{0.5,0.5,0.5}\scriptscriptstyle\,(\pm\,0.0356)} | 0.2722{\color[rgb]{0.5,0.5,0.5}\scriptscriptstyle\,(\pm\,0.0272)} | 0.3056{\color[rgb]{0.5,0.5,0.5}\scriptscriptstyle\,(\pm\,0.0000)} | 0.2878{\color[rgb]{0.5,0.5,0.5}\scriptscriptstyle\,(\pm\,0.0152)} |
| Lidar | \mathbf{0.8836}{\color[rgb]{0.5,0.5,0.5}\scriptscriptstyle\,(\pm\,0.0185)} | \mathbf{0.9838}{\color[rgb]{0.5,0.5,0.5}\scriptscriptstyle\,(\pm\,0.0072)} | \mathbf{0.9336}{\color[rgb]{0.5,0.5,0.5}\scriptscriptstyle\,(\pm\,0.0107)} | \mathbf{0.9513}{\color[rgb]{0.5,0.5,0.5}\scriptscriptstyle\,(\pm\,0.0124)} | \mathbf{0.9962}{\color[rgb]{0.5,0.5,0.5}\scriptscriptstyle\,(\pm\,0.0036)} | \mathbf{0.9736}{\color[rgb]{0.5,0.5,0.5}\scriptscriptstyle\,(\pm\,0.0069)} |
| Group | Configuration | Fields | Top-1 | Top-1 (pp) | Top-3 | MRR |
| Reference | Full method | 958 | 0.9352 | – | 0.9884 | 0.9631 |
| Feature view | Instance only | 756 | 0.7500 | 0.9005 | 0.8377 | |
| Distribution only | 202 | 0.8981 | 0.9884 | 0.9426 | ||
| Distribution statistic | Instance + mean | 857 | 0.9190 | 0.9838 | 0.9529 | |
| Instance + standard deviation | 857 | 0.9097 | 0.9792 | 0.9454 | ||
| Probe pair | Without Verify | 724 | 0.8819 | 0.9745 | 0.9315 |
| Clean | A1: Identity obfuscation | A2: Output-format control | |||||
|---|---|---|---|---|---|---|---|
| Harness | Method | Acc. | MRR | Acc. | MRR | Acc. | MRR |
| mini-swe-agent [ 31 ] | LLMmap [ 26 ] | 0.5667 | 0.7639 | 0.4667 ( ) | 0.7000 ( ) | 0.3333 ( ) | 0.5533 ( ) |
| ZeroPrint [ 30 ] | 0.3333 | 0.5111 | 0.0000 ( ) | 0.3944 ( ) | 0.1667 ( ) | 0.4222 ( ) | |
| MET [ 10 ] | 0.5000 | 0.6444 | 0.1667 ( ) | 0.4500 ( ) | 0.1667 ( ) | 0.4083 ( ) | |
| One Token [ 3 ] | 0.3333 | 0.4167 | 0.1667 ( ) | 0.2500 ( ) | 0.0000 ( ) | 0.0000 ( ) | |
| Lidar | 1.0000 | 1.0000 | 0.8889 ( ) | 0.9444 ( ) | 1.0000 | 1.0000 | |
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
| Family | # | Models |
|---|---|---|
| GPT [ 24 ] | 7 | gpt-5.1, gpt-5.2, gpt-5.4, gpt-5.5, gpt-5.6-luna, gpt-5.6-sol, gpt-5.6-terra |
| Claude [ 2 ] | 7 | claude-fable-5, claude-opus-4-6, claude-opus-4-7, claude-opus-4-8, claude-opus-5, claude-sonnet-4-6, claude-sonnet-5 |
| Qwen [ 35 ] | 7 | Qwen3.5-122B-A10B, Qwen3.5-27B, Qwen3.6-Plus, Qwen3.7-Max, Qwen3.7-Plus, Qwen3.8-2.4T-A95B, Qwen3.8-27B |
| GLM [ 12 ] | 6 | GLM-4.7, GLM-5, GLM-5-Turbo, GLM-5.1, GLM-5.2, GLM-5.3 |
| DeepSeek [ 9 ] | 4 | DeepSeek-V3.1, DeepSeek-V3.2, DeepSeek-V4-Flash, DeepSeek-V4-Pro |
| Kimi [ 17 ] | 3 | Kimi-K2.5, Kimi-K2.6, Kimi-K3 |