Reliable tool use requires more than triggering a mechanism or matching a query to an API description. Before selecting a specific tool, an agent must first infer the capability requirements implied by the user query. In this paper, we investigate whether these query-side capability requirements are linearly decodable from LLM hidden representations prior to generation, and how this hidden-state accessibility compares with explicit verbal classification. We introduce TACIT, a framework that decomposes external requirements along three fundamental axes: Source, Transformation, and World Effect, defining eight structurally distinct capability classes. Using 1,600 balanced training queries from benchmarks, synthetic examples, and new domain scenarios, we train linear probes on pre-generation hidden states from four open-weight LLM families. Our empirical results demonstrate that fine-grained capability structures are linearly decodable with high accuracy across all models. Crucially, however, we expose a representation-to-verbalization gap: these same models are significantly less reliable when asked to explicitly classify the same queries in natural language. This disconnect indicates that information about required external capabilities is linearly accessible in LLM hidden representations but not reliably expressed, a phenomenon we define as "structured but silent."
Figures & tables
Figure 1: Overview of the TACIT probing framework. (I) Given a user query, we extract hidden states at the final prompt token before generation. Separate linear probes, with a layer selected on validation data for each axis, predict Source, Transformation, and World Effect, jointly defining eight capability classes. (II) We examine these predictions through two analyses: (A) robustness to misleading surface cues, comparing probes with lexical and sentence-embedding baselines on deceptive queries; and (B) the representation-to-verbalization gap, comparing probe predictions with the same model’s explicit verbal classifications.
Figure 2: Overview of TACIT dataset composition and splits. Training combines 800 queries sampled from the benchmark-derived corpus with 800 domain-scenario queries. The domain-scenario dataset supplies 160 validation and 720 test queries; another 200 benchmark-source held-out queries complete the 920-query main test set. Each scenario family contains one query per TACIT class and remains within a single split.
Class
S
T
W
Representative Example Query
(A,R,O)
A
R
O
Which chemical element has the symbol Fe?
(A,T,O)
A
T
O
Ruel has four books of 10 stamps and six books of 15 stamps. How many stamps does Ruel have?
(E,R,O)
E
R
O
What is Cristiano Ronaldo’s current club?
(E,T,O)
E
T
O
Which department has the largest number of employees?
(A,R,M)
A
R
M
Can you set a new alarm for me at 17:15?
(A,T,M)
A
T
M
Calculate 18% tip on my $85 dinner bill and add the total amount to my expense report.
Table 1: Taxonomy overview and representative example queries for the eight TACIT capability classes.
Probe Axis
Best Layer
Relative Depth
Accuracy (%)
Source
L14
43.8%
97.28
Transformation
L7
21.9%
97.28
World Effect
L16
50.0%
99.24
Joint (3-axis)
per-axis best layers
94.02
Table 2: Best-layer linear probing accuracy on Llama 3.1 8B Instruct (main test set, N=920 ). Relative depth is l/32 . Joint accuracy requires all three axis predictions to be correct.
Capability Class
S
T
W
Joint Accuracy (%)
(A,R,O)
A
R
O
91.30
(A,T,O)
A
T
O
98.26
(E,R,O)
E
R
O
96.52
(E,T,O)
E
T
O
90.43
(A,R,M)
A
R
M
93.91
(A,T,M)
A
T
M
94.78
Table 3: Per-class joint accuracy on Llama 3.1 8B Instruct (main test set, N=115 per class).
Axis Pair
cos(wi,wj)
Source × Transformation
−0.016
Source × World Effect
−0.009
Transformation × World Effect
−0.030
Table 4: Cosine similarity between best-layer probe weight vectors on Llama 3.1 8B Instruct.
Target-axis Accuracy
Method
Source
Trans.
W-Eff.
Joint
BoW
61.04
44.29
79.45
38.64
N-gram
64.94
44.29
76.71
35.91
E5
55.84
55.71
75.34
37.27
Llama 3.1 8B probe
72.73
67.14
97.26
60.45
Table 5: Accuracy on the deceptive evaluation set (%). Source, Transformation, and World Effect columns evaluate queries targeting the respective axis ( N=77 , 70 , and 73 ). Joint accuracy is computed over all 220 queries and requires all three axis predictions to be correct.
Accuracy
Probe
four-shot
Gap (pp)
Source
97.28
78.80
+18.48
Transformation
97.28
88.80
+8.48
World Effect
99.24
98.04
+1.20
Joint (3-axis)
94.02
68.04
+25.98
Table 6: Probe and four-shot verbalization accuracy on Llama 3.1 8B Instruct (main test set, N=920 ). The gap is probe accuracy minus verbalization accuracy in percentage points.
Main Test Joint Accuracy (%)
Deceptive-Set Joint Accuracy (%)
Model
Probe
BoW
N-gram
E5
four-shot
Probe
BoW
N-gram
E5
Llama 3.1 8B
94.02
83.26
85.98
84.13
68.04
60.45
38.64
35.91
37.27
Qwen 3 8B
92.50
83.26
85.98
84.13
59.67
62.27
38.64
35.91
37.27
Gemma 2 9B
94.78
83.26
85.98
84.13
69.46
60.91
38.64
35.91
37.27
Mistral 7B v0.3
92.50
83.26
85.98
84.13
54.35
58.18
38.64
35.91
37.27
Table 7: Cross-model joint three-axis accuracy on the main test set ( N=920 ) and the deceptive evaluation set ( N=220 ). The text-only baselines have the same scores for each model because they use the same query data independently of the probed LLM. The four-shot column reports each model’s verbal classification on the main test set.
Figure 3: Four-way breakdown of joint three-axis correctness for the probe and four-shot verbalization on Llama 3.1 8B Instruct ( N=920 ). Bars show the fraction of test queries in each outcome, with counts shown above them.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Component
Generation model
Generation role
Review model
Review role
Original corpus and benchmark-source held-out set
GPT-4.1-mini
Provisional labeling of BFCL candidates, synthetic supplementation, and adaptation of τ -bench contexts.
Gemini 2.5 Flash
Three-axis label prediction and query-quality checks.
Domain-scenario dataset
GPT-5-mini
Generation of eight-class scenario families and targeted revision of rejected queries.
Gemini 2.5 Flash
Item-level label and quality checks, plus family-level coherence and coverage checks.
Deceptive evaluation set
GPT-5
Generation of deceptive queries with a specified target axis, gold label, and misleading surface cue.
Claude Sonnet 5
Blind three-axis labeling and checks of label correctness, naturalness, ambiguity, grounding, and deceptive-cue validity.
Appendix
Table 8: Roles of the generation and automated review models. For benchmark-derived data, labeling refers to assigning initial TACIT labels to existing queries, adaptation refers to rewriting benchmark task contexts into requests matching a target TACIT class, and supplementation refers to generating additional queries where class coverage is insufficient. Original benchmark queries are not generated by these models.
Class
Original-corpus source
Selection rationale
Held-out source
Selection rationale
(A,R,O)
TriviaQA
Stable factual questions requiring retrieval of available knowledge.
SimpleQA
Factual questions from a separate benchmark source.
(A,T,O)
GSM8K
Self-contained mathematical problems requiring transformation of supplied information.
MATH
Mathematical problems from a separate benchmark source.
Time-sensitive questions from a separate benchmark source.
(E,T,O)
Spider
Structured-data questions requiring operations over external records.
WikiSQL
Table questions involving aggregation over external records.
Mutate classes
BFCL v3 and supplementation
Class-matching state-changing requests, supplemented where benchmark coverage is insufficient.
τ -bench contexts and supplementation
Requests adapted from airline and retail task contexts, with additional examples to cover all four Mutate classes.
Appendix
Table 9: Benchmark sources informing the original corpus and the 200-query benchmark-source held-out set. The domain-scenario dataset is constructed separately using the domain specifications in Table 10 .
Domain group
Design inspiration
Domains
Count
Seen domains
SGD-style task-oriented settings Rastogi et al. (2020)
Banking; payment and transfer; calendar; events; flights; hotels; restaurants; rental cars; local transport; home services; movies; music and media.
12
BFCL-style API workflows
Email and messaging; file and cloud storage; smart home; software development and DevOps.
4
Common application and API workflows
Shopping cart; subscription management; project management; database and reporting.
4
Held-out domains
τ -bench-style support settings
Retail order support; airline customer support.
2
Common application and API workflows
Healthcare appointments; education and library; government services.
3
Appendix
Table 10: The 25 domains in the domain-scenario dataset: 20 for training, validation, and in-domain testing, and 5 for held out within this component. Design inspirations guide domain cards, not copied queries; held-out domains may overlap with topics or operations in the original corpus.