Structured but Silent: Probing Capability Requirements in LLM Hidden States
Organizations: Soongsil University
Abstract
Reliable tool use requires more than triggering a mechanism or matching a query to an API description. Before selecting a specific tool, an agent must first infer the capability requirements implied by the user query. In this paper, we investigate whether these query-side capability requirements are linearly decodable from LLM hidden representations prior to generation, and how this hidden-state accessibility compares with explicit verbal classification. We introduce TACIT, a framework that decomposes external requirements along three fundamental axes: Source, Transformation, and World Effect, defining eight structurally distinct capability classes. Using 1,600 balanced training queries from benchmarks, synthetic examples, and new domain scenarios, we train linear probes on pre-generation hidden states from four open-weight LLM families. Our empirical results demonstrate that fine-grained capability structures are linearly decodable with high accuracy across all models. Crucially, however, we expose a representation-to-verbalization gap: these same models are significantly less reliable when asked to explicitly classify the same queries in natural language. This disconnect indicates that information about required external capabilities is linearly accessible in LLM hidden representations but not reliably expressed, a phenomenon we define as "structured but silent."
Figures & tables
| Class | S | T | W | Representative Example Query |
|---|---|---|---|---|
| (A,R,O) | A | R | O | Which chemical element has the symbol Fe? |
| (A,T,O) | A | T | O | Ruel has four books of 10 stamps and six books of 15 stamps. How many stamps does Ruel have? |
| (E,R,O) | E | R | O | What is Cristiano Ronaldo’s current club? |
| (E,T,O) | E | T | O | Which department has the largest number of employees? |
| (A,R,M) | A | R | M | Can you set a new alarm for me at 17:15? |
| (A,T,M) | A | T | M | Calculate 18% tip on my $85 dinner bill and add the total amount to my expense report. |
| Probe Axis | Best Layer | Relative Depth | Accuracy (%) |
|---|---|---|---|
| Source | L14 | 43.8% | 97.28 |
| Transformation | L7 | 21.9% | 97.28 |
| World Effect | L16 | 50.0% | 99.24 |
| Joint (3-axis) | per-axis best layers | 94.02 | |
| Capability Class | S | T | W | Joint Accuracy (%) |
|---|---|---|---|---|
| (A,R,O) | A | R | O | 91.30 |
| (A,T,O) | A | T | O | 98.26 |
| (E,R,O) | E | R | O | 96.52 |
| (E,T,O) | E | T | O | 90.43 |
| (A,R,M) | A | R | M | 93.91 |
| (A,T,M) | A | T | M | 94.78 |
| Axis Pair | |
|---|---|
| Source Transformation | |
| Source World Effect | |
| Transformation World Effect |
| Target-axis Accuracy | ||||
|---|---|---|---|---|
| Method | Source | Trans. | W-Eff. | Joint |
| BoW | 61.04 | 44.29 | 79.45 | 38.64 |
| N-gram | 64.94 | 44.29 | 76.71 | 35.91 |
| E5 | 55.84 | 55.71 | 75.34 | 37.27 |
| Llama 3.1 8B probe | 72.73 | 67.14 | 97.26 | 60.45 |
| Accuracy | Probe | four-shot | Gap (pp) |
|---|---|---|---|
| Source | 97.28 | 78.80 | +18.48 |
| Transformation | 97.28 | 88.80 | +8.48 |
| World Effect | 99.24 | 98.04 | +1.20 |
| Joint (3-axis) | 94.02 | 68.04 | +25.98 |
| Main Test Joint Accuracy (%) | Deceptive-Set Joint Accuracy (%) | ||||||||
| Model | Probe | BoW | N-gram | E5 | four-shot | Probe | BoW | N-gram | E5 |
| Llama 3.1 8B | 94.02 | 83.26 | 85.98 | 84.13 | 68.04 | 60.45 | 38.64 | 35.91 | 37.27 |
| Qwen 3 8B | 92.50 | 83.26 | 85.98 | 84.13 | 59.67 | 62.27 | 38.64 | 35.91 | 37.27 |
| Gemma 2 9B | 94.78 | 83.26 | 85.98 | 84.13 | 69.46 | 60.91 | 38.64 | 35.91 | 37.27 |
| Mistral 7B v0.3 | 92.50 | 83.26 | 85.98 | 84.13 | 54.35 | 58.18 | 38.64 | 35.91 | 37.27 |
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
| Component | Generation model | Generation role | Review model | Review role |
|---|---|---|---|---|
| Original corpus and benchmark-source held-out set | GPT-4.1-mini | Provisional labeling of BFCL candidates, synthetic supplementation, and adaptation of -bench contexts. | Gemini 2.5 Flash | Three-axis label prediction and query-quality checks. |
| Domain-scenario dataset | GPT-5-mini | Generation of eight-class scenario families and targeted revision of rejected queries. | Gemini 2.5 Flash | Item-level label and quality checks, plus family-level coherence and coverage checks. |
| Deceptive evaluation set | GPT-5 | Generation of deceptive queries with a specified target axis, gold label, and misleading surface cue. | Claude Sonnet 5 | Blind three-axis labeling and checks of label correctness, naturalness, ambiguity, grounding, and deceptive-cue validity. |
| Class | Original-corpus source | Selection rationale | Held-out source | Selection rationale |
|---|---|---|---|---|
| (A,R,O) | TriviaQA | Stable factual questions requiring retrieval of available knowledge. | SimpleQA | Factual questions from a separate benchmark source. |
| (A,T,O) | GSM8K | Self-contained mathematical problems requiring transformation of supplied information. | MATH | Mathematical problems from a separate benchmark source. |
| (E,R,O) | FreshQA | Fast-changing factual questions requiring external information. | RealTimeQA | Time-sensitive questions from a separate benchmark source. |
| (E,T,O) | Spider | Structured-data questions requiring operations over external records. | WikiSQL | Table questions involving aggregation over external records. |
| Mutate classes | BFCL v3 and supplementation | Class-matching state-changing requests, supplemented where benchmark coverage is insufficient. | -bench contexts and supplementation | Requests adapted from airline and retail task contexts, with additional examples to cover all four Mutate classes. |
| Domain group | Design inspiration | Domains | Count |
| Seen domains | SGD-style task-oriented settings Rastogi et al. (2020) | Banking; payment and transfer; calendar; events; flights; hotels; restaurants; rental cars; local transport; home services; movies; music and media. | 12 |
| BFCL-style API workflows | Email and messaging; file and cloud storage; smart home; software development and DevOps. | 4 | |
| Common application and API workflows | Shopping cart; subscription management; project management; database and reporting. | 4 | |
| Held-out domains | -bench-style support settings | Retail order support; airline customer support. | 2 |
| Common application and API workflows | Healthcare appointments; education and library; government services. | 3 |
| Layer | Source (A/E) | Transformation (R/T) | World Effect (O/M) |
|---|---|---|---|
| Llama 3.1 8B Instruct (32 layers; joint validation accuracy = 0.9813) | |||
| Best layers: S = L14 (43.75%), T = L7 (21.88%), W = L16 (50.00%) | |||
| L1 | 0.9000 | 0.9250 | 0.8688 |
| L2 | 0.9438 | 0.9500 | 0.8813 |
| L3 | 0.9000 | 0.9188 | 0.8875 |
| L4 | 0.9000 | 0.9688 | 0.8875 |