From High Recall to High Utility: Dataset-Adaptive Post-Processing of LLM-Generated Customer Intents
Organizations: Cisco Systems, Inc., San Jose, CA, USA
Abstract
Large language models can extract useful signals from heterogeneous enterprise data, but high-recall extraction often produces outputs that are duplicated, uneven in granularity, semantically overlapping, or too numerous for downstream systems and human reviewers to use effectively. We present a dataset-adaptive post-processing architecture developed for Customer Intent Extraction (CIE), where unstructured customer language is transformed into stable, traceable intent units. The approach separates recall-oriented extraction from utility-oriented reduction. Source-specific preprocessing first isolates evidence from multimodal plans, sparse operational records, and structured opportunity data. Candidate intents are then standardized and deduplicated, optionally enriched with metadata for embedding computation, represented in a shared semantic vector space, and grouped using a clustering strategy selected according to the candidate set's characteristics. Cluster-level keywords provide an explainability layer, while singleton reassignment requires agreement between embedding and keyword similarity. Finally, constrained language-model aggregation produces one concise intent per cluster without introducing unsupported concepts, and the resulting unit retains provenance, clustering, embedding, and generation metadata. This treats post-processing not as cosmetic cleanup, but as a semantic reduction layer converting high-recall LLM outputs into reusable enterprise intelligence. We also describe two downstream applications: Machine-Generated Intents, which infer likely objectives for customers lacking direct evidence from peer customers with similar profiles, and intent-guided semantic retrieval and mapping, which uses the stable intent as a query against a downstream decision space, illustrated here by mapping customer intents to business outcomes.
Figures & tables
| Dimension | E2E plans | CS Console | Sales Data Sources |
|---|---|---|---|
| Primary form | Slides, tables, merged cells, structured narrative | Large sets of fielded records and free-form entries | Related structured tables containing intent, opportunity, and product information |
| Main failure mode | Loss of layout and table relationships | Sparse fields, low-signal updates, duplicates, missing product context | Relevant context distributed across tables and inconsistent product representations |
| Core preparation | Target-chapter selection, multimodal table extraction, structured validation | Product-scope filtering, regex rejection, empty-field handling, exact deduplication | Table joins, product mapping, feature engineering, schema normalization |
| Output | Compact structured context package | Filtered records grouped by customer and product | Structured intent/opportunity record with enriched product context |
| Step | Stage | Purpose | Key output |
|---|---|---|---|
| 1 | Standardization | Remove formatting variation, exact duplicates, empty items, and metadata inconsistency | Clean candidate intent objects |
| 2–3 | Context padding and embeddings | Enrich short/ambiguous text for embedding only and map candidates to a common semantic vector space | Padded representation, unchanged display text, embedding and model version |
| 4–5 | Adaptive clustering | Select and apply an interpretable/scalable grouping strategy according to candidate-set characteristics | Cluster assignments, centroids, method and parameters |
| 6 | Keyword extraction | Describe shared cluster meaning and provide an explainability signal | Normalized cluster keyword set |
| 7 | Outlier handling | Preserve unique intents or cautiously repair uncertain singleton assignments | Updated clusters plus retained singletons |
| 8 | Constrained aggregation | Produce one faithful, non-redundant statement per cluster | Combined intent plus generation metadata |
| Approach | Role and Characteristics | Strengths | Weaknesses |
|---|---|---|---|
| Bi-Encoder | Used for initial candidate retrieval. Intent and outcome texts are encoded independently, enabling efficient vector-similarity search over large candidate spaces. Model selection was informed by Massive Text Embedding Benchmark (MTEB) sentence-similarity benchmarks [ 14 ] . | Fast, scalable, and allows outcome embeddings to be precomputed and reused. | Lower ranking precision because intent–outcome interactions are not evaluated jointly. |
| Cross-Encoder | Applied only to shortlisted candidates. Each intent–outcome pair is evaluated jointly, allowing richer semantic interaction and more accurate ranking. | High semantic accuracy and stronger pairwise ranking. | More computationally expensive because each candidate pair must be evaluated separately. |
| Hybrid (chosen) | Uses the bi-encoder for high-recall candidate retrieval and the cross-encoder for final re-ranking of the shortlisted candidates. | Combines efficient retrieval with high-precision semantic ranking. | Introduces slightly greater pipeline complexity than either model used independently. |
| Mechanism | Rule |
|---|---|
| Score-gap detection | Ranked bi-encoder similarities are scanned for a sufficiently large drop between consecutive candidates; a detected drop is treated as evidence that semantic relevance is beginning to degrade and is used to restrict the candidate set before reranking, rather than relying on a single fixed similarity cutoff. |
| Maximum outcomes returned | No more than five business outcomes are returned per customer intent, bounding the output to prevent noisy long tails. |
| Final ordering | Determined by the cross-encoder ranking over the shortlisted candidates, not by the bi-encoder similarity used for initial retrieval. |
| Displayed similarity scores | Derived from the bi-encoder scores and adjusted, where necessary, to decrease monotonically with the final cross-encoder rank, so that displayed values remain consistent with the output ordering. |
| Approach | F1@5 | MRR | Latency (s) |
|---|---|---|---|
| Bi-encoder | 51.09 | 0.76 | 0.056 |
| Cross-encoder | 54.82 | 0.87 | 0.098 |
| Hybrid | 54.82 | 0.87 | 0.060 |
| Zero-shot classifier | 47.00 | 0.85 | 1.213 |
| Metric | Mean | Standard deviation | Interpretation |
|---|---|---|---|
| Faithfulness | 93.5% | 7.9% | Consistency of the generated response with the retrieved source context. |
| Faithfulness vs. ground truth | 66.1% | 31.9% | Faithfulness when evaluated relative to the manually established reference output. |
| Semantic similarity | 48.5 | 10.7 | Semantic resemblance between the generated output and the reference answer. |
| Answer correctness | 37.4% | 14.4% | Combined assessment of semantic and factual agreement with the reference answer. |
| Evaluation target | Evaluation scope | Result | Interpretation |
|---|---|---|---|
| Intent extraction | Five customers | 78% average accuracy | SMEs assessed whether individual inferred intents were supported and relevant for each customer. |
| Outcome relevancy | 20 customers across four products | 100% | SMEs confirmed that all recommended business outcomes were relevant to the respective customers. |
| Peer relevancy | 20 customers across four products | 81.6% | Most selected peers were judged relevant; remaining errors were associated primarily with inaccuracies in customer industry metadata used for peer grouping. |