Foundation model recommender systems require user context that can be consumed by large language models, reasoned over, and refined through natural-language interaction. Traditional behavioral embedding vectors remain highly effective for retrieval and ranking, but they are opaque to users and not natively expressed for language model workflows. We present Textual User Taste, a system that generates structured natural-language taste profiles from listening behavior, interaction signals, content metadata, and optional user feedback, and deploys them to millions of Spotify users. We describe the end-to-end production lifecycle required to generate, evaluate, optimize, and maintain these representations at industrial scale, including prompt development and compression, user steering, and integration with downstream personalization systems. Because no unique ground-truth taste profile exists, we introduce a multi-faceted evaluation framework to evaluate taste profiles as a production representation: they carry user-specific predictive signal independently, and when integrated with behavioral embeddings, improve MRR by 0.6% for future-track prediction and NDCG@7 by 2.2% for search ranking. Our evaluation also reveals that taste profiles support positive natural-language steering, while exposing important limitations, including challenges with negation and short-term temporal adaptation. These findings position taste profiles not as replacements for behavioral embeddings, but as an interpretable and steerable interface between evolving user context and foundation-model recommender systems.
Figures & tables
Condition
Tokens
Factuality
Valid output
Uncompressed
∼ 6,800
8.66
100%
LLMLingua 2×
3,428
8.61
100%
LLMLingua 4×
1,803
7.80
73.1%
LLMLingua 8×
1,202
3.31
16.5%
Table 1. Instruction compression trade-offs. Factuality is scored from 1 to 10.
Representation
MRR
AUC
NDCG@10
Behavioral
0.956
0.957
0.932
Taste profile
0.937
0.946
0.912
Taste profile + behavioral
0.962
0.963
0.941
Gated fusion
0.961
0.962
0.940
Table 2. Future-track ranking (five-run mean).
Representation
MRR
AUC
NDCG@10
Behavioral
0.956
0.957
0.932
Taste profile
0.937
0.946
0.912
Taste profile + behavioral
0.962
0.963
0.941
Gated fusion
0.961
0.962
0.940
Table 2. Future-track ranking (five-run mean).
Context
NDCG @7 Uplift
None
baseline
Taste profile
+0.9%
Behavioral
+1.7%
Combined
+2.2%
Table 3. Search ranking; uplift is relative to no context.