Paper ID: 2311.10083
Characterizing Tradeoffs in Language Model Decoding with Informational Interpretations
Chung-Ching Chang, William W. Cohen, Yun-Hsuan Sung
We propose a theoretical framework for formulating language model decoder algorithms with dynamic programming and information theory. With dynamic programming, we lift the design of decoder algorithms from the logit space to the action-state value function space, and show that the decoding algorithms are consequences of optimizing the action-state value functions. Each component in the action-state value function space has an information theoretical interpretation. With the lifting and interpretation, it becomes evident what the decoder algorithm is optimized for, and hence facilitating the arbitration of the tradeoffs in sensibleness, diversity, and attribution.
Submitted: Nov 16, 2023