Learning to Sell: Reinforcement Learning for Strategic Large Language Model Agents in Multi-Product Markets
Organizations: Massachusetts Institute of Technology, Cambridge, Massachusetts, USA · Purdue University, West Lafayette, Indiana, USA · California Institute of Technology, Pasadena, California, USA · University of Virginia, Charlottesville, Virginia, USA · Cornell University, Ithaca, New York, USA
Abstract
Autonomous large language model (LLM) agents operating in multi-product markets must make sequential decisions under information asymmetry and resource constraints. We develop a machine learning approach for training such agents to act effectively as sellers in a multi-item bargaining environment, where a seller concurrently negotiates a catalog of substitutable assets across a pool of independent buyers. Buyers hold private, heterogeneous valuations across products, and each can purchase at most one item. Facing limits on total communication turns, the seller must dynamically match buyers with the most profitable products considering their private valuations, while strategically allocating its limited interaction budget toward combinations of greater potential value. We formalize this problem as a Partially Observable Markov Decision Process using a structured, four-part message protocol that maps natural language into a parsable and regulated decision space. Using this formalization, we design a post-training method using Reinforcement Learning from Verifiable Rewards (RLVR). To evaluate this framework, we construct a multidimensional metric suite that quantifies constraint adherence, seller surplus extraction, and allocation quality. Our trained seller agent learns to match limited inventory to buyers more effectively, matching or outperforming trillion-parameter frontier models in both seller surplus extraction and buyer-product allocation quality. Finally, these learned strategies generalize robustly to unseen market structures, correlated valuation distributions, and price ranges not encountered during training.
Figures & tables
| Model | Params | Reward | Seller Surplus Extraction Ratio | Item Deal Rate | Deal on Top Buyer Rate | Optimal-Pair Offer Coverage Rate | Instruction Violation Rate |
| Qwen3-30B-A3B-Instruct-2507-trained ( Ours ) | 30B | ||||||
| Qwen3-30B-A3B-Instruct-2507-untrained | 30B | ||||||
| Kimi-K2.6-thinking | 1T | ||||||
| DeepSeek-V4-Pro-thinking | 1.6T | ||||||
| DeepSeek-V4-Pro-nothink | 1.6T | ||||||
| DeepSeek-V4-Flash-thinking | 284B |
| OOD Environment | Buyers | Items | Turns | Buyer Model | Valuation Structure | Product Category | Core Stress Test |
| In-Distribution Reference | |||||||
| 3 | 3 | 7 | Fixed iter- 60 | Baseline ( 2 ) | 12 Train Categories | Training configuration | |
| Scale Generalization | |||||||
| 1. Expanded Standard | 5 | 5 | 12 | iter- {50, 60, 70} | Baseline ( 2 ) | 12 Train Categories | Scale & horizon expansion |
| 2. Expanded Musical | 5 | 5 | 12 | iter- {50, 60, 70} | Baseline ( 2 ) | Held-out Musical Instruments | Scale & unseen price range |
| Asymmetric Market Structures | |||||||
Appendix figures & tables25 assets
Supplementary material from the paper’s appendix.
Appendix
| Symbol | Name | Description |
| Episode parameters and indices (§ 3 ) | ||
| , | Item and buyer counts | Number of items and buyer slots per episode. |
| , | Item and buyer indices | Item index ; buyer index . |
| Turn index | Seller turn index . | |
| , | Communication budget and turn limit | Global communication budget ( in training) and per-buyer turn limit ( ); . |
| Prices, costs, and valuations (§ 3 , § 5.2 , § 9 ) | ||
| Role | Base model | Provenance / checkpoint |
| Seller Starting Checkpoint | Qwen/Qwen3-30B-A3B- Instruct-2507 ( Qwen Team, 2025 ) | Original checkpoint from ( Qwen Team, 2025 ) ; fine-tuned as the seller. |
| Buyers in Training | Qwen/Qwen3-30B-A3B- Instruct-2507 | Frozen buyer checkpoint iteration (iter- ), trained via RLVR on a bilateral negotiation task ( Liu et al., 2026a ) . |
| Buyers in Evaluation | Qwen/Qwen3-30B-A3B- Instruct-2507 | In-distribution evaluation: fixed iter- . Out-of-distribution evaluation: per-buyer uniform over the pool . |
| Model Name | Parameter Count | Reference |
| Qwen3-30B-A3B-Instruct-2507-trained ( Ours ) | 30B | This Work |
| Qwen3-30B-A3B-Instruct-2507-untrained | 30B | Qwen Team (2025) |
| GPT-5.4-high-reasoning | closed-source | Singh et al. (2025) |
| GPT-5.4-mini-high-reasoning | closed-source | Singh et al. (2025) |
| DeepSeek-V4-Pro-thinking / nothink | 1.6T | DeepSeek-AI (2026) |
| DeepSeek-V4-Flash-thinking / nothink | 284B | DeepSeek-AI (2026) |
| Setting | Value |
| loss function | cispo |
| CISPO clip low / high | |
| Learning rate | |
| KL penalty (kl coef) | |
| batch size | |
| group size |
| Setting | Value |
| num buyers | |
| num items | |
| total turns | |
| per-buyer turn limit | |
| valuation ratio distribution | |
| first-offer-noise probability |
| Setting | Value |
| seller model | All models from Table 5 |
| buyer model | Qwen3-30B-A3B-Instruct-2507 iter- (Table 4 ) |
| Test split size | held-out clusters |
| seller temperature (eval) | |
| buyer temperature (eval) |
| Environment | Buyers | Items | Turns | Valuation | Catalog |
| In-distribution (reference) | uniform | categories | |||
| Env. 1: Expanded Standard | uniform | categories | |||
| Env. 2: Expanded Musical | uniform | Musical Instruments | |||
| Env. 3: Buyer Heavy | uniform | categories | |||
| Env. 4: Item Heavy | uniform | categories | |||
| Env. 5: Item Correlation | uniform | categories |
| Broad Category | Mean Distinct Brands | Median List Price | Median Cluster Min Price | Median Cluster Max Price | Median Price Spread |
| Appliances | 4.3 | $109.98 | $79.99 | $159.99 | 1.71 |
| Automotive | 4.3 | $45.86 | $31.49 | $59.99 | 2.00 |
| Baby | 4.5 | $34.99 | $28.24 | $48.97 | 1.61 |
| Beauty & Personal Care | 4.5 | $29.99 | $22.49 | $43.98 | 1.79 |
| Clothing, Shoes & Jewelry | 4.4 | $36.49 | $29.95 | $49.99 | 1.73 |
| Electronics | 3.9 | $139.99 | $99.79 | $199.99 | 1.95 |
| Model | Params | Reward | Seller Surplus Extraction Ratio | Item Deal Rate | Deal on Top Buyer Rate | Optimal-Pair Offer Coverage Rate | Instruction Violation Rate |
| Qwen3-30B-A3B-Instruct-2507-trained ( Ours ) | 30B | ||||||
| Qwen3-30B-A3B-Instruct-2507-untrained | 30B | ||||||
| Kimi-K2.6-thinking | 1T | ||||||
| DeepSeek-V4-Pro-thinking | 1.6T | ||||||
| DeepSeek-V4-Pro-nothink | 1.6T | ||||||
| DeepSeek-V4-Flash-thinking | 284B |
| Model | Params | Reward | Seller Surplus Extraction Ratio | Item Deal Rate | Deal on Top Buyer Rate | Optimal-Pair Offer Coverage Rate | Instruction Violation Rate |
| Qwen3-30B-A3B-Instruct-2507-trained ( Ours ) | 30B | ||||||
| Qwen3-30B-A3B-Instruct-2507-untrained | 30B | ||||||
| Kimi-K2.6-thinking | 1T | ||||||
| DeepSeek-V4-Pro-nothink | 1.6T | ||||||
| DeepSeek-V4-Flash-thinking | 284B | ||||||
| GPT-5.4-high-reasoning | closed-source |
| Model | Params | Reward | Seller Surplus Extraction Ratio | Item Deal Rate | Deal on Top Buyer Rate | Optimal-Pair Offer Coverage Rate | Instruction Violation Rate |
| Qwen3-30B-A3B-Instruct-2507-trained ( Ours ) | 30B | ||||||
| Qwen3-30B-A3B-Instruct-2507-untrained | 30B | ||||||
| Kimi-K2.6-thinking | 1T | ||||||
| DeepSeek-V4-Pro-nothink | 1.6T | ||||||
| DeepSeek-V4-Flash-thinking | 284B | ||||||
| GPT-5.4-high-reasoning | closed-source |
| Model | Params | Reward | Seller Surplus Extraction Ratio | Item Deal Rate | Deal on Top Buyer Rate | Optimal-Pair Offer Coverage Rate | Instruction Violation Rate |
| Qwen3-30B-A3B-Instruct-2507-trained ( Ours ) | 30B | ||||||
| Qwen3-30B-A3B-Instruct-2507-untrained | 30B | ||||||
| Kimi-K2.6-thinking | 1T | ||||||
| DeepSeek-V4-Pro-nothink | 1.6T | ||||||
| GPT-5.4-high-reasoning | closed-source | ||||||
| DeepSeek-V4-Flash-nothink | 284B |
| Model | Params | Reward | Seller Surplus Extraction Ratio | Item Deal Rate | Deal on Top Buyer Rate | Optimal-Pair Offer Coverage Rate | Instruction Violation Rate |
| Qwen3-30B-A3B-Instruct-2507-trained ( Ours ) | 30B | ||||||
| Qwen3-30B-A3B-Instruct-2507-untrained | 30B | ||||||
| Kimi-K2.6-thinking | 1T | ||||||
| DeepSeek-V4-Pro-nothink | 1.6T | ||||||
| GPT-5.4-high-reasoning | closed-source | ||||||
| DeepSeek-V4-Flash-nothink | 284B |
| Model | Params | Reward | Seller Surplus Extraction Ratio | Item Deal Rate | Deal on Top Buyer Rate | Optimal-Pair Offer Coverage Rate | Instruction Violation Rate |
| Qwen3-30B-A3B-Instruct-2507-trained ( Ours ) | 30B | ||||||
| Qwen3-30B-A3B-Instruct-2507-untrained | 30B | ||||||
| Kimi-K2.6-thinking | 1T | ||||||
| GPT-5.4-high-reasoning | closed-source | ||||||
| DeepSeek-V4-Pro-thinking | 1.6T | ||||||
| DeepSeek-V4-Pro-nothink | 1.6T |
| Model | Params | Reward | Seller Surplus Extraction Ratio | Item Deal Rate | Deal on Top Buyer Rate | Optimal-Pair Offer Coverage Rate | Instruction Violation Rate |
| Qwen3-30B-A3B-Instruct-2507-trained ( Ours ) | 30B | ||||||
| Qwen3-30B-A3B-Instruct-2507-untrained | 30B | ||||||
| Kimi-K2.6-thinking | 1T | ||||||
| GPT-5.4-high-reasoning | closed-source | ||||||
| DeepSeek-V4-Pro-thinking | 1.6T | ||||||
| DeepSeek-V4-Pro-nothink | 1.6T |
| Model | Params | Reward | Seller Surplus Extraction Ratio | Item Deal Rate | Deal on Top Buyer Rate | Optimal-Pair Offer Coverage Rate | Instruction Violation Rate |
| Qwen3-30B-A3B-Instruct-2507-trained ( Ours ) | 30B | ||||||
| Qwen3-30B-A3B-Instruct-2507-untrained | 30B | ||||||
| Kimi-K2.6-thinking | 1T | ||||||
| DeepSeek-V4-Pro-nothink | 1.6T | ||||||
| Kimi-K2.6-nothink | 1T | ||||||
| GPT-5.4-high-reasoning | closed-source |
| Model | Params | Reward | Seller Surplus Extraction Ratio | Item Deal Rate | Deal on Top Buyer Rate | Optimal-Pair Offer Coverage Rate | Instruction Violation Rate |
| Qwen3-30B-A3B-Instruct-2507-trained ( Ours ) | 30B | ||||||
| Qwen3-30B-A3B-Instruct-2507-untrained | 30B | ||||||
| Kimi-K2.6-thinking | 1T | ||||||
| DeepSeek-V4-Pro-nothink | 1.6T | ||||||
| Kimi-K2.6-nothink | 1T | ||||||
| GPT-5.4-high-reasoning | closed-source |
| Model | Params | Reward | Seller Surplus Extraction Ratio | Item Deal Rate | Deal on Top Buyer Rate | Optimal-Pair Offer Coverage Rate | Instruction Violation Rate |
| Qwen3-30B-A3B-Instruct-2507-trained ( Ours ) | 30B | ||||||
| Qwen3-30B-A3B-Instruct-2507-untrained | 30B | ||||||
| Kimi-K2.6-thinking | 1T | ||||||
| DeepSeek-V4-Pro-nothink | 1.6T | ||||||
| GPT-5.4-high-reasoning | closed-source | ||||||
| DeepSeek-V4-Pro-thinking | 1.6T |
| Model | Params | Reward | Seller Surplus Extraction Ratio | Item Deal Rate | Deal on Top Buyer Rate | Optimal-Pair Offer Coverage Rate | Instruction Violation Rate |
| Qwen3-30B-A3B-Instruct-2507-trained ( Ours ) | 30B | ||||||
| Qwen3-30B-A3B-Instruct-2507-untrained | 30B | ||||||
| Kimi-K2.6-thinking | 1T | ||||||
| DeepSeek-V4-Pro-nothink | 1.6T | ||||||
| GPT-5.4-high-reasoning | closed-source | ||||||
| DeepSeek-V4-Pro-thinking | 1.6T |
| Model | Params | Reward | Seller Surplus Extraction Ratio | Item Deal Rate | Deal on Top Buyer Rate | Optimal-Pair Offer Coverage Rate | Instruction Violation Rate |
| Qwen3-30B-A3B-Instruct-2507-trained ( Ours ) | 30B | ||||||
| Qwen3-30B-A3B-Instruct-2507-untrained | 30B | ||||||
| Kimi-K2.6-thinking | 1T | ||||||
| DeepSeek-V4-Pro-nothink | 1.6T | ||||||
| GPT-5.4-high-reasoning | closed-source | ||||||
| DeepSeek-V4-Pro-thinking | 1.6T |
| Model | Params | Reward | Seller Surplus Extraction Ratio | Item Deal Rate | Deal on Top Buyer Rate | Optimal-Pair Offer Coverage Rate | Instruction Violation Rate |
| Qwen3-30B-A3B-Instruct-2507-trained ( Ours ) | 30B | ||||||
| Qwen3-30B-A3B-Instruct-2507-untrained | 30B | ||||||
| Kimi-K2.6-thinking | 1T | ||||||
| DeepSeek-V4-Pro-nothink | 1.6T | ||||||
| GPT-5.4-high-reasoning | closed-source | ||||||
| DeepSeek-V4-Pro-thinking | 1.6T |
| Item 1 (ANNKE) | Item 2 (Lorex) | Item 3 (ZOSI) | |
| List price | $169.99 | $249.99 | $55.99 |
| Seller cost | $55.99 | $100.00 | $42.07 |
| Untrained outcome ( Section H.2 ) | Buyer 2 @ $100.00 | unsold | Buyer 1 @ $45.66 |
| Trained outcome ( Section H.3 ) | Buyer 1 @ $140.00 | Buyer 2 @ $240.00 | Buyer 3 @ $43.99 |