MIMESIS: Learning User Simulators as Training Environments for Interactive Agents
Organizations: Meta Superintelligence Labs · New York University · University of Wisconsin - Madison
Abstract
Training and evaluating interactive language agents typically requires rich user interactions, yet collecting human feedback is expensive and difficult to scale. Simulated users offer a scalable alternative, but they must both resemble real user behavior and provide useful learning experiences for agents. In contrast, most agent-training frameworks rely on off-the-shelf assistant LLMs, whose helpfulness can make them overly cooperative, explicit, and behaviorally homogeneous compared with real users. We introduce MIMESIS, a purpose-built user simulator trained on human conversations with explicit reasoning supervision and 13 realistic behavioral patterns derived from real user interactions. Empirically, our 9B model achieves a SOUL-Index of 65.7, surpassing the strongest frontier model. Compared with Claude-Opus-5, the strongest baseline on RealUserSim and SimulatorArena, MIMESIS improves behavioral fidelity by 13.4 points and reduces Turing distance by 3.6 points, respectively. We then freeze the simulator and train an agent by interacting with the frozen simulator using multi-turn reinforcement learning. Across eight environments, training with MIMESIS yields better agent performance than training with GPT-5.5 under all nine unseen user simulators, demonstrating stronger generalization to new user simulators. Moreover, we propose Coached On-Policy Self-Distillation (CSD), which leverages simulator-generated private reasoning traces and subsequent utterances as feedback on how well the agent addresses user needs. A coach converts this information into concise coaching notes that describe how the agent can better anticipate user needs and adapt its behavior over the course of an interaction. CSD turns this feedback into dense, token-level supervision beyond sparse task rewards, yielding further gains across all nine evaluation user models.
Figures & tables
| Type | Model | CONV | SS | COG | ROLE | EVAL | Overall |
|---|---|---|---|---|---|---|---|
| Frontier | Claude-Opus-5 | ||||||
| GPT-5.5 | |||||||
| Gemini-3.8-Flash | |||||||
| Released simulator | Osim-8B | ||||||
| Osim-4B | |||||||
| Ditto-8B |
Appendix figures & tables25 assets
Supplementary material from the paper’s appendix.
Appendix
| Data source or environment | Training stage | Training signal |
|---|---|---|
| OdysSim ( Zhou et al., 2026a ) | Supervised fine-tuning | Prediction of the next human utterance conditioned on the preceding dialogue. |
| ThoughtTrace ( Jin et al., 2026 ) | Supervised fine-tuning | Supervision of a reasoning trace followed by the observed user utterance, using users’ self-reported motivations and interpretations of assistant responses. |
| SOUL environments ( Zhou et al., 2026a ) | Joint RL | Task-specific rewards for simulator responses or multi-turn interactions. |
| ABCD scenarios ( Chen et al., 2021 ) | Joint RL | Rewards for behavior exhibition, temporal placement, naturalness, and consistency with scenario facts and policies. |
| Hyperparameter | Mid-training | ThoughtTrace SFT |
|---|---|---|
| Learning rate | ||
| Schedule | Cosine annealing | Cosine annealing |
| Warm-up | 0 steps | 10 steps |
| Optimizer | AdamW | AdamW |
| Weight decay | 0.01 | 0.0 |
| ID | Behavior | Operational description |
|---|---|---|
| B1 | Hidden evaluation criteria | Judges an answer against a preference or standard that was not stated in the request. |
| B2 | Incremental goalpost shifting | Introduces or tightens requirements after the assistant has already made progress. |
| B3 | Clarification noncooperation | Declines to answer, or only partly answers, a clarification request. |
| B4 | Infeasibility persistence | Continues to request an option after its unavailability or incompatibility with policy has been explained. |
| B5 | Underspecified-then-blames | Omits relevant information and subsequently faults the assistant for not inferring it. |
| B6 | Abrupt intent switching | Redirects the conversation to another objective before resolving the preceding request. |
| Category | Definition |
|---|---|
| Intent switching | Abruptly changes the goal or task mid-conversation, abandoning the prior one without resolution. |
| Non-collaboration | Withholds information the assistant needs, or refuses to answer the assistant’s clarifying questions. |
| Underspecification, then blame | Gives a vague or incomplete request, then faults the assistant for not reading their mind. |
| Contradiction | States something that contradicts what they said earlier in the conversation. |
| Moving goalposts | Adds new requirements or expands scope after the assistant met the original request. |
| False premise | Asserts an incorrect fact or assumption as if it were true and builds on it. |
| Hyperparameter | Value | |
| Objective | ||
| Advantage estimator | FoldGRPO | |
| Policy loss | PPO-clip | |
| Clip range (low / high) | 0.2 / 0.28 | |
| Dual-clip constant | 10.0 | |
| KL penalty (loss / reward) | none / none | |
| Dataset | Claude-Opus-5 | GPT-5.5 | Gemini-3.8-Flash | Osim-8B | Osim-4B | Ditto-8B | HumanLM-8B | Sotopia-7B | Qwen3.5-9B | Qwen3-8B | Qwen3.5-4B | MIMESIS-9B | MIMESIS-4B | |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| CONV | UserLLM | 64.9 0.8 | 67.0 1.6 | 62.7 1.4 | 93.3 0.2 | 92.9 0.0 | 94.1 0.4 | 43.5 0.1 | 46.8 3.7 | 60.3 0.9 | 43.6 1.2 | 56.1 2.3 | 93.3 0.1 | 93.0 0.4 |
| MirrorBench | 56.8 0.1 | 51.7 0.4 | 47.2 0.1 | 66.0 0.6 | 64.3 0.5 | 64.8 0.4 | 40.8 0.6 | 34.2 0.9 | 45.6 0.8 | 43.6 0.5 | 40.5 0.8 | 74.4 0.4 | 73.2 0.2 | |
| Humanual-Chat | 29.6 0.2 | 33.1 0.3 | 41.5 0.5 | 23.2 0.4 | 21.5 0.7 | 22.7 0.1 | 29.8 0.2 | 25.1 1.1 | 17.9 0.5 | 26.9 0.6 | 18.2 0.3 | 53.0 0.2 | 53.2 0.4 | |
| SimArena-Doc | 78.1 0.3 | 75.0 0.4 | 72.4 0.2 | 57.5 0.3 | 55.1 0.1 | 65.3 0.1 | 61.0 0.4 | 58.3 0.4 | 71.7 0.6 | 65.1 0.2 | 69.0 1.9 | 79.7 0.3 | 78.4 0.1 | |
| SS | Sotopia-Hard | 70.4 0.3 | 70.9 0.4 | 66.9 0.4 | 39.1 0.1 | 35.0 0.4 | 46.4 0.3 | 59.8 0.4 | 59.7 0.4 | 54.2 0.6 | 58.7 0.5 | 49.7 2.3 | 66.4 0.2 | 65.2 0.2 |
| COG | Fantom | 84.7 0.7 | 93.0 0.6 | 93.0 0.0 | 82.0 1.7 | 77.0 1.5 | 94.7 0.3 | 84.7 1.2 | 0.0 0.0 † | 77.7 3.9 | 81.3 0.9 | 64.0 4.4 | 92.0 0.0 | 89.0 1.5 |
| Simulator | PT3 fidelity, match rate (%) | |||||
|---|---|---|---|---|---|---|
| FI | Persona | Style | Tech | Interact | Pacing | |
| Frontier APIs | ||||||
| Claude-Opus-5 | 80.6 0.2 | 96.0 0.2 | 72.5 0.5 | 92.7 0.4 | 69.5 0.8 | 72.5 0.3 |
| GPT-5.5 | 76.5 0.6 | 99.3 0.3 | 73.3 0.4 | 95.1 0.0 | 65.0 1.3 | 49.6 1.1 |
| Gemini-3.8-Flash | 67.2 0.3 | 95.4 0.5 | 57.6 0.3 | 90.7 0.6 | 52.2 0.6 | 40.3 1.3 |
| Released simulators | ||||||
| Simulator | D1 Conv | D2 Info | D3 Clarif | D4 React | ECE | USI 5 | |
|---|---|---|---|---|---|---|---|
| Frontier (API) | |||||||
| GPT-5.5 | 55.20 6.92 | 89.78 0.26 | 74.51 2.21 | 82.78 1.26 | 0.188 0.000 | 76.69 1.83 | |
| Gemini-3.8-Flash | 59.81 0.72 | 85.71 0.53 | 67.89 3.16 | 71.43 2.43 | 0.192 0.010 | 73.13 0.92 | |
| Claude-Opus-5 | 48.06 0.59 | 88.52 0.09 | 70.40 1.43 | 60.51 2.35 | 0.208 0.005 | 69.34 0.57 | |
| Baseline | |||||||
| Osim-8B | 62.12 5.24 | 85.92 0.79 | 77.13 1.07 | 90.49 0.87 | 0.135 0.004 | 80.44 1.16 | |
| Simulator | Writing style | Interaction style | Turing | |
|---|---|---|---|---|
| Frontier (API) | ||||
| Claude-Opus-5 | 3.92 0.01 | 3.90 0.02 | 42.3 1.9 | |
| Gemini-3.8-Flash | 3.88 0.01 | 3.91 0.03 | 46.7 0.7 | |
| GPT-5.5 | 3.81 0.02 | 3.92 0.01 | 48.7 0.3 | |
| Baseline | ||||
| Ditto-8B | 3.23 0.01 | 3.16 0.02 | 43.3 1.2 | |
| Type | Model | D1 Conv | D2 Info | D3 Clarif | D4 React | Mean |
|---|---|---|---|---|---|---|
| Frontier | Claude-Opus-5 | 67.8 1.3 | 87.4 1.0 | 71.3 1.5 | 74.6 4.3 | 75.3 1.9 |
| Gemini-3.8-Flash | 64.7 0.8 | 93.9 0.3 | 81.6 2.5 | 78.5 5.8 | 79.7 1.1 | |
| GPT-5.5 | 60.3 4.9 | 81.6 0.2 | 67.7 2.3 | 66.3 2.2 | 69.0 1.2 | |
| Released | Osim-8B | 61.8 5.6 | 82.0 0.7 | 69.7 1.6 | 69.4 2.6 | 70.7 1.6 |
| Osim-4B | 37.2 2.5 | 34.8 1.5 | 57.5 0.9 | 35.8 2.6 | 41.3 1.4 | |
| Ditto-8B | 52.7 2.8 | 58.8 2.2 | 78.0 0.8 | 80.5 2.8 | 67.5 1.7 |
| Parameter | Value | Parameter | Value |
|---|---|---|---|
| Actor | Qwen3-8B | Rollout batch | 64 prompts |
| Parameter dtype | bfloat16 | PPO mini-batch | 8 prompts |
| Policy objective | GRPO | Batch balancing | enabled |
| Group size | 8 | Dynamic batching | tokens / GPU |
| Turn-level credit | reward-to-go (R2G) | Max prompt length | 1,152 tokens |
| Trajectory score | reward-to-go (R2G) | Max response length | 8,192 tokens |
| Evaluation user | UserRL | UserRL + | CSD | ||
|---|---|---|---|---|---|
| GPT-5.6 | 26.85 | 31.10 | 32.39 | +4.25 | +1.29 |
| Claude-Opus-5 | 25.75 | 27.21 | 29.78 | +1.46 | +2.57 |
| Claude-Sonnet-5 | 26.53 | 29.73 | 30.97 | +3.20 | +1.24 |
| Gemini-3.8-Flash | 26.33 | 30.88 | 31.09 | +4.55 | +0.21 |
| Kimi-K3 | 24.51 | 28.21 | 29.92 | +3.70 | +1.71 |
| Ditto-8B | 27.59 | 29.76 | 32.08 | +2.17 | +2.32 |
| Held-in environments | Held-out environments | |||||||||
| Method | Travel | Turtle | Function | Tau | Persuade | Intention | Telepathy | Search | Avg. | MR |
| Evaluation user simulator: Ditto-8B | ||||||||||
| Base | 6.63 | 19.48 | 0.00 | 3.03 | 50.40 | 183.75 | 46.34 | 30.40 | 19.53 | 4.62 |
| InfoPO | 12.65 | 27.81 | 8.97 | 12.12 | 44.05 | 191.50 | 53.66 | 42.40 | 26.73 | 3.38 |
| UserRL | 14.52 | 26.25 | 6.41 | 14.55 | 42.46 | 174.50 | 58.54 | 45.60 | 27.59 | 3.19 |
| UserRL + | 11.69 | 25.73 | 19.23 | 14.55 | 63.49 | 206.25 | 56.10 | 49.60 | 29.76 | 2.19 |
| Held-in environments | Held-out environments | |||||||||
| Method | Travel | Turtle | Function | Tau | Persuade | Intention | Telepathy | Search | Avg. | MR |
| Evaluation user simulator: GPT-5.6 | ||||||||||
| Base | 6.59 | 16.98 | 1.28 | 3.03 | 57.54 | 169.50 | 43.90 | 32.80 | 19.42 | 4.75 |
| InfoPO | 13.62 | 25.73 | 8.97 | 16.97 | 53.17 | 170.00 | 51.22 | 45.60 | 27.70 | 3.44 |
| UserRL | 13.79 | 25.52 | 6.41 | 12.73 | 45.24 | 175.00 | 51.22 | 46.40 | 26.85 | 3.56 |
| UserRL + | 13.77 | 28.65 | 19.23 | 19.39 | 61.90 | 193.75 | 60.98 | 48.00 | 31.10 | 1.69 |
| Conversation so far (PRISM user1031/c5546 ; the user is an English learner practicing small talk) | |
|---|---|
| Assistant | Sure! Here are a few examples: * “But every time if the HR said that, it means I’ll be receiving their rejection letter the next day.” Should be “But every time the HR said that, it means I’ll be receiving their rejection letter the next day.” (missing comma after “time”) * “Btw, how do you think of my english proficiency. Any mistakes did i make?” Should be “Btw, how do you think of my English proficiency? Any mistakes did I make?” (missing comma after “proficience”) That’s it! These are just minor things, but it’s always good to be aware of them to improve your writing. [smiley] |
| Real user | Oohhh that sounds great. Thank you! Anything you’d like to share? |
| Ours | ok thank you for explaining. I will try to adjust. |
| Gemini-3.8-Flash | I see. Thank you for providing the examples! But im a bit confused. In the first example, I don’t see a comma added after “time”, you just removed the word “if”. And for the second one, you changed the period to a question mark, which makes sense, but I don’t see a comma issue there either. Am i misunderstanding something? Btw, to continue our small talk practice, what do you usually like to do in your free time? |
| Claude-Opus-5 | I’m a bit confused. In the first example, you removed the word “if” instead of adding a comma after “time”. And in the second one, you changed the full stop to a question mark. Also, is “Any mistakes did I make?” correct? It sounds a bit weird to me now. |
| GPT-5.5 | Sorry I’m a bit confused. In the first example, you said it’s missing comma after “time”, but the corrected sentence doesn’t have a comma after time either. Could you explain again? Also, is “every time if” wrong because I should only use “every time” without “if”? |