RLTL;DR: Self-improvement by Internalizing Self-generated Feedback
Organizations: Apple
Abstract
The common paradigm of reinforcement learning with verifiable rewards (RLVR) is to let agents make multiple attempts at a task, and optimize towards the successful ones. This becomes problematic in the realms of self-improvement, where tasks are so difficult that the agent has a low or even no chance of success, and where there are no teacher models or example solutions to distill from. In this paper, we introduce RLTL;DR. After each failed attempt, we show the policy the verifier outputs and let it write its own feedback, in the form of a single TL;DR insight. The next rollout is conditioned on all previous insights, and we sequentially sample rollouts until a solution is found. Moreover, we enable backpropagation on the in-context insights to internalize a direct task to insight mapping. On challenging tool-calling and coding datasets (filtered to Pass@128=0), standard GRPO training of a Qwen 3.5 9B Thinking policy stays flat at a Pass@1 of 0% to 1%. RLTL;DR breaks through this learning barrier, achieving a Pass@1 of 14-31% with insights in context during training and, crucially, 12-13% when no insight is in context at eval time. We identify that the key is the task to insight internalization. To study this further, we reduce our approach to SFTL;DR, training only on (task, insight) tuples, without showing or backpropagating on any rollouts. Training on only 4k of these tuples recovers almost the full performance of RLTL;DR and classical SFT on full rollouts. This demonstrates a promising compacted training paradigm of the form "on this sort of task, keep this sort of thing in mind", which we hope to inspire future research on.
Figures & tables
| SAPI | Appworld | Leetcode | |
| GRPO | 57.8% | 33.7% | 55.1% |
| SGE | 57.7% | 33.5% | 56.4% |
| RLTF-SD | 53.0% | 38.9% | 42.2% |
| RLTL;DR | 91.1% | 61.7% | 49.1% |
| Forward Tokens | Backward Tokens | Train Pass@1 | Eval Pass@1 | |
| RLTL;DR | M | M | 21.5% | 18.9% |
| RLTL;DR, | M | M | 9.3% | 14.0% |
| RLTL;DR, | M | M | 6.1% | 13.2% |
| RLTL;DR, no GRPO loss, only | M | k | 20.0% | 17.0% |
| SFT on all 3.5k full rollouts | M | M | 28.0% | 20.9% |
| SFT on 1k rollouts | M | M | 27.1% | 21.0% |
| Setting | Train Pass@1 | Eval Pass@1 | ||
| Default (TL;DR format, self-generated with thinking) | 16 | 21.4% | 19.2% | |
| Insight content | Diagnostic paragraph | 16 | 20.3% | 18.6% |
| Summary + Diagnostic paragraph | 16 | 13.8% | 16.0% | |
| Summary + Diagnostic paragraph + Corrected code | 16 | 15.7% | 17.3% | |
| Insight generator | Student, non-thinking | 16 | 20.1% | 19.8% |
| Student, no information about failed unit tests (only success) | 16 | 0 3.4% | 10.0% |
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
| Appworld/SAPI | Leetcode | |
| Rollout collection phase | ||
| Unique tasks per rollout collection phase | 8 | 128 |
| Min attempts per task | 8 | 8 |
| Avg attempts per task (due to async collection) | 21 | 26 |
| Max steps per attempt | 50 | 1 |
| Max think tokens per step | 768 | 0 |
| Share | Correction | Examples of insight |
| 15.3% | Wrong capitalization | • Use the uppercase Orange color enum instead of lowercase orange when updating the list color. • Use uppercase PM instead of lowercase pm when setting the alarm’s AM/PM marker. • Use the Priority enum value instead of a string when setting the reminder priority. |
| 11.8% | Unrecorded action | • You need to mark the article as read rather than just viewing it. • Use the unified search API instead of the ticker search to create a history record that contains the query text. • You need to create a download event for the mobile listing instead of using the generic task completion function. |
| 11.2% | Extra parameter | • Call complete_task without passing an answer parameter for non-QA tasks. |
| 5.5% | Unsaved setting | • You need to set the default chart horizon preference instead of just viewing a chart. • The list sorting preference needs to be saved to the database, not just the sort configuration updated. |
| 2.1% | Over-filtering | • Remove the route_type filter to see all hiking trails instead of getting no results. |
| 1.2% | Argument position | • Pass the access token as the first positional argument, not as a keyword argument. |
| Failure type: answer argument passed to a non-QA task | |
| Diagnostic paragraph | “The task failed because the agent called complete_task() with an answer parameter, but non-QA tasks should not include an answer. The verifier only checks that the BPM filter was set to 125–125, the playlist search results were recorded, and the embed code exists in the database — not the submission method.” |
| Corrected code | apis.supervisor.complete_task() |
| TL;DR insight | “Remove the answer parameter from the complete task call.” |
| Failure type: item not persisted to saved for later | |
| Diagnostic paragraph | “The rollout failed because the product was not properly persisted to the saved_for_later collection as verified by the database. The agent called save_for_later with only product_id and note , but did not pass the access_token parameter which is required for authentication. The verifier checks the database directly via show_saved_for_later(db) , not API responses.” |
| Corrected code | apis.shop.save_for_later(product_id=1, |
| Method | Train loop | bf16 precision | Gini( ) | top-1% share of | ||
| RLTL;DR | RL | 7.4 | 0.79 | 2.11% | 0.99 | 96.88% |
| RLTL;DR, | RL | 3.2 | 0.44 | 1.58% | 0.99 | 99.66% |
| RLTL;DR, | RL | 2.4 | 0.34 | 1.58% | 0.99 | 99.67% |
| RLTL;DR, no GRPO loss, only | RL | 9.3 | 0.89 | 2.51% | 0.99 | 94.07% |
| SFT on all 3.5k full rollouts | SFT | 1.3 | 36.12 | 30.42% | 0.87 | 35.47% |
| SFT on 1k rollouts | SFT | 0.8 | 22.72 | 24.86% | 0.88 | 36.33% |
| Setting | No-insight success (%) |
| RLTF-FM (insight generator SFT, all insight) | 6.2 |
| Self-play insight SFT (insight generator SFT, effective insight only) 1 1 1 This run’s training crashed at update 75, so its value is averaged over the 75 available updates rather than the full 80. | 7.7 |
| RLTL;DR (ours) | 14.8 |
| RLTL;DR + RLTF-FM | 15.0 |
| RLTL;DR + Self-play insight SFT | 15.7 |
| Ablation | Eval Pass@1 |
| Insert insights if running success rate, | 14.2% |
| Insert insights if running success rate, | 15.3% |
| Insert insights if running success rate, | 17.1% |
| Insert insights if 3 of first 10 rollouts successful, | 19.2% |
| Insert if running success rate, | 19.6% |
| Split Advantages | 18.9% |