We present onPanda, an interactive tool for efficiently annotating LLM alignment data and agent trajectories. onPanda adopts token-level correction as its core interaction: while reading a model response, the annotator locates the first inappropriate token and either picks a substitute from the model's candidate tokens or types the correct text via free-form editing. The system then truncates everything after that position and continues generation from the corrected prefix, repeating this locate-correct-continue loop until a satisfactory response is obtained. This mechanism lets annotators precisely steer model outputs at low cost: a small controlled study suggests that onPanda reduces median annotation time by 52% over manual post-editing. Since the vast majority of tokens in the final response are generated by the model itself, the resulting data largely preserves the model's sampling distribution and is well suited for constructing on-policy SFT and preference data. Furthermore, the token-level corrections recorded during annotation provide fine-grained supervision with precise positions and naturally paired positive--negative samples. onPanda also connects to external tools and harnesses, enabling interactive trajectory annotation in realistic environments. In addition, we release Panda-CVL, a dataset annotated with onPanda, together with a benchmark for token-level correction.
Figures & tables
Figure 1: The token-level correction interface of onPanda. (a) Each token’s generation probability is color-coded below it: greener means higher, redder means lower. (b) Hovering over the first inappropriate token ␣triangles pops up the top-20 candidate tokens at this position; the annotator selects the more appropriate ␣shapes . (c) Replace and continue: the system replaces ␣triangles with ␣shapes (highlighted with a red background), truncates the subsequent content, and lets the model continue generating from the corrected prefix. (d) Free-form editing: when no candidate fits, the annotator double-clicks the erroneous ␣paralle to open an input box and replaces it with the correct ␣trapezoid . (e) The model continues generating after ␣trapezoid , finally yielding a response that meets the SFT quality bar ( is_good=Y ). (f) The session preserves all intermediate versions: (a), (c), and (e) correspond to Dialogs [1, 2, 3]; the final Dialog 3 is the positive sample ( is_good=Y ), while Dialogs 1 and 2 are retained as negatives.
Figure 2: Annotating an agent trajectory with onPanda. Top left: the user prompt; top right: the connected MCP servers and tools. Left: the assistant response is rendered by the response template into the model’s native token stream, where special tokens (e.g., </think> ) are directly visible and correctable; the annotator corrects tokens in the reasoning and in the tool-call arguments, then approves execution. Right: the modified token stream is parsed in real time into reasoning and tool_calls , displayed separately, with the returned tool results shown at the bottom right.
Tool
Paradigm
Time (s) ↓
Win rate ↑
PPL
Δ PPL
SFT Cov. ↑
Pref. pairs ↑
NASA-TLX ↓
Argilla
rank 4 candidates
336 (684)
28.6%
1.161
- 0.83%
52% (11)
6.00
5.4
POTATO
post-editing
681 (711)
54.8%
1.596
+ 36.31%
100% (21)
0.95
6.8
onPanda
token-level correction
330 ( 516 )
66.7%
1.181
+ 0.86%
100% (21)
7.43
3.1
Table 1: Comparison of annotation paradigms in the controlled study. Time is reported in seconds as the median annotation time over the 21 prompts, with the mean in parentheses. Win rate is the fraction of LLM-judged pairwise comparisons each method wins against the other two. PPL is scored by the rollout model; Δ PPL is the relative change from the re-sampling baseline (1.171); smaller absolute values indicate higher on-policy fidelity. SFT Cov. is the fraction of prompts yielding a valid SFT response, and Pref. pairs is the mean number of derivable preference pairs per prompt. NASA-TLX, from the user study, is on our adapted 0–10 scale (lower is better).
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Metric
Vision
Audio
Agentic
Scale
Annotation sessions
25,596
105,143
1,257
Annotators
32
42
24
Total human-hours
3,557.4
18,871.3
523.5
Annotation time P25 (min) †
0.64
0.95
17.82
Annotation time P50 (min) †
1.65
1.93
31.31
Appendix
Table 2: Production annotation statistics for the Vision, Audio, and Agentic onPanda deployments. † denotes statistics reported per annotation session. Some projects require annotating multiple SFT instances from a single session to increase data throughput. Furthermore, a single agent trajectory contains multi-turn tool calls that can serve as SFT data. Due to the nascent state of agentic annotation and unoptimized rollout models, stricter control is required, leading to increased manual typing.
Model
Format
GoodAcc
Loc.-NG
Corr.-NG
F1
Reasoning
Doubao-Seed-2.1-pro ByteDance Seed (2026)
99.01
9.66
21.64
13.70
11.33
GPT-5.5 (xhigh) OpenAI (2026a)
93.96
53.37
15.26
10.18
17.09
GPT-5.6-sol (xhigh) OpenAI (2026b)
99.98
16.56
19.95
13.09
14.63
GPT-6 (xhigh) OpenAI (2026c)
99.65
17.38
24.46
15.83
16.57
Kimi-K2.6 Moonshot AI (2026)
92.41
26.40
18.47
10.59
15.11
Appendix
Table 3: Scores on the Panda-CVL test set (%). The best result in each column is shown in bold.
To align a Large Language Model (LLM), most existing methods collect explicit human feedback and train a reward model to predict the human preference based on the response text. These existing methods have two key limitations. First, the users rarely provide explicit feedback for LLM responses, which makes the high-quality preference annotation expensive to collect. Second, the methods do not leverage implicit human feedback, which has proven vital to the economic moats of Internet giants. To quantify the value of implicit feedback, we build a new dataset called IFLLM, which collects 1336 multi-turn questions from the 59 Mechanical Turk workers, their mouse trajectories, and eye gazing points to the LLMs' responses from their webcams. IFLLM shows that the users have very diverse types of gazing behavior and mouse trajectories. Our reward model based on the implicit user feedback boosts the accuracy of the text-based reward model from 55% to 64% and nearly triples the relative response quality improvements after applying the DPO to eight LLMs, demonstrating the value of implicit feedback in the wild. Our data collection website, dataset, and codes can be found at https://github.com/themehulpatwari/llm-implicit-feedback/.
Haw-Shiuan Chang, Jeffrey Gomez, Mehul Patwari +2
University of Massachusetts, Amherst, USA · York University, Canada
Large Language Models (LLMs) are increasingly used for zero-shot annotation and LLM-as-a-judge tasks, yet their reliability hinges on how model-internalized priors interact with user-provided instructions. We investigate three dimensions of this interaction: (1) how an LLM's familiarity with data and task definitions affects performance, (2) the extent to which additional information in prompts can correct zero-shot errors ("decision stickiness"), and (3) model susceptibility to misaligned task definitions. Through experiments on toxicity detection across diverse datasets (spanning social media, gaming, news, and forums) using both dense and mixture-of-experts models, we find that nearly two-thirds of zero-shot errors are resistant to correction, with an overall rescue rate (fraction of initial errors corrected by prompting) of only 34.8%. High-confidence errors prove especially resistant to correction. When given misaligned definitions, LLMs follow them while maintaining confidence levels unchanged from the aligned condition. Crucially, we introduce Definition-Specific Familiarity (DSF), which measures alignment between a model's internal concept and the task definition. After controlling for dataset-level confounds, DSF shows a positive association with model performance (partial r = +0.41), while three distinct memorization metrics (ROUGE-L, BERTScore, and embedding cosine similarity) all fail to show a positive association. These findings show the limitations of prompt-based correction in annotation tasks, highlighting the importance of definition alignment over text-level memorization.
Etienne Casanova, Rafal Kocielnik, R. Michael Alvarez
California Institute of Technology, Pasadena, CA, USA.
As large language models (LLMs) are increasingly deployed in real-world applications, alignment is no longer governed by a single universal notion of safety or helpfulness, but instead by provider- or application-specific model specifications. These specifications are typically long, structured, and frequently updated, yet existing alignment pipelines lack a systematic mechanism to operationalize them as training signals. In this paper, we propose specification-grounded alignment, a new alignment paradigm that treats provider-authored model specifications as the primary alignment target rather than abstract principles or static benchmarks. To instantiate this paradigm, we introduce SpecAlign, a framework that synthesizes alignment data directly from specification documents. SpecAlign combines structured rule annotation, controllable specification instantiation, and multi-agent adversarial data synthesis to generate fine-grained, boundary-aware preference pairs that capture both compliant behaviors and meaningful specification violations. Experiments across multiple model specifications and backbone models demonstrate that training with SpecAlign consistently improves rule compliance while preserving general capabilities and avoiding over-conservative behavior. These results suggest that grounding alignment in explicit model specifications enables rapid, precise, and scalable adaptation of LLM behavior to evolving policy requirements.
Wenjie Wang, Yue Huang, Zhengqing Yuan +6
University of Notre Dame · Carnegie Mellon University · LMU Munich +1