Human-in-the-Loop Annotation

Latest papers 95

Mar 22, 2026cs.CL

Multi-Perspective LLM Annotations for Valid Analyses in Subjective Tasks

Large language models are increasingly used to annotate texts, but their outputs reflect some human perspectives better than others. Existing methods for correcting LLM annotation error assume a single ground truth. However, this assumption fails in subjective tasks where disagreement across demographic groups is meaningful. Here we introduce Perspective-Driven Inference, a method that treats the distribution of annotations across groups as the quantity of interest, and estimates it using a small human annotation budget. We contribute an adaptive sampling strategy that concentrates human annotation effort on groups where LLM proxies are least accurate. We evaluate on politeness and offensiveness rating tasks, showing targeted improvements for harder-to-model demographic groups relative to uniform sampling baselines, while maintaining coverage.
Feb 23, 2026cs.SD

AuditoryHuM: Auditory Scene Label Generation and Clustering using Human-MLLM Collaboration

Manual annotation and clustering of audio datasets is labour intensive. We introduce AuditoryHuM, a training-free framework for the unsupervised discovery and clustering of auditory scene labels using human-Multimodal Large Language Model (MLLM) collaboration. Leveraging MLLMs, our framework generates contextually relevant labels for audio data. To ensure label quality and mitigate hallucinations, zero-shot learning (Human-CLAP) quantifies the alignment between generated text labels and raw audio. A targeted human-in-the-loop intervention, refines only the lowest aligned pairs. The discovered labels form an interpretable alignment vector to group audio into cohesive clusters. The framework was evaluated across three auditory scene datasets (ADVANCE, AHEAD-DS, and TAU 2019), achieving a 96.1% reduction in human labour for ADVANCE during testing. Downstream models trained on our clusters exhibit enhanced classification accuracy vs the baseline (0.79 to 0.84) in clean acoustic environments, though performance scales down in dense, complex soundscapes. The project page: https://github.com/Australian-Future-Hearing-Initiative
Dec 14, 2025cs.LG

Active Learning with Imperfect Labels: Optimal Labeler Assignment and Sample Selection

Active Learning (AL) is commonly used in applications where labeling data is expensive or time-consuming. In practice, however, labels are often noisy due to varying labeler expertise and annotation uncertainty, especially for complex or ambiguous samples. Learning from such imperfectly labeled data can degrade classifier performance. We propose an AL framework that explicitly accounts for label noise by optimally assigning labelers and selecting samples to minimize labeling error. Our approach, called OLAS (Optimal Labeler Assignment and Sampling), uses a noise model that depends on both labeler accuracy and model uncertainty to guide these decisions. We develop two tractable optimization formulations: one for assigning samples to labelers to minimize worst-case noise, and another for selecting samples while controlling overall label noise. Theoretical results provide closed-form solutions under mild conditions. Empirical evaluations on benchmark datasets and a real-world warranty claim classification problem show that OLAS achieves the highest or near-highest classification accuracy among existing AL strategies across most settings, using only a single label per sample.
Sep 22, 2025cs.CL

Interactive In-Meeting Speaker Correction with Human Feedback

Most automatic speech processing systems operate in ``open loop'' mode without user feedback about who said what, yet human-in-the-loop workflows can potentially enable higher accuracy. We propose an LLM-assisted in-meeting speaker correction system that lets users fix speaker attribution errors through brief corrective feedback. After performing streaming ASR and diarization, the system presents concise LLM-generated summaries to help users identify important speaker errors, and it incorporates user feedback by updating the speaker-attributed transcript and adding online speaker enrollments. To make this workflow effective despite errors in speech processing, LLM analysis, and user feedback, we developed several mechanisms to identify the intended correction more precisely. Further, we built an LLM-driven user feedback simulation to evaluate the workflow reprodubilty and at scale. Applied to the AMI headset test set, our system substantially reduces the DER from a streaming baseline (Google ASR + ECAPA) by 31.99% and speaker substitution error by 52.68%. Results of a pilot usability study suggest several avenues to improve the user experience.
Date pendingcs.CL

Decomposing LLM-Judge Uncertainty to Target Expert Labels

An LLM judge evaluates outputs at scale. Experts should label only where it is least sure. Its natural escalation signal conflates two uncertainties: aleatoric, real disagreement in the expert pool, which labels cannot reduce, and epistemic, the judge's ignorance, which labels do reduce. A small Bayesian model separates them: a regression on labels already collected learns how far to trust a black-box judge's prediction. Both components follow as simple formulas, with no sampling or further judge calls. The components isolate on a real LLM judge against exactly known truth, and stated confidence is no guide to its actual error. On real human disagreement (ChaosNLI) the epistemic ranking removes 83% more error than total uncertainty for the same expert labels, though simply escalating the least-labelled items does as well there. We demonstrate we can estimate where a judge is ignorant rather than where experts genuinely disagree, and propose using this to direct expert labelling. Code and data are available at https://github.com/composo-ai/judge-uncertainty-decomposition.