Human-AI Collaboration

Recent momentum

emerging

0 papers in the last 28 days · 0.0% of indexed attention

Twelve weeks of publication activity for this topic as it is defined today.

Weekly history

Recent digests

What was published in this field, kept on the site without email delivery.

Period ending 2026-09-21

24 new papers

A weekly snapshot of new work published in Human-AI Collaboration.

Period ending 2026-09-14

8 new papers

A weekly snapshot of new work published in Human-AI Collaboration.

Period ending 2026-09-07

24 new papers

A weekly snapshot of new work published in Human-AI Collaboration.

Inside this field

Focused directions

577 papers

Latest in Human-AI Collaboration

Feb 19, 2026cs.CL

Modeling Distinct Human Interaction in Web Agents

Despite rapid progress in autonomous web agents, human involvement remains essential for shaping preferences and correcting agent behavior as tasks unfold. However, current agentic systems lack a principled understanding of when and why humans intervene, often proceeding autonomously past critical decision points or requesting unnecessary confirmation. In this work, we introduce the task of modeling human intervention to support collaborative web task execution. We collect CowCorpus, a dataset of 400 real-user web navigation trajectories containing over 4,200 interleaved human and agent actions. We identify four distinct patterns of user interaction with agents -- hands-off supervision, hands-on oversight, collaborative task-solving, and full user takeover. Leveraging these insights, we train language models (LMs) to anticipate when users are likely to intervene based on their interaction styles, yielding a 61.4-63.4% improvement in intervention prediction accuracy over base LMs. Finally, we deploy these intervention-aware models in live web navigation agents and evaluate them in a user study, finding a 36.8% increase in user-rated agent usefulness. Together, our results show structured modeling of human intervention leads to more adaptive, collaborative agents.
Faria Huq, Zora Zhiruo Wang, Zhanqiu Guo +6
Feb 12, 2026cs.HC

Understanding Persuasive Interactions between Generative Social Agents and Humans: The Knowledge-based Persuasion Model (KPM)

Generative social agents (GSAs) use artificial intelligence to autonomously communicate with human users in a natural and adaptive manner. Currently, there is a lack of theorizing regarding interactions with GSAs, and likewise, few guidelines exist for studying how they influence user attitudes and behaviors. Consequently, we propose the Knowledge-based Persuasion Model (KPM) as a novel theoretical framework. According to the KPM, a GSA's self-, user-, and context-knowledge drives its persuasive behavior, which in turn shapes attitudes and behaviors of a responding human user. Building on existing research, the model offers a structured approach to studying interactions with GSAs, supporting the development of agents that motivate rather than manipulate human users. Accordingly, the KPM encourages the integration of responsible GSAs that adhere to social norms and ethical standards with the goal of increasing user wellbeing. A preliminary evaluation study, as well as implications of the KPM for research and application domains such as healthcare and education are discussed.
Stephan Vonschallen, Friederike Eyssel, Theresa Schmiedel
Feb 2, 2026cs.RO

PokeNet: Learning Kinematic Models of Articulated Objects from Human Observations

Articulation modeling enables robots to learn joint parameters of articulated objects for effective manipulation which can then be used downstream for skill learning or planning. Existing approaches often rely on prior knowledge about the objects, such as the number or type of joints. Some of these approaches also fail to recover occluded joints that are only revealed during interaction. Others require large numbers of multi-view images for every object, which is impractical in real-world settings. Furthermore, prior works neglect the order of manipulations, which is essential for many multi-DoF objects where one joint must be operated before another, such as a dishwasher. We introduce PokeNet, an end-to-end framework that estimates articulation models from a single human demonstration without prior object knowledge. Given a sequence of point cloud observations of a human manipulating an unknown object, PokeNet predicts joint parameters, infers manipulation order, and tracks joint states over time. PokeNet outperforms existing state-of-the-art methods, improving joint axis and state estimation accuracy by an average of over 27% across diverse objects, including novel and unseen categories. We demonstrate these gains in both simulation and real-world environments.
Anmol Gupta, Weiwei Gu, Omkar Patil +2
Jan 21, 2026cs.AI

Agentic AI for Commercial Insurance Underwriting with Adversarial Self-Critique

Commercial insurance underwriting is a labor-intensive process that requires manual review of extensive documentation to assess risk and determine policy pricing. While AI offers substantial efficiency improvements, existing solutions lack comprehensive reasoning and internal mechanisms to ensure reliability in regulated, high-stakes environments. Full automation remains impractical and inadvisable when human judgment and accountability are critical. This study presents a decision-negative, human-in-the-loop agentic system that incorporates an adversarial self-critique mechanism as a bounded safety architecture for regulated underwriting workflows. In this system, a critic agent challenges the primary agent's conclusions prior to submitting recommendations to human reviewers. This internal system of checks and balances addresses a critical gap in AI safety for regulated workflows. Additionally, the research develops a formal taxonomy of failure modes to characterize potential errors by decision-negative agents. This taxonomy provides a structured framework for risk identification and management in high-stakes applications. Experimental evaluation using 500 expert-validated underwriting cases demonstrates that the adversarial critique mechanism reduces AI hallucination rates from 11.3% to 3.8% and increases decision accuracy from 92% to 96%. At the same time, the framework enforces strict human authority over all binding decisions by design. These findings indicate that adversarial self-critique supports safer AI deployment in regulated domains and offers a model for responsible integration where human oversight is indispensable.
Joyjit Roy, Samaresh Kumar Singh
Jan 20, 2026cs.CV

On the Role of Rotation Equivariance in Monocular 2D-to-3D Human Pose Lifting

We consider monocular 3D human pose estimation (HPE), where the goal is to predict 3D human skeletal joints from a single 2D image, typically via 2D keypoint detection followed by 2D-to-3D lifting. Despite their success, we find that current lifting models exhibit strong performance degradation under rotations. We systematically study rotation equivariance across lifting architectures with increasing degrees of equivariant inductive bias, and investigate whether equivariance should be enforced by architectural design or learned from data, given the inherent ambiguity of monocular 2D-to-3D lifting. Utilising common HPE benchmarks, we demonstrate that rotation equivariance can be effectively learned via rotation-based data augmentation applied jointly to input and output poses. Compared with non-augmented models, this reduces error on rotated poses by over 30% across the benchmarks, with reductions of up to 72%. Moreover, augmentation-based models achieve up to 15% lower error than methods that are fully equivariant by design, while providing up to 37×\times faster inference. We further show that these findings generalise to real-world human pose sequences involving full-body rotations and to a state-of-the-art HPE model.
Pavlo Melnyk, Cuong Le, Urs Waldmann +2
Dec 1, 2025cs.CV

Generative Action Tell-Tales: Assessing Human Motion in Synthesized Videos

Despite rapid advances in video generative models, robust metrics for evaluating visual and temporal correctness of complex human actions remain elusive. Critically, existing pure-vision encoders and Multimodal Large Language Models (MLLMs) are strongly appearance-biased, lack temporal understanding, and thus struggle to discern intricate motion dynamics and anatomical implausibilities in generated videos. We tackle this gap by introducing a novel evaluation metric derived from a learned latent space of real-world human actions. Our method first captures the nuances, constraints, and temporal smoothness of real-world motion by fusing appearance-agnostic human skeletal geometry features with appearance-based features. We posit that this combined feature space provides a robust representation of action plausibility. Given a generated video, our metric quantifies its action quality by measuring the distance between its underlying representations and this learned real-world action distribution. For rigorous validation, we develop a new multi-faceted benchmark specifically designed to probe temporally challenging aspects of human action fidelity. Through extensive experiments, we show that our metric achieves substantial improvement of more than 68% compared to existing state-of-the-art methods on our benchmark, performs competitively on established external benchmarks, and has a stronger correlation with human perception. Our in-depth analysis reveals critical limitations in current video generative models and establishes a new standard for advanced research in video generation.
Xavier Thomas, Youngsun Lim, Ananya Srinivasan +2
Nov 24, 2025cs.CY

Towards Synergistic Teacher-AI Interactions with Generative Artificial Intelligence

Generative artificial intelligence (GenAI) is increasingly used in education, posing significant challenges for teachers adapting to these changes. GenAI offers unprecedented opportunities for accessibility, scalability and productivity in educational tasks. However, the automation of teaching tasks through GenAI raises concerns about reduced teacher agency, potential cognitive atrophy, and the broader deprofessionalisation of teaching. Drawing findings from prior literature on AI in Education, and refining through a recent systematic literature review, this chapter presents a conceptualisation of five levels of teacher-AI teaming: transactional, situational, operational, praxical and synergistic teaming. The framework aims to capture the nuanced dynamics of teacher-AI interactions, particularly with GenAI, that may lead to the replacement, complementarity, or augmentation of teachers' competences and professional practice. GenAI technological affordances required in supporting teaming, along with empirical studies, are discussed. Drawing on empirical observations, we outline a future vision that moves beyond individual teacher agency toward collaborative decision-making between teachers and AI, in which both agents engage in negotiation, constructive challenge, and co-reasoning that enhance each other's capabilities and enable outcomes neither could realise independently. Further discussion of socio-technical factors beyond teacher-AI teaming is also included to streamline the synergy of teachers and AI in education ethically and practically.
Mutlu Cukurova, Wannapon Suraworachet, Qi Zhou +1
Nov 6, 2025cs.AI

Large language models replicate and predict human cooperation across experiments in game theory

Large language models (LLMs) are increasingly deployed as decision-making agents in high-stakes domains and as imitators of human behavior in the social and behavioral sciences. Yet how closely LLMs mirror human decision-making remains poorly understood. This gap is critical: misalignment could produce harmful outcomes in practice, while failure to replicate human behavior renders LLMs ineffective as social simulators. Here, we address this gap by replicating large-scale game-theoretic experiments and by introducing a systematic prompting and probing framework for machine-behavioral evaluation. We test three open models typically used to power agents (Llama, Mistral, and Qwen). Across 121 dyadic games spanning four classical game types, Llama reproduces human cooperation patterns with high fidelity, while Qwen aligns closely with Nash equilibrium predictions. Characterizing models through behavioral phenotyping, we find that humans and Llama share an envious decision profile, while Qwen and Mistral exhibit different profiles. An attention-based analysis of payoff salience reveals Llama processes payoff information in a structured, layer-dependent manner absent in Qwen and Mistral, suggesting a mechanistic basis for its closer alignment with human behavior. Population-level behavioral replication is achieved without persona-based prompting, simplifying the simulation process. Extending the experimental parameter space beyond the original human-tested games, we generate and preregister testable hypotheses for novel game configurations. Our findings demonstrate appropriately configured LLMs can replicate aggregate human behavioral patterns, exhibit human-like decision phenotypes, and enable systematic exploration of unexplored experimental spaces, offering a complementary approach to traditional behavioral research that generates new empirical predictions about human social decision-making.
Andrea Cera Palatsi, Samuel Martin-Gutierrez, Ana S. Cardenal +1
Oct 30, 2025cs.AI

Human-AI Complementarity: A Goal for Amplified Oversight

Human feedback is critical for aligning AI systems to human values. As AI capabilities improve and AI is used to tackle more challenging tasks, verifying quality and safety becomes increasingly challenging. This paper explores how we can leverage AI to improve the quality of human oversight. We focus on an important safety problem that is already challenging for humans: fact-verification of AI outputs. We find that combining AI ratings and human ratings based on AI rater confidence is better than relying on either alone. Giving humans an AI fact-verification assistant further improves their accuracy, but the type of assistance matters. Displaying AI explanation, confidence, and labels leads to over-reliance, but just showing search results and evidence fosters more appropriate trust. These results have implications for Amplified Oversight -- the challenge of combining humans and AI to supervise AI systems even as they surpass human expert performance.
Rishub Jain, Sophie Bridgers, Lili Janzer +3
Oct 29, 2025q-bio.NC

Gravity-Awareness: Deep Learning Models and LLM Simulation of Human Awareness in Altered Gravity

Earth s gravity fundamentally shapes human behaviour. The brain encodes this force as an internal model of gravity, enabling the prediction and interpretation of gravitational effects during perception and action. Understanding how this model adapts to altered gravity is critical for predicting human performance in spaceflight. We present a computational framework for modelling neurophysiological adaptation across diverse gravitational environments. The framework has two components trained on open-access data from altered-gravity studies, particularly parabolic flights. The first component (CorticalG) employs a lightweight multilayer perceptron neural network to predict gravity-dependent changes in EEG frequency bands, estimating cortical state under different gravitational loads. The second component (PhysioG) uses independent Gaussian process models to capture broader physiological responses, including heart rate variability, electrodermal activity, and motor control. To complement the quantitative modelling, we simulated subjective experience across gravitational environments using the Large Language Model (LLM) Claude 3.5 Sonnet. Physiological outputs prompted the model to generate narratives describing alertness, bodily awareness, and cognitive state across zero gravity, partial gravity of the Moon and Mars, and hypergravity. This framework provides a novel approach for investigating human adaptation to spaceflight. It offers a predictive tool to assess performance and resilience, supporting the design of future space exploration missions.
Bakytzhan Alibekov, Alina Gutoreva, Elisa Raffaella-Ferre
Oct 29, 2025cs.CY

Human Resilience in the AI Era -- What Machines Can't Replace

AI is changing work and decision making faster than many institutions can adapt their operating practices. We argue that this adaptation gap makes human resilience a core capability for the AI era. We define resilience as the capacity to absorb disruption while preserving effective action and human agency around core purposes. The framework operates at three interacting levels. Psychological resilience keeps a person goal-directed under stress. Social resilience makes trusted support and correction available across a group. Organizational resilience turns detected problems into learning and recovery. We connect established resilience and technostress research with direct AI-in-the-loop experiments. General resilience is trainable, while AI-specific causal evidence is still emerging. Direct AI studies show that assistance can raise productivity and spread expertise. Other experiments show improved expressed empathy and more calibrated reliance. We translate these findings into a practical agenda for AI education, workplace design, governance, and evaluation. The central proposal is socio-technical: structural safeguards define the operating boundary, while resilient people and institutions provide adaptive capacity when conditions change.
Shaoshan Liu, Anina Schwarzenbach, Yiyu Shi
Oct 28, 2025cs.CV

IBIS: A Hybrid Inception-BiLSTM and SVM Ensemble for Robust Doppler-based Human Activity Recognition

Wi-Fi sensing is a leading technology for Human Activity Recognition (HAR), offering a non-intrusive and cost-effective solution for healthcare and smart environments. Despite its potential, existing methods struggle with domain shift issues, often failing to generalize to unseen environments due to overfitting. This paper proposes IBIS, a robust ensemble framework combining Inception-Bidirectional Long Short-Term Memory (BiLSTM) for feature extraction and Support Vector Machine (SVM) for classification of Doppler signatures. The proposed architecture specifically targets generalization capabilities. Experimental results on multiple datasets show that IBIS achieves 95.40% accuracy, delivering a 7.58% performance gain compared to standard architectures in cross-scenario evaluations on external datasets. The analysis confirms that IBIS effectively mitigates environmental dependency in Wi-Fi-based HAR.
Alison M. Fernandes, Hermes I. Del Monego, Bruno S. Chang +3
Oct 22, 2025cs.SE

Human-Agent Collaborative Paper-to-Page Crafting

In the quest for scientific progress, communicating research is as vital as the discovery itself. Yet, researchers are often sidetracked by the manual, repetitive chore of building project webpages to make their dense papers accessible. While automation has tackled static slides and posters, the dynamic, interactive nature of webpages has remained an unaddressed challenge. To bridge this gap, we reframe the problem, arguing that the solution lies not in a single command, but in a collaborative, hierarchical process. We introduce AutoPage\textbf{AutoPage}, a novel multi-agent system that embodies this philosophy. AutoPage deconstructs paper-to-page creation into a coarse-to-fine pipeline from narrative planning to multimodal content generation and interactive rendering. To combat AI hallucination, dedicated "Checker" agents verify each step against the source paper, while optional human checkpoints ensure the final product aligns perfectly with the author's vision, transforming the system from a mere tool into a powerful collaborative assistant. To rigorously validate our approach, we also construct PageBench\textbf{PageBench}, the first benchmark for this new task. Experiments show AutoPage not only generates high-quality, visually appealing pages but does so with remarkable efficiency in under 15 minutes for less than $0.1. Code and dataset will be released at \href\href{https://mqleet.github.io/AutoPage_ProjectPage/}{Webpage}.
Qianli Ma, Siyu Wang, Yilin Chen +7
Oct 10, 2025cs.LG

CHUCKLE -- When Humans Teach AI To Learn Emotions The Easy Way

Curriculum learning (CL) structures training from simple to complex samples, facilitating progressive learning. However, existing CL approaches for emotion recognition often rely on heuristic, data-driven, or model-based definitions of sample difficulty, neglecting the difficulty for human perception, a critical factor in subjective tasks like emotion recognition. We propose CHUCKLE (Crowdsourced Human Understanding Curriculum for Knowledge Led Emotion Recognition), a perception-driven CL framework that leverages annotator agreement and alignment in crowd-sourced datasets to define sample difficulty, under the assumption that clips challenging for humans are similarly hard for neural networks. Experimental results suggest that CHUCKLE enhances the performance of LSTMs and Transformers over non-curriculum baselines, while reducing the number of gradient updates, thereby enhancing both training efficiency and model robustness in both subject-dependent and subject-independent settings.
Ankush Pratap Singh, Houwei Cao, Yong Liu
Sep 25, 2025cs.HC

Understanding Role Switching in Human-AI Collaboration through Multimodal Behavioral Signals

Human-AI collaboration often requires dividing complex tasks into complementary subtasks. As a task unfolds, users may want to shift which subtasks they perform and which their AI partner performs, in response to evolving task demands and perceptions of the AI's capabilities. In this work, we investigate whether behavioral signals can reveal such changes during a sequential decision-making task. We conducted a study using hand-and-brain chess, where, on each turn, participants chose either to select the piece type (brain) while their AI partner chose the move (hand), or to choose the move after the AI selected the piece type. Across 21 chess players, this yielded more than 1,100 decisions to retain their role from the previous turn or switch to the other role. Players generally retained their current roles across turns. When participants did switch, they exhibited more exploratory gaze patterns and role switches were associated with lower subsequent move quality. Using these behavioral and task-specific signals, we trained a classifier to distinguish switch from stay decisions, achieving a PR-AUC of 0.56 (compared to a random baseline of 0.39). Feature-set ablations showed that gaze and task-specific features contributed most to model performance. Role switching was not a simple choice between controlling and delegating the task as both roles required participants to perform one subtask and delegate the other. However, interviews showed that many participants perceived selecting the piece type as giving them greater control. Perceived AI ability, relative subtask difficulty, and desired influence over the direction of play shaped these role preferences. These findings suggest that behavioral cues could help intelligent systems recognize when users want to reallocate complementary responsibilities during collaboration.
Avinash Ajit Nargund, Arthur Caetano, Kevin Yang +7
Sep 12, 2025cs.CL

Human Psychometric Questionnaires Mischaracterize LLM Behavior

We examine whether human psychometric questionnaires can serve as reliable tools for characterizing and predicting LLM behavior in everyday user interactions. We analyze eight open-source LLMs by comparing their value and personality profiles derived from two different methods: Likert self-reports on established questionnaires (PVQ-40/21 and BFI-44/10) and generation probabilities over value-laden responses to everyday user queries. The two profiles diverge substantially. Within-construct item consistency, often cited as evidence of stable LLM dispositions, disappears in generation probabilities. We find that established questionnaire items contain explicit lexical cues that allow models to recognize the target construct and respond in alignment-consistent, socially desirable ways, whereas realistic user queries contain far less recognizable cues. In addition, demographic persona prompts shift models' responses to human questionnaires in ways consistent with real human patterns, but no such shifts appear in the generation setting, highlighting that human questionnaires overestimate LLMs' ability to faithfully reproduce expected psychological traits when role-playing demographic personas. Overall, our study indicates that questionnaire scores alone should not be treated as evidence of LLMs' response tendencies in realistic user interactions, and supports generation-probability profiling with ecologically valid items as a complementary behavioral measure.
Woojung Song, Dongmin Choi, Yoonah Park +3
Aug 7, 2025cs.RO

Do Robots Really Need Anthropomorphic Hands? A Comparison of Human and Robotic Hands

Human manipulation skills represent a pinnacle of voluntary motor functions, requiring the coordination of many degrees of freedom and the processing of high-dimensional sensor input to achieve remarkable dexterity. Thus, this study investigates whether the human hand, with its associated biomechanical properties, sensors, and control mechanisms, is an ideal that should be strived for in robotics. Do robots need anthropomorphic hands? First, characteristics of the human hand in terms of biomechanics and perception are extracted to compare them with currently commercially available robotic hands. From this comparison, research questions are derived that connect manipulation system complexity to skill repertoire size and dexterity. These questions are addressed through a systematic literature review, analyzing the manipulation capabilities demonstrated in 125 papers published between 2019 and 2025. Although complex five-fingered hands are often considered the ultimate goal for robotic manipulators, they are not necessary for all tasks. Findings indicate that in-hand manipulation does not benefit from anthropomorphic hand design, as simpler mechanisms are sufficient; however, mechanism complexity correlates with the breadth of manipulation tasks a hand can perform. Sensor integration and intelligent manipulation strategies remain underexplored, which may be due to a misalignment with hand design: instead of replicating the number of fingers and degrees of freedom, focusing on robustness and softness would allow more intelligent control and learning to exploit environmental contacts and integrate more sensors. Finally, the article argues for standardized evaluation criteria to enable the systematic comparison of hand designs and manipulation systems.
Alexander Fabisch, Wadhah Zai El Amri, Chandandeep Singh +1
Aug 5, 2025cs.AI

InqEduAgent: Adaptive AI Learning Partners with Gaussian Process Augmentation

Collaborative partnerships play a crucial role in inquiry-oriented education. However, most learning partners are currently assigned through experience-driven heuristics or rule-based machine assistants, which often result in limited knowledge expansion and low adaptability. To address these challenges, this study introduces InqEduAgent, an LLM-empowered generative agent framework designed to simulate and select adaptive learning partners for inquiry-based learning. InqEduAgent integrates a Gaussian process-augmented matching mechanism to model the cognitive and evaluative characteristics of learners, allowing adaptive partner selection based on prior knowledge patterns. Comprehensive experiments demonstrate that InqEduAgent consistently achieves superior performance across diverse learning scenarios and large language model configurations. This study advances human-AI collaborative learning by enabling intelligent pairing between human- and AI-based learning partners, and contributes to adaptive user modeling and personalized recommendation within Web-based educational environments.
Wen-Xi Yang, Tian-Fang Zhao, Guan Liu
Jul 4, 2025cs.HC

Interaction Techniques that Encourage Longer Prompts Can Improve Psychological Ownership when Writing with AI

Writing longer prompts for an AI assistant to generate a story increases psychological ownership, a user's feeling that the writing belongs to them. To encourage users to write longer prompts, we evaluated two interaction techniques that modify the prompt entry interface of chat-based generative AI assistants: pressing and holding the prompt submission button, and continuously moving a slider up and down when submitting a short prompt. A within-subjects experiment investigated the effects of such techniques on prompt length and psychological ownership, and results showed that these techniques increased prompt length and led to higher psychological ownership than baseline techniques. A second experiment further augmented these techniques by showing AI-generated suggestions for how the prompts could be expanded. This further increased prompt length, but did not lead to improvements in psychological ownership. Our results show that simple interface modifications like these can elicit more writing from users and improve psychological ownership.
Nikhita Joshi, Daniel Vogel
Jun 11, 2025cs.CV

MetricHMSR:Metric Human Mesh and Scene Recovery from Monocular Images

We introduce MetricHMSR, a novel framework for recovering metric human meshes and 3D scenes from a single monocular image. Existing methods struggle to recover metric scale due to monocular scale ambiguity and weak-perspective camera assumptions. Moreover, their fully coupled feature representations make it difficult to disentangle local pose from global translation, often requiring multi-stage pipelines that introduce accumulated errors. To address these challenges, we propose MetricHMR (Metric Human Mesh Recovery), which incorporates a bounding camera ray map representation to provide explicit metric cues for human reconstruction,together with a Human Mixture-of-Experts (HumanMoE) that dynamically routes image features to specialized experts, enabling the disentangled perception of local human pose and global metric position. Leveraging the recovered metric human as a geometric anchor, we further refine monocular metric depth estimation to achieve more accurate 3D alignment between humans and scenes.Comprehensive experiments demonstrate that our method achieves state-of-the-art performance on both human mesh recovery and metric human-scene reconstruction. Project Page: https://Metaverse-AI-Lab-THU.github.io/MetricHMSR.
Chentao Song, He Zhang, Haolei Yuan +4
Jun 11, 2025cs.HC

"Is This Really a Human Peer Supporter?": Misalignments Between Peer Supporters and Experts in LLM-Supported Interactions

Mental health is a growing global concern, prompting interest in AI-driven solutions to expand access to psychosocial support. \emph{Peer support}, grounded in lived experience, offers a valuable complement to professional care. However, variability in training, effectiveness, and definitions raises concerns about quality, consistency, and safety. Large Language Models (LLMs) present new opportunities to enhance peer support interactions, particularly in real-time, text-based interactions. We present and evaluate an AI-supported system with an LLM-simulated distressed client (\client{}), context-sensitive LLM-generated suggestions (\suggestions{}), and real-time emotion visualisations. 2 mixed-methods studies with 12 peer supporters and 6 mental health professionals (i.e., experts) examined the system's effectiveness and implications for practice. Both groups recognised its potential to enhance training and improve interaction quality. However, we found a key tension emerged: while peer supporters engaged meaningfully, experts consistently flagged critical issues in peer supporter responses, such as missed distress cues and premature advice-giving. This misalignment highlights potential limitations in current peer support training, especially in emotionally charged contexts where safety and fidelity to best practices are essential. Our findings underscore the need for standardised, psychologically grounded training, especially as peer support scales globally. They also demonstrate how LLM-supported systems can scaffold this development--if designed with care and guided by expert oversight. This work contributes to emerging conversations on responsible AI integration in mental health and the evolving role of LLMs in augmenting peer-delivered care.
Kellie Yu Hui Sim, Roy Ka-Wei Lee, Kenny Tsu Wei Choo
May 28, 2025cs.AI

Modeling and Optimizing User Preferences in AI Copilots: A Comprehensive Survey and Taxonomy

AI copilots represent a new generation of AI-powered systems designed to assist users, particularly knowledge workers and developers, in complex, context-rich tasks. As these systems become more embedded in daily workflows, personalization has emerged as a critical factor for improving usability, effectiveness, and user satisfaction. Central to this personalization is preference optimization: the system's ability to detect, interpret, and align with individual user preferences. While prior work in intelligent assistants and optimization algorithms is extensive, their intersection within AI copilots remains underexplored. This survey addresses that gap by examining how user preferences are operationalized in AI copilots. We investigate how preference signals are sourced, modeled across different interaction stages, and refined through feedback loops. Building on a comprehensive literature review, we define the concept of an AI copilot and introduce a taxonomy of preference optimization techniques across pre-, mid-, and post-interaction phases. Each technique is evaluated in terms of advantages, limitations, and design implications. By consolidating fragmented efforts across AI personalization, human-AI interaction, and language model adaptation, this work offers both a unified conceptual foundation and a practical design perspective for building user-aligned, persona-aware AI copilots that support end-to-end adaptability and deployment.
Saleh Afzoon, Ali Shahsavandi, Phuong Thao Huynh +4
May 22, 2025cs.AI

Serious Games: Human-AI Interaction, Evolution, and Coevolution

The serious games between humans and AI have only just begun. Evolutionary Game Theory (EGT) models the competitive and cooperative strategies of biological entities. EGT could help predict the potential evolutionary equilibrium of humans and AI. The objective of this work was to examine EGT models relevant to human-AI interaction, evolution, and co-evolution. Of thirteen EGT models considered, three were examined: the Hawk-Dove Game, Iterated Prisoner's Dilemma, and the War of Attrition. This selection was based on the widespread acceptance and clear relevance of these models to potential human-AI evolutionary dynamics and co-evolutionary trajectories. The Hawk-Dove Game predicts balanced mixed-strategy equilibria based on the costs of conflict. Iterated Prisoner's Dilemma suggests that repeated interaction may lead to cognitive co-evolution. The War of Attrition suggests that competition for resources may result in strategic co-evolution, asymmetric equilibria, and conventions on sharing resources. Each model was examined from the perspective of human and AI decision-making, from psychological and biological perspectives, and from an AI viewpoint. AI is being shaped by human input and is evolving in response to it. So too, neuroplasticity allows the human brain to evolve in response to stimuli. If humans and AI converge in future, what might be the result of human neuroplasticity combined with an ever-evolving AI? There are profound ethical and cognitive implications. EGT may provide a suitable framework to understand and predict human-AI interaction, evolution, and co-evolution. However, future research should extend beyond EGT and explore additional frameworks, empirical validation methods, and interdisciplinary perspectives. In the spirit of further exploration, an illustrative computational simulation is provided.
Nandini Doreswamy, Louise Horstmanshof
May 21, 2025cs.RO

Human Supervisor Workload Prediction: Lag Horizon Selection

Teleoperation systems must be aware of the human's workload during missions to maintain operator performance. Prior work employed wearable physiological sensor response metrics to estimate current human workload; however, these estimates only enable robots to respond to under- or overload conditions reactively. Current human workload prediction approaches are limited to very short prediction horizons and fail to investigate variable lag horizons' impact on those predictions. This manuscript investigates physiological sensor driven human workload prediction focusing on the impact of lag horizons on both univariate and multivariate time series forecasting models, with longer prediction horizons than the workload prediction state-of-the-art (i.e., > 30 seconds using Long Short-Term Memory networks). Models were trained using data from a 64 participant non-sedentary supervisory environment NASA Multi-Attribute Task Battery-II human subjects evaluation. A key finding is that univariate workload predictions required 240 second lag horizons, whereas multivariate workload predictions sufficed with 120 second lag horizons. This finding indicates additional workload components reduce lag horizon requirements, enabling more efficient models with longer prediction horizons.
Mark-Robin Giolando, Julie A. Adams
May 16, 2025cs.CV

HumaniBench: A Human-Centric Framework for Large Multimodal Models Evaluation

Although recent large multimodal models (LMMs) show impressive progress on vision language tasks, their alignment with human centered (HC) principles such as fairness, ethics, inclusivity, empathy, and robustness is often overlooked. Existing LMM benchmarks are largely accuracy-agnostic. We present HumaniBench, a unified framework for characterizing HC alignment across realistic, socially grounded visual contexts. It contains 32,000 expert-verified image-question pairs from real-world news imagery, each mapped to one or more HC principles through explicit metrics. Comparing 15 state of the art LMMs reveals consistent trade -offs: proprietary systems lead on ethics, reasoning, and empathy, while open-source models show superior visual grounding and resilience. All models show persistent gaps in fairness and multilingual inclusivity. Chain-of-thought prompting and test-time scaling yield 8to 12 % gains on several HC dimensions. HumaniBench enables fine-grained analysis of alignment trade-offs not captured by conventional multimodal benchmarks. https://vectorinstitute.github.io/humanibench/
Shaina Raza, Aravind Narayanan, Vahid Reza Khazaie +6
Apr 27, 2025cs.HC

Toward AI standardization: A triadic human-ai collaboration framework for multi-level autonomous mobility

The goal of the current study is to introduce a triadic human-AI collaboration framework that could be applied in transportation systems such as automated vehicles, micromobility systems, and vehicle teleoperation. Previous standards, such as SAE Levels of Automation, have focused on defining automation levels based on who controls the vehicle. However, it is still not clear how human users and AI should collaborate in real time, especially in dynamic driving contexts where roles can shift frequently. To fill this gap, this study proposed a triadic human-AI collaboration framework with three AI roles: Advisor, Co-Pilot, and Guardian. These roles can dynamically adapt to human needs based on real-time data, such as mental states and environmental conditions. The Advisor AI offers informational support without direct intervention. The Co-Pilot AI provides partial intervention when needed, with the goal of sharing control with humans. The Guardian AI performs emergency overrides if necessary. The use cases for these AI roles in micromobility devices, such as e-scooters, are presented to demonstrate how these roles can influence user preferences and trust. Overall, the study takes a first step toward a universal role-based collaborative framework for AI standardization and explores how AI technologies can be embedded in future transportation systems while considering human interactions.
Gaojian Huang, Wei-Hsiang Lo, Guannan Liu +2
Feb 26, 2025cs.AI

Accelerating scientific discovery with Co-Scientist

Scientific discovery is driven by scientists generating novel hypotheses for complex problems that undergo rigorous experimental validation. To augment this process, we introduce Co-Scientist, a multi-agent AI system built on Gemini for structured scientific thinking and hypothesis generation. Co-Scientist aims to help scientists discover new original knowledge. Conditioned on their research objectives and prior scientific evidence, it formulates demonstrably novel research hypotheses for experimental verification. The system's design involves agents continuously generating, critiquing and refining hypotheses accelerated by scaling test-time compute. Key contributions include: (1) a multi-agent architecture with an asynchronous task execution framework for flexible compute scaling; (2) a tournament evolution process for self-improving hypotheses generation. Automated evaluations show continued benefits of test-time compute scaling, improving hypothesis quality over time. While general purpose, we focus the validation in three biomedical applications: drug repurposing, novel target discovery, and explaining mechanisms of anti-microbial resistance. Specifically, Co-Scientist helped identify new drug repurposing candidates and synergistic combination therapies for acute myeloid leukemia, which were validated through in vitro experiments. These real-world validations demonstrate the potential of Co-Scientist to accelerate scientific discovery and usher in an era of AI empowered scientists.
Juraj Gottweis, Wei-Hung Weng, Alexander Daryin +48
Feb 17, 2025cs.HC

Toward Metaphor-Fluid Conversation Design for Voice User Interfaces

Metaphors play a critical role in shaping user experiences with Voice User Interfaces (VUIs), yet existing designs often rely on static, human-centric metaphors that fail to adapt to diverse contexts and user needs. This paper introduces Metaphor-Fluid Design, a novel approach that dynamically adjusts metaphorical representations based on conversational use-contexts. We compare this approach to a Default VUI, which characterizes the present implementation of commercial VUIs commonly designed around the persona of an assistant, offering a uniform interaction style across contexts. In Study 1 (N=130), metaphors were mapped to four key use-contexts-commands, information seeking, sociality, and error recovery-along the dimensions of formality and hierarchy, revealing distinct preferences for task-specific metaphorical designs. Study 2 (N=91) evaluates a Metaphor-Fluid VUI against a Default VUI, showing that the Metaphor-Fluid VUI enhances perceived intention to adopt, enjoyment, and likability by aligning better with user expectations for different contexts. However, individual differences in metaphor preferences highlight the need for personalization. These findings challenge the one-size-fits-all paradigm of VUI design and demonstrate the potential of Metaphor-Fluid Design to create more adaptive and engaging human-AI interactions.
Smit Desai, Jessie Chin, Dakuo Wang +2
Feb 13, 2025cs.CL

Thinking beyond the anthropomorphic paradigm benefits LLM research

Anthropomorphism, or the attribution of human traits to technology, is an automatic and unconscious response that occurs even in those with advanced technical expertise. In this position paper, we analyze hundreds of thousands of research articles to present empirical evidence of the prevalence and growth of anthropomorphic terminology in research on large language models (LLMs). We argue for challenging the deeper assumptions reflected in this terminology -- which, though often useful, may inadvertently constrain LLM development -- and broadening beyond them to open new pathways for understanding and improving LLMs. Specifically, we identify and examine five anthropomorphic assumptions that shape research across the LLM development lifecycle. For each assumption (e.g., that LLMs must use natural language for reasoning, or that they should be evaluated on benchmarks originally meant for humans), we demonstrate empirical, non-anthropomorphic alternatives that remain under-explored yet offer promising directions for LLM research and development.
Lujain Ibrahim, Myra Cheng
Jan 15, 2025cs.CV

Human Pose-Constrained UV Map Estimation

UV map estimation is used in computer vision for detailed analysis of human posture or activity. Previous methods assign pixels to body model vertices by comparing pixel descriptors independently, without enforcing global coherence or plausibility in the UV map. We propose Pose-Constrained Continuous Surface Embeddings (PC-CSE), which integrates estimated 2D human pose into the pixel-to-vertex assignment process. The pose provides global anatomical constraints, ensuring that UV maps remain coherent while preserving local precision. Evaluation on DensePose COCO demonstrates consistent improvement, regardless of the chosen 2D human pose model. Whole-body poses offer better constraints by incorporating additional details about the hands and feet. Conditioning UV maps with human pose reduces invalid mappings and enhances anatomical plausibility. In addition, we highlight inconsistencies in the ground-truth annotations.
Matej Suchanek, Miroslav Purkrabek, Jiri Matas
Jan 15, 2025cs.CL

LLM-based Human Simulations Have Not Yet Been Reliable

Large Language Models (LLMs) are increasingly employed for simulating human behaviors across diverse domains. However, our position is that current LLM-based human simulations remain insufficiently reliable, as evidenced by significant discrepancies between their outcomes and authentic human actions. Our investigation begins with a systematic review of LLM-based human simulations in social, economic, policy, and psychological contexts, identifying their common frameworks, recent advances, and persistent limitations. This review reveals that such discrepancies primarily stem from inherent limitations of LLMs and flaws in simulation design, both of which are examined in detail. Building on these insights, we propose a systematic solution framework that emphasizes enriching data foundations, advancing LLM capabilities, and ensuring robust simulation design to enhance reliability. Finally, we introduce a structured algorithm that operationalizes the proposed framework, aiming to guide credible and human-aligned LLM-based simulations. To facilitate further research, we provide a curated list of related literature and resources at https://github.com/Persdre/awesome-llm-human-simulation.
Qian Wang, Jiaying Wu, Zichen Jiang +6
Dec 24, 2024cs.CL

Learning from Many Voices: Literary MT Using Multi-Reference Human and Synthetic Data

Unlike many other texts, literary works are often translated multiple times. We investigate strategies for leveraging these multi-reference datasets to improve literary machine translation. We propose a filtering framework based on semantic similarity to identify source texts whose references display meaningful variation while remaining faithful. We find that fine-tuning with medium to high semantic similarity data substantially outperforms low semantic similarity data. Moreover, using medium and high semantic similarity data achieves comparable or better performance than using the full unfiltered data. Synthetic translations generated by LLMs are economical and convenient alternatives to human expert translations; however, we find fine-tuning on human expert translations outperforms fine-tuning on synthetically augmented data in automatic metrics and human evaluations, demonstrating the indispensable value of human expert translations for fine-tuning literary machine translation models.
Si Wu, John Wieting, David A. Smith
Sep 16, 2024cs.CV

PSHuman: Photorealistic Single-image 3D Human Reconstruction using Cross-Scale Multiview Diffusion and Explicit Remeshing

Detailed and photorealistic 3D human modeling is essential for various applications and has seen tremendous progress. However, full-body reconstruction from a monocular RGB image remains challenging due to the ill-posed nature of the problem and sophisticated clothing topology with self-occlusions. In this paper, we propose PSHuman, a novel framework that explicitly reconstructs human meshes utilizing priors from the multiview diffusion model. It is found that directly applying multiview diffusion on single-view human images leads to severe geometric distortions, especially on generated faces. To address it, we propose a cross-scale diffusion that models the joint probability distribution of global full-body shape and local facial characteristics, enabling detailed and identity-preserved novel-view generation without any geometric distortion. Moreover, to enhance cross-view body shape consistency of varied human poses, we condition the generative model on parametric models like SMPL-X, which provide body priors and prevent unnatural views inconsistent with human anatomy. Leveraging the generated multi-view normal and color images, we present SMPLX-initialized explicit human carving to recover realistic textured human meshes efficiently. Extensive experimental results and quantitative evaluations on CAPE and THuman2.1 datasets demonstrate PSHumans superiority in geometry details, texture fidelity, and generalization capability.
Peng Li, Wangguandong Zheng, Yuan Liu +9
Aug 27, 2024cs.CY

Measuring Human Contribution in AI-Assisted Content Generation

With the growing prevalence of generative artificial intelligence (AI), an increasing amount of content is no longer exclusively generated by humans but by generative AI models with human guidance. This shift presents notable challenges for the delineation of originality due to the varying degrees of human contribution in AI-assisted works. This study raises the research question of measuring human contribution in AI-assisted content generation and introduces a framework to address this question that is grounded in information theory. By calculating mutual information between human input and AI-assisted output relative to self-information of AI-assisted output, we quantify the proportional information contribution of humans in content generation. Our experimental results demonstrate that the proposed measure effectively discriminates between varying degrees of human contribution across multiple creative domains. We hope that this work lays a foundation for measuring human contributions in AI-assisted content generation in the era of generative AI.
Yueqi Xie, Tao Qi, Jingwei Yi +7
Mar 6, 2015cs.IT

Information entropy as an anthropomorphic concept

According to E.T. Jaynes and E.P. Wigner, entropy is an anthropomorphic concept in the sense that in a physical system correspond many thermodynamic systems. The physical system can be examined from many points of view each time examining different variables and calculating entropy differently. In this paper we discuss how this concept may be applied in information entropy; how Shannon's definition of entropy can fit in Jayne's and Wigner's statement. This is achieved by generalizing Shannon's notion of information entropy and this is the main contribution of the paper. Then we discuss how entropy under these considerations may be used for the comparison of password complexity and as a measure of diversity useful in the analysis of the behavior of genetic algorithms.
Panteleimon Rodis
Date pendingcs.CL

Inverse Turing Bench: Evaluating Language Models as Judges of Human vs. AI Dialogue

As AI systems integrate into online spaces, differentiating them from humans in conversations is increasingly important. We present Inverse Turing Bench, a benchmark that evaluates LLMs and other models on their ability to differentiate humans and AI in multi-turn text. The benchmark provides a collection of paired dialogue transcripts, wherein one dialogue is between two humans and the other is between a human and an AI. The task is to correctly identify which dialogue is human-only vs. human-AI. We evaluated a preliminary set of models against this benchmark, and found that GPTZero, Claude Opus-4.6, and GPT-5.5 achieve the highest accuracy: 89.41%, 77.92%, and 75.94% respectively. Our results suggest that statistical approaches to detection have semantic blind spots, but semantic approaches are susceptible to persona-prompting. Our work speaks to the Inverse Turing Test and motivates human-AI differentiation as a critical capability for AI systems. Our live benchmark can be found at https://huggingface.co/spaces/roc-hci/Inverse-Turing-Bench-Leaderboard.
William Hager, Ishika Rathi, Masum Hasan +1
Date pendingcs.CR

Builder, Defender, Breaker: Measurable Independence and Bounded Autonomy When Generative Models Build, Defend and Test Software

Generative models now write application code, harden and monitor it, and probe it for exploitable flaws, so that one family of models increasingly plays builder, defender and breaker at once. The prevailing view treats full autonomy as the natural end point of assistance. This article argues for a narrower and more defensible position than a blanket requirement for human oversight. We define the shared generative substrate as the set of upstream dependencies (training corpus, model family, alignment procedure, vendor, toolchain) that induce correlated errors, and we define independence as a measurable property of pairs of lifecycle roles rather than an attribute of any participant. A coincident-failure model in the tradition of Eckhardt and Lee shows why organisational independence, the proxy on which verification standards have relied, no longer implies statistical independence once roles share a substrate; it also shows that heterogeneous models, deterministic analysers and formal verification can restore independence for the fault classes they cover. What remains for humans is specific: authority over the specification against which every machine oracle is judged, accountability that current governance instruments do not permit to be delegated, and last-resort authority to halt. We translate this into an operational framework with five autonomy levels, three distinct human roles and five decision criteria (consequence, reversibility, time-criticality, verifiability and adversarial exposure), apply it to build, defend and test actions, and specify how independence and oversight effectiveness can be measured. The central claims are stated as testable hypotheses with a protocol for refuting them.
Mohamed Chahine Ghanem