AI Trustworthiness

Momentum

28 papers in the last four weeks, up 133% on the four weeks before. 0.3% of all new papers.

Jul 6Week of Sep 21

Latest papers 228

All topics
CardsList
  1. What Robots Do Matters More Than What They Look Like: Task Context Shapes Trust in Educational HRI

    Jun 12, 2026Anna-Maria Velentza, Konstantina Nikou, Anne-Gwenn Bosser +1Human-Robot InteractionRobot Systems

  2. Does Reasoning Preserve Alignment? On the Trustworthiness of Large Reasoning Models

    Jun 9, 2026Prajakta Kini, Avinash Reddy, Souradip Chakraborty +4Large Language Model ReliabilityLarge Reasoning Models

  3. Generative Explainability for Next-Generation Networks: LLM-Augmented XAI with Mutual Feature Interactions

    Jun 9, 2026Kiarash Rezaei, Omran Ayoub, Sebastian Troia +3Explainable Artificial IntelligenceXai

  4. From Context-Aware to Conflict-Aware: Generalizing Contrastive Decoding for Knowledge Conflict in LLMs

    Jun 9, 2026Runze Jiang, Taiqiang Wu, Yan Wang +2Contrastive DecodingAI Trustworthiness

  5. The ACUTE Protocol: Operationalizing Language Model Activations for Better Calibration, Utility, and Trust

    Jun 5, 2026Nishant Subramani, Palash Goyal, Yiwen Song +4Confidence EstimationAI Trustworthiness

  6. Exploring Agentic Tool-Calling Decisions via Uncertainty-Aligned Reinforcement Learning

    Jun 5, 2026Yijin Zhou, Linqian Zeng, Xiaoya Lu +4Agentic Reinforcement LearningOffline Reinforcement Learning

  7. Coding with "Enemy": Can Human Developers Detect AI Agent Sabotage?

    Jun 4, 2026Jingheng Ye, Huiqi Zou, Simon Yu +1SabotageCoding Agents

  8. From Agent Traces to Trust: A Survey of Evidence Tracing and Execution Provenance in LLM Agents

    Jun 3, 2026Yiqi Wang, Jiaqi Zhang, Taotao Cai +8Large Language Model AgentsAI Trustworthiness

  9. Toward Pre-Deployment Assurance for Enterprise AI Agents: Ontology-Grounded Simulation and Trust Certification

    Jun 2, 2026Thanh Luong Tuan, Abhijit SanyalAgentic DeploymentsPre-Deployment Safety Assessments

  10. RoboTrustBench: Benchmarking the Trustworthiness of Video World Models for Robotic Manipulation

    Jun 1, 2026Huiqiong Li, Jiayu Wang, Zhiting Mei +3Video World ModelsRobotic Manipulation

  11. Trust Functions: Near-Lossless Weak-to-Strong Generalization by Learning When to Trust the Weak Teacher

    May 31, 2026Arda Uzunoglu, Alvin Zhang, Daniel KhashabiLearnabilityTeacher

  12. Silent Failures in Federated Personalization of Foundation Models

    May 31, 2026YongKyung Oh, Alex BuiFederated LearningAI Trustworthiness

  13. Can You Trust What You See? Human and AI Detection of Synthetic Legal Evidence

    May 29, 2026Jinzhe Tan, Ali Ekber Cinar, Karim BenyekhlefArtificial Intelligence DetectionGenerative Artificial Intelligence

  14. Personalized to Persuade: The Effects of Contextualization and Warmth on Trust and Reliance in Conversational AI

    May 29, 2026Mert Yazan, Suzan Verberne, Frederik Bungaran Ishak SitumeangPersuasionArtificial Intelligence Literacy

  15. Immuno-VLM: Immunizing Large Vision-Language Models via Generative Semantic Antibodies for Open-World Trustworthiness

    May 29, 2026Xiang Fang, Wanlong Fang, Wei JiOutliersAI Trustworthiness

  16. A Study on Question-Answer Dataset for LLM Safety Evaluation with a Focus on Illegal Activities

    May 28, 2026Kenji Imamura, Masao Ideuchi, Atsushi FujitaLarge Language Model SafetyAI Trustworthiness

  17. Verified Misguidance: Measuring Structural Citation Failures in Search-Augmented LLMs

    May 27, 2026Yongsik Seo, Wooseok Jeong, Eunyoung Kim +2CitationsCredibility

  18. AI, Take the Wheel: What Drives Delegation and Trust in Human-Computer Cooperative Question Answering?

    May 27, 2026Maharshi Gor, Yoo Yeon Sung, Yu Hou +4Human-Ai CollaborationTrustworthy Artificial Intelligence

  19. Defending LLM-based Multi-Agent Systems Against Cooperative Attacks with Sentence-Level Rectification

    May 27, 2026Yaoyang Luo, Zhi Zheng, Ziwei Zhao +5Model-Based Multi-Agent SystemsMalicious Agents

  20. I Hear, Therefore I Trust: A Socio-Technical Investigation of Humans as Synthetic Speech Detectors

    May 27, 2026Lelia Erscoi, Tomi KinnunenUtterancesAI Trustworthiness

  21. Evaluating Local Explainability Metrics for Machine Learning Models on Tabular Data

    May 26, 2026Tomás Pereira, João Vitorino, Eva Maia +1ExplainabilityExplainable AI Methods

  22. Trust Region Q Adjoint Matching

    May 26, 2026Yonghoon Dong, Kyungmin Lee, Changyeon Kim +2Flow PoliciesQ-Learning

  23. Credibility-Aware Learning and Control for Safe USV Navigation under Perception Uncertainty

    May 26, 2026Yuhang Zhang, Shuqi Chai, Yukang Zhang +5Autonomous Underwater VehiclesCollision Avoidance

  24. Quality Without Usefulness: LLM-Generated XAI Narratives as Trust Heuristics Rather Than Decision Aids

    May 26, 2026Fabian Lukassen, Jan Herrmann, Christoph Weisser +3XaiExplainable Artificial Intelligence

  25. The Timing Dependencies of Trust: Speed, Accuracy, and cBCI Neuro-Decoupling in Human-AI Teams

    May 25, 2026Christopher Baker, Stephen Hinton, Akashdeep Nijjar +4Human-Ai CollaborationTeamwork

  26. Evo-Attacker: Memory-Augmented Reinforcement Learning for Long-Horizon Tool Attacks on LLM-MAS

    May 25, 2026Bingyu Yan, Xiaoming Zhang, Jinyu Hou +4Multi-Agent EvolutionModel-Based Multi-Agent Systems

  27. Eroding Trust in Real Speech: A Large-Scale Study of Human Audio Deepfake Perception

    May 21, 2026Nicolas M. Müller, Wei Herng ChoongAudio Deepfake DetectionFake News

  28. Adversarial Trust Poisoning in Vehicular Collaborative Perception

    May 21, 2026Yutong Liu, Chenyi Wang, Ming F. Li +1Collaborative PerceptionPoisoning

  29. Understanding Perspectives of Patients, Caregivers and Clinicians towards Emerging Collaborative-decision Making Technologies

    May 20, 2026Ray-Yuan Chung, Athena Ortega, Zixuan Xu +7TechnologyAI Trustworthiness

  30. Trusted Weights, Treacherous Optimizations? Optimization-Triggered Backdoor Attacks on LLMs

    May 20, 2026Yifei Wang, Yida Yang, Tianlin Li +3Attacker Large Language ModelAI Trustworthiness

  31. Rethinking Visual Attribution for Chest X-ray Reasoning in Large Vision Language Models

    May 19, 2026Guangzhi Xiong, Qiao Jin, Sanchit Sinha +2Medical Vision-Language ModelsChest X-Ray

  32. Towards Trust Calibration in Socially Interactive Agents: Investigating Gendered Multimodal Behaviors Generation with LLMs

    May 19, 2026Lucie Galland, Chloé Clavel, Magalie OchsHuman-Ai InteractionAI Trustworthiness

  33. Trust It or Not: Evidential Uncertainty for Feed-Forward 3D Reconstruction with Trust3R

    May 19, 2026Zihao Zhu, Wenyuan Zhao, Nuo Chen +2Feed-Forward 3D ReconstructionRobust Geometric Model Estimation

  34. A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook

    May 18, 2026Kaiwen Luo, Zhenhong Zhou, Leo Wang +31Large Audio Language ModelsLarge Language Model Safety

  35. Trustworthy Agent Network: Trust in Agent Networks Must Be Baked In, Not Bolted On

    May 18, 2026Yixiang Yao, Yuhang Yao, Xinyi Fan +5Model-Based Multi-Agent SystemsTrustworthy Artificial Intelligence

  36. Exploring Trust Calibration in XAI - The Impact of Exposing Model Limitations to Lay Users

    May 18, 2026Alfio Ventura, Tim Katzke, Jan Corazza +1XaiAI Trustworthiness

  37. Trust No Tool: Evaluating and Defending LLM Agents under Untrusted Tool Feedback

    May 17, 2026Lecheng Yan, Ruizhe Li, Xicheng Han +5Untrusted ContentLarge Language Model Safety

  38. Towards trustworthy agentic AI: a comprehensive survey of safety, robustness, privacy, and system security

    May 17, 2026Jinhu Qi, Muzhi Li, Jiahong Liu +9Trustworthy Artificial IntelligenceAI Trustworthiness

  39. Right Predictions, Misleading Explanations: On the Vulnerability of Vision-Language Model Explanations

    May 15, 2026Narges Babadi, Hadis KarimipourExplainable AI MethodsHeatmap

  40. To Trust or Not to Trust: Authors' Response to AI-based Reviews

    May 15, 2026César Leblanc, Lukas PicekPeer ReviewAuthorship

  41. The End of Trust: How Agentic AI Breaks Security Assumptions

    May 14, 2026Osama Zafar, Alexander Nemecek, Erman AydaySecurityAI Trustworthiness

  42. Uncertainty-Aware 3D Position Refinement for Multi-UAV Systems

    May 13, 2026Hosam Alamleh, Damir PulatovMulti-UavIndoor Localization

  43. TRUST-TAEA: A trustworthiness-guided two-archive evolutionary algorithm with variable-grouping sparse search for large-scale multi-objective optimization

    May 13, 2026Junyi Cui, Chao Min, Stanisław Migórski +2Multi-Objective OptimizationEvolutionary Optimization