physics.soc-phSep 14, 2026

Assessment of Non-Institutional AI Tool Usage Among Clinicians

Authors: Sarah PungitoreJarrod Mosier

Abstract

Generative artificial intelligence (AI) tools are increasingly accessible and have the potential to improve efficiency across clinical workflows. However, clinicians may also use non-institutional AI tools that are not provided, managed, or governed by their healthcare institutions, creating potential concerns related to privacy, security, accuracy, and clinician-AI interaction. Little is known about how clinicians currently use these tools for work-related tasks. We conducted a descriptive survey of clinicians recruited from the University of Arizona College of Medicine-Tucson and Banner University Medical Center-Tucson between May 20 and June 26, 2026. Participants reported their use of AI tool categories and the frequency with which they used AI for specific tasks across five workload categories: administrative work, clinical work, research, studying/continued education, and teaching. Forty-four respondents completed the survey. Forty-three respondents reported using AI for at least one work-related task during the preceding 6 months. Conversational AI and clinical decision support/diagnostic AI were the most used tool categories, each reported by 28 respondents. Administrative and clinical tasks demonstrated the most frequent use. AI was also used for higher-risk activities, including diagnostic assistance and clinical decision support. Non-institutional AI use was common among surveyed clinicians and extended across a broad range of work-related activities, including tasks with potential implications for clinical reasoning and patient care. Further research is needed to characterize how clinicians use these tools, how they evaluate AI-generated outputs, and how AI can be safely and effectively integrated into clinical workflows.

Explore similar work

Jun 3, 2026cs.CY

How Indian Dermatologists are Utilizing Artificial Intelligence for Clinical Practice and Workflow Management: A Nationwide Survey with a Special Focus on atopic dermatitis

Background: Dermatology AI has mainly focused on image-based diagnosis, while chronic disease workflows have received less attention. We surveyed Indian dermatologists to map routine clinical challenges, with a focus on atopic dermatitis (AD), and assess current AI use. Methods: A nationwide cross-sectional survey commissioned by the Society for Eczema Studies included 377 practicing Indian dermatologists. The survey assessed clinical challenges, AD workflow barriers, AI use, adoption barriers, and ethical concerns. Analyses used descriptive statistics, chi-square tests, false discovery rate correction, and multivariable logistic regression. Results: Patient adherence (61.3%) and treatment planning in difficult or refractory cases (57.0%) were reported more often than diagnostic uncertainty (48.0%). In AD care, severity scoring was reported as a challenge by 47.7% and had the lowest satisfaction among measured workflow areas. Current AI use was reported by 49.9%, most often involving general large language models for literature synthesis, documentation, and academic tasks rather than specialized image analysis. Barriers differed by experience: dermatologists with more than 20 years of practice more often cited lack of training, while those with 5 years or less more often cited lack of clinical utility after trying AI tools. AI users were more likely than non-users to report concern about patient self-misdiagnosis and anxiety, which remained significant after adjustment for experience and academic affiliation. Conclusion: Respondents reported using general-purpose AI mainly for cognitive and administrative tasks, while their clinical needs centered on chronic disease management and AD workflow support. Clinician-supervised workflow tools may be more useful than standalone diagnostic applications.
Dipayan Sengupta, Saumya Panda, Sandipan Dhar +3
Apr 30, 2026cs.AI

Consumer Attitudes Towards AI in Digital Health: A Mixed-Methods Survey in Australia

AI applications are increasingly being introduced into digital health. While technical performance has advanced rapidly, successful deployment mainly depends on consumer attitudes, especially to patient-facing applications. However, most existing research examines consumer attitudes towards healthcare AI at an abstract level rather than in response to concrete artefacts. We report a mixed-methods survey study in Australia (N=275) examining consumer readiness, acceptance, trust, and risk perceptions of healthcare AI, combined with a scenario-based evaluation of an AI-generated versus clinician-written consultation summary. Participants expressed moderate optimism and strong perceived usefulness and ease of use, but also substantial concerns about accuracy, safety, and data use. In the scenario task, the AI-generated summary was strongly preferred for quality, empathy, and overall usefulness, yet identification of the AI summary was near chance. Findings show that consumers judge AI through concrete communication quality and visible human governance, underscoring the need for clinically supervised deployment frameworks beyond technical performance alone.
Wei Zhou, Rashina Hoda, Joycelyn Ling
Jun 27, 2026cs.AI

Expert Evaluation of Clinical AI Tools on Real Point-of-Care Clinical Queries

Physicians now pose millions of clinical questions to AI tools each week, yet these tools are evaluated largely on hypothetical or exam-style questions, not those actually asked in practice. We report a blinded evaluation built on 620 Real-world Point-Of-Care Queries (Real-POCQi) submitted to the OpenEvidence (OE) platform by physicians spanning 30 specialties, as well as 187 questions from HealthBench. 149 practicing physicians across 36 states made head-to-head comparisons between answers from three frontier general-purpose models (Claude Opus 4.8, Gemini 3.1 Pro, and GPT-5.5) and a specialized clinical tool (OE), with graders matched to each question's specialty. When comparing answers along five dimensions relevant to clinical decision support -- accuracy, clinical utility, source quality, verifiability, & completeness -- physicians scored the specialized tool highest on all axes; in the primary analysis on Real-POCQi, win differences (margins between win and loss rates) ranged from 25 to 39 percentage points (p<0.001). Results remained consistent in sensitivity analyses stratifying by citation display, answer length, OE-user status, and Real-POCQi versus HealthBench. In parallel, LLM judges were found to systematically differ from expert judges, though both generally agreed on the best model. These findings underscore two conclusions: (i) AI tool evaluations should reflect real-world query distributions and use expert judges that mirror the specialization defining modern medicine and (ii) the consistent advantage of the specialized tool over general-purpose models does not necessarily mean that the latter cannot serve similar purposes, but that targeted engineering and customization can yield meaningful gains in performance for its users. We release Real-POCQi as a public benchmark, as well as the prespecified statistical analysis for reproducing results of this study.
Jean Feng, Vishal Patel, Patrick Heagerty +5