AI-generated faces can be difficult to distinguish from real ones, leaving viewers to rely on source labels when judging an image. Yet prior work has made it difficult to separate the effects of what an image actually is from what viewers are told it is. We validated faces as AI-generated or human in an online study (N=169), then crossed actual source (AI, human) with label (none, Made with AI, Made by a human) in a lab study $N=30), recording event-related potentials (ERPs) and gaze. ERP responses were equivalent for AI-generated and real faces, but varied with the label: labels drew early attention (N2), while labels that conflicted with the face's actual source prompted re-evaluation of the face (P3). Affective processing and initial gaze orienting were unchanged, but labels altered visual exploration. We provide a validated stimulus set and evidence that attributed origin shapes face processing, with implications for disclosure design.
Figures & tables
Figure 1 . Observers cannot tell an AI-generated face from a real one, so the disclosure label attached to it becomes the only available cue to its origin. In the figure, two faces of identical apparent quality are depicted, one AI-generated and one real, and are shown with the source labels used in our study. The scanpath illustrates how participants moved between the face and the label while EEG and eye movements were recorded. Faces carrying a label drew more and shorter fixations than unlabelled ones. We found no clear difference between AI-generated and real faces on any of our measures, while the origin attributed by the label was reflected in how a face was processed. Image generated with Gemini Nanobanana 2.0 . Two illustrated portraits, each carrying a small corner tag reading Made with AI'' or Made by a human''. Circles marking eye fixations are distributed across each face and on each tag, joined by lines into a scanpath that runs to the eye of a head shown in profile, with an EEG waveform drawn across the skull. At right, two rows of a forest plot share a vertical zero line: the row labelled what the image is'' has its point estimate on zero with the interval crossing it, and the row labelled what the label says'' has its point estimate to the right of zero with the interval clear of it.
Figure 2 . Example images of AI-generated faces used in Study 1.
Figure 3 . Perceived AI-likeness vs. rating variability. Relationship between mean AI-likeness rating and rating variability (SD) across the 320 AI-generated faces. Blue points mark the 82 images whose lower 95% CI ≥60 , satisfying the inclusion criterion for the validated set. The dashed line indicates the CI-based classification boundary.
Figure 4 . Summary of perceptual consistency metrics. (a) shows how mean ratings distribute around the inclusion threshold; (b) illustrates the achieved precision of participants’ judgments via CI widths.
Figure 5 . Sorted mean ratings with 95% confidence intervals. Mean AI-likeness ratings for all 320 AI-generated faces, ordered by mean value. Error bars represent the 95% CI per image, and the dashed line at 60 marks the classification boundary used to identify the 82 reliably AI-looking stimuli.
Figure 6 . Trial structure. Each trial began with a fixation cross (1000 ms), followed by a blank inter-stimulus interval (1000 ms plus a random jitter of 250, 500, or 750 ms) to reduce temporal expectancy and ERP overlap. A face was then shown for 4000 ms, either AI-generated or human and carrying a Made with AI , Made by a human , or no label, while EEG and eye movements were recorded. After the face offset, participants rated on a continuous scale whether the image was created with the help of generative AI and reported their confidence. Trials followed one another in fully randomized order.
Figure 7 . Example images used in Study 2. In the first row, all images are AI-generated faces, representing all three label possibilities from left to right: Made with AI, Made by a human, No label. In the second row, all images are real humans, also representing the above mentioned labeling possibilities.
Figure 8 . An example of how the label was positioned in the bottom right corner, following Gamage et al. (2025) ’s work.
Figure 9 . Posterior effects on the three ERP components. Each row is a contrast: the marginal difference between AI- and human-sourced faces ( Source ), and the congruent and incongruent labelling conditions, each relative to unlabelled faces. Points are posterior medians and bars 95% credible intervals; a row is drawn dark when its interval excludes zero and grey when it crosses zero, with the probability of direction ( pd ) at right. Source sits close to zero in every component, while the label-related movement appears in the congruence contrasts: the N2 grows with a label present and the P3 with label–face conflict, whereas the LPP is unmoved by either factor.
Figure 10 . Observed amplitude means with 95% within-subject confidence intervals (Cousineau–Morey), by image source and labelling condition, for each component; amplitudes are averaged over the component region of interest (frontocentral for the N2, centroparietal for the P3 and LPP). The AI and human traces overlap at every level, showing that actual source did not separate the response. The condition axis carries what movement there is: the N2 becomes more negative once a label is present, the P3 rises under incongruence and most so for AI faces, and the LPP stays flat.
Figure 11 . N2 grand average waveform. Averages for the six conditions (AI/Human face × Made with AI/Made by a human/No label), measured over fronto-central electrodes (Fz, Cz, FC1, FC2); the shaded region (200–350 ms) marks the analysis window. AI and human faces produced the same response. Any face carrying a label, whether the label matched the face or not, produced a larger response than a face with no label at all, though the difference is small. The earliest sensitivity is therefore to the presence of a source cue, not to its agreement with the face, which is evaluated at a later stage.
Figure 12 . P3 grand average waveform. Averages for the six conditions (AI/Human face × Made with AI/Made by a human/No label), measured over centro-parietal electrodes (Pz, CPz, Cz, CP1, CP2, P3, P4); the shaded region (300–500 ms) marks the analysis window. AI and human faces produced the same response. An AI face labelled “Made by a human”, or a human face labelled “Made with AI”, produced a larger response than a face with no label at all.
Figure 13 . LPP grand average waveform. Stimulus-locked averages for the six experimental conditions (AI/Human source × Made with AI/Made by a human/No label), computed over centro-parietal electrodes (Pz, CPz, Cz, CP1, CP2, P1, P2). The shaded region (400–800 ms) indicates the LPP analysis window. A sustained positivity rises from roughly 200 ms and is maintained across the window in every condition. The six traces overlap throughout, and the models estimate no difference by actual source or by label congruence. Whatever a disclosure label did to earlier processing did not carry into this later affective stage.
Figure 14 . Posterior effects of image source, disclosure label, and their interaction on each eye-tracking measure. Points are posterior medians and bars the 95% CrIs; the upper axis restates the log-scale coefficients as percentage change. Effects whose interval excludes zero are dark, those overlapping zero are grey, and pd is given at right. Labels credibly shifted how faces were explored while leaving the speed of initial orienting unchanged, and image source registered only in fixation duration.
Figure 15 . Observed cell means with 95% within-subject confidence intervals (Cousineau–Morey) for each measure, by image source and label presence, on the natural scale of each measure; total dwell time is omitted as redundant. The panels show the pattern behind the modelled effects: fixations were shorter and more numerous when a label was present, the fixation-count shift fell almost entirely on AI faces, and the timing of the first fixation overlapped across all conditions.
Across social and online platforms, people are increasingly exposed to AI-generated images. As a consequence, the task of distinguishing AI-generated from authentic images is becoming a central challenge for information ecosystems. While humans perform better than chance, accuracy falls short of many operational needs. Initial evidence shows that visually oriented training can improve deepfake detection but does not improve participants' ability to identify real images as real. Here, we investigate the efficacy of a brief training intervention for intelligence analysts employed by the United States government in 2024. We conducted a counterbalanced within-subject randomized experiment in which we showed participants real and AI-generated images varying in pose complexity and scene context and asked them whether each image was real or AI-generated, both before and after an expert delivered a 30-minute training that pointed out patterns in seven real and 50 AI-generated images. We collected 2,544 image-level judgments from 32 intelligence analysts. We find training increased overall accuracy by 9 percentage points (95% CI: [2.7, 15.4]) from a baseline of 72%. We find the improvement is driven by a 14.2 percentage point increase in accuracy for real images (95% CI: [0.7, 27.7]). Through a careful experimental setup that curated matched pairs of real and AI-generated images across pose complexity categories, we reveal how these trainings influence people with different levels of digital forensics and generative AI experience and identify the kind of image-based content where this training intervention appears to be most effective. Ultimately, these results provide causal evidence that a brief, structured training can improve human judgment across a diverse array of real and AI-generated images, informing organizational responses to AI-generated visual misinformation.
Department of Computer Science, Northwestern University · Ryan Institute on Complexity, Northwestern University · Research Directorate and AI Security Center, National Security Agency +1
As generative AI makes polished prose cheap to produce, users can no longer rely on fluency as a proxy for truth. We call this failure mode the Fluency Trap: users trust fluent hallucinations while also discounting accurate content once it is disclosed as AI-generated. Binary ``Made with AI'' labels respond with authorship disclosure, but they do not show what supports a claim. We propose Provenance Density, an evidence-visualization interface that shows the density of verified claims in a text. In a user study with 81 participants, an idealized Provenance Density interface produced a large discernment gap between truth and fabrication (+4.15 points, d=1.82), whereas participants given no signal showed no detectable discrimination. A technical audit with 200 samples shows that retrieval density alone is insufficient; unexpectedly, the Consistency Veto carries most of the discriminative signal on dynamic queries. As AI-generated content becomes indistinguishable from human writing, effective transparency must move from authorship disclosure toward evidence visualization.
Qing Zhang, Yifei Huang, Juyoung Lee +2
The University of Tokyo · Institute of Industrial Science, The University of Tokyo · Korea Advanced Institute of Science and Technology +2
AI systems are increasingly deployed in conversational settings where users may be uncertain whether they are speaking with a human or an AI. Despite mounting regulatory attention to this known safety risk, existing evaluations of AI disclosure are typically English-only, based on machine-generated questions, and restricted to text. We present RealityTest to comprehensively test whether AI systems disclose their identity when asked. The benchmark is the first large-scale multimodal and multilingual evaluation, grounded in human data on how people actually encounter and question AI identity in the real-world. Alongside the benchmark, we release the underlying dataset of 3,152 identity-probing queries collected from ~750 participants across 49 countries and five languages, in text and speech scenarios. We find that only 31% of people ask about identity directly in ambiguous scenarios, and that the questions people ask are far more diverse than machine-generated queries. We test 17 text and 6 speech models, and find substantial variation in disclosure behaviour. However, a single suppression instruction reduces disclosure rates to below 30%, even in the best-performing models. Validating our investment in diverse, human-grounded evaluation data, we find that how the question is phrased and the context of the conversation matter more for disclosure than which model is being tested. Safety evaluations built on narrow or synthetic query sets risk mischaracterising how models behave in realistic deployment settings.