cs.AIJun 22, 2026

AI Exposure Scores: what they measure, what they miss, and what comes next

Authors: Campbell LundThomas EuyangZanele MunyikwaMarzieh Fadaee

Organizations: Cohere Labs

Abstract

A set of exposure scores calculated in 2023 has become a central empirical input to the future of work debate. Produced by Eloundou et al. (2023) and referred to here as the GPTs are GPTs scores, they define exposure as the share of occupational tasks a large language model can assist with. This work is a genuine methodological contribution, but as the scores travel from the time and place they were produced, the limitations the authors named do not always travel with them. Two gaps have widened as a result. The first is structural, between what static exposure scores measure and what policy questions actually require. Taking the diffusion of these scores as a case study, we show how their temporal, geographic, and ontological limitations compound in policy-facing analyses, and we survey five families of research responding to these limits: dynamic and benchmark-based measures, ensemble methods, task-framework extensions, worker-centered metrics, and adoption and usage data. The second gap is the one we argue needs more attention: the coordination between researchers and policymakers. The policy-relevant work which ask who is harmed, who benefits, how, and when, continues to reference the static GPTs are GPTs scores without engagement with the methodological updates that would let these questions be answered more reliably. We then ask what additional steps towards navigating uncertainty remain: ex-post frameworks and the deliberate, political work of reimagining what futures are worthy of building towards are. Closing the research-policy gap is a shared task: policymakers must widen their evidence base, engage workers as epistemic partners, and shift from prediction to preparedness; researchers must build data infrastructure, adopt participatory methods, and write with policymakers in mind. Better measurement matters, but it will not close the second gap alone.

Explore similar work

May 20, 2026cs.AI

Who Uses AI? Platform Selection and the Measurement of Occupational AI Exposure

Conversation logs from AI platforms are increasingly used to measure occupational exposure to artificial intelligence, but the users observed in these logs are not the workforce. We show that platform-derived exposure scores combine task-level AI applicability with the occupational composition of the platform's user base. Holding the empirical design fixed, changing only the platform input changes the post-ChatGPT employment coefficient by a factor of 1.9, and consumer and enterprise channels within the same vendor disagree in sign. We formalize the resulting non-classical measurement error, decompose it into between- and within-occupation selection, and construct workforce-reweighted partial-identification bounds. Reweighting to Bureau of Labor Statistics employment shares attenuates estimates by 42 to 93 percent. The bias captures augmentation among observed users more directly than substitution in the workforce.
Michelle Yin, Burhan Ogut
Sep 14, 2026cs.HC

Beyond AI Literacy: A Structured Review and Exploratory Meta-Analysis of Measures for Competent Generative-AI Use

Researchers assessing competent generative-AI use at work must choose among self-reports, objective tests, and measures of oversight and reliance. We conducted a structured, seeded review of 24 focal empirical publications, starting from the 2024 COSMIN-based review and adding a targeted update through 17 August 2026. We grouped the measures into four domains: knowledge and use, epistemic oversight, reliance calibration, and operational control of tool-using agents. In an exploratory meta-analysis, we pooled three direct subjective-objective correlations from one research program (REML r = .055; Hartung-Knapp 95% CI [-.047, .156]; combined reported N = 2,765). We could not resolve a discrepancy between the largest study's reported correlation and p-value, leaving its weight uncertain. Adding a synthetic mean of 12 cross-factor correlations from a fourth study gave r = .079 (95% CI [-.025, .181]). This sensitivity analysis concerns a broader comparison. From this small evidence base, we cannot establish a population correlation, validate workplace cutoffs, or justify substituting self-ratings for performance scores. We identified tests of foundation knowledge (AICOS-S and GLAT) and measures of verification, reliance, trust, and dependency. We found no validated individual-level instrument in the focal corpus that tests the full combination of agent scope, permissions, recovery, state isolation, independent review, and evidence-based closure; some cover subsets. We propose a four-layer workplace battery with non-compensatory decision rules, but have not tested its thresholds or whether it improves on other assessment approaches.
Daniele Veri'
Aug 19, 2026cs.AI

Who Delegates to AI? Evidence from Agent Configurations in Github

A growing body of literature measures the extent to which occupations are exposed to AI, yet existing measures capture where AI could perform tasks rather than whether workers have actually adopted it. We introduce a distinct tier of exposure, delegated exposure, which records whether a worker has committed a task to AI by embedding it into a structured workflow. We operationalize this concept through the Agentic Adoption Index (AAI), measuring how closely an occupation's tasks align with the agentic routines that practitioners have built and shared. Using semantic embeddings of roughly 888,000 agent skill specifications from public GitHub repositories, we compute their similarity to nearly 18,000 O*NET task statements and aggregate these scores to the occupational level. We present three main findings. First, the occupations where task delegation concentrates differ sharply from those identified as most vulnerable by pre-AI automation frameworks. Second, the AAI aligns more closely with measures of technical capability than with measures of current conversational LLM use. Third, for occupations requiring a bachelor's degree or less, the AAI increases alongside average wage levels; however, this relationship reverses for occupations requiring a master's degree or higher, where adoption declines among higher earners. These patterns replicate on an independently collected corpus of agent skills from the Manus Skills Marketplace. This lower adoption among highly educated, high-earning workers may reflect tasks that inherently resist advance specification or professional discretion over the pacing of workflow codification. Distinguishing these mechanisms will require longitudinal measurement.
Hyeongjae Lee, Jihyang Cheon, Lanu Kim