cs.CLSep 24, 2026

From Policy Documents to Structured Survey Responses: Evaluating Large Language Models for Policy Monitoring

Authors: Carolyn Cole, Matthias Deschryvere, Toqeer Ehsan, Arash Hajikhani

Organizations: Reliable Intelligence Team, VTT Technical Research Centre of Finland Ltd., 02150 Espoo, Finland

Abstract

Science, technology, and innovation policies are crucial for competitiveness, yet their diversity and scale make them difficult to map and monitor consistently. Existing approaches rely heavily on manual survey efforts, which are costly and challenging to scale across countries. Large language models (LLMs) enable new possibilities for extracting and structuring information from long and unstructured policy documents. This paper presents an application of LLMs as "AI respondents" for generating structured survey responses from policy texts. We develop a data extraction pipeline based on long-context in-context learning to map information from public web sources into predefined survey categories, including policy instruments, target groups, and thematic areas. The pipeline integrates a validation step using a secondary LLM to assess relevance and evidence, alongside comparisons with human-provided responses. Using a multi-country dataset, we evaluate the alignment between LLM-generated and human-generated outputs through overlap measures and cross-validation. Results show that LLMs achieve high agreement for structured indicators (84-95%), while differences remain in free-text fields, where models tend to provide more detailed procedural descriptions. These findings highlight the potential of hybrid human-AI workflows for policy monitoring, improving both efficiency and scalability while maintaining the need for human validation and contextual interpretation.

Figures & tables

Appendix figures & tables3 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Can Large Language Models Anticipate Behavioral Responses to Social Policies? A Case of Pension Enrollment Prediction among China's Flexible Workers

    Sep 4, 2026Yumiao Li, Peixin Liu, Donglin Di +2Large Language Model Policy OptimizationAgent Behavior Modeling

  2. Can Large Language Models Revolutionize Survey Research? Experiments with Disaster Preparedness Responses

    May 19, 2026Yan Wang, Ziyi Guo, Christopher McCartySurveyLarge Language Model Evaluation

  3. Poli-Bias: Understanding and Measuring Large Language Model Biases in International Political Conflicts

    Aug 6, 2026Massi-Nissa Abboud, Aladin Djuhera, Elena Cabrio +1Large Language Model BiasGeopolitical Grounding Interacts