From Policy Documents to Structured Survey Responses: Evaluating Large Language Models for Policy Monitoring
Organizations: Reliable Intelligence Team, VTT Technical Research Centre of Finland Ltd., 02150 Espoo, Finland
Abstract
Science, technology, and innovation policies are crucial for competitiveness, yet their diversity and scale make them difficult to map and monitor consistently. Existing approaches rely heavily on manual survey efforts, which are costly and challenging to scale across countries. Large language models (LLMs) enable new possibilities for extracting and structuring information from long and unstructured policy documents. This paper presents an application of LLMs as "AI respondents" for generating structured survey responses from policy texts. We develop a data extraction pipeline based on long-context in-context learning to map information from public web sources into predefined survey categories, including policy instruments, target groups, and thematic areas. The pipeline integrates a validation step using a secondary LLM to assess relevance and evidence, alongside comparisons with human-provided responses. Using a multi-country dataset, we evaluate the alignment between LLM-generated and human-generated outputs through overlap measures and cross-validation. Results show that LLMs achieve high agreement for structured indicators (84-95%), while differences remain in free-text fields, where models tend to provide more detailed procedural descriptions. These findings highlight the potential of hybrid human-AI workflows for policy monitoring, improving both efficiency and scalability while maintaining the need for human validation and contextual interpretation.
Figures & tables
| Web text | The Government is increasingly looking to utilize artificial intelligence to make, or assist in making, administrative decisions…compatible with core administrative law principles such as transparency, accountability, legality, and procedural fairness. The objective of this Directive is to ensure that Automated Decision Systems are deployed in a manner that reduces risks to Canadians and federal institutions … Completing an Algorithmic Impact Assessment prior to the production of any Automated Decision System … Applying the relevant requirements prescribed in …developing processes so that the data and information used …are tested for unintended data biases … providing a meaningful explanation to affected individuals of how and why the decision was made … Contracted third-party vendor with a related specialization …Data and information on the use of Automated Decision Systems…are made available to the public, where appropriate… |
|---|---|
| Codes | PIs: PI027, PI032, TGs: TG16, TG23, TG29, THs: TH89 |
| Text fields | Description: Federal policy instrument providing a risk-based approach… |
| Objectives: Automated decision systems deployed by federal… | |
| Label defs. | PI027 : Governance | Standards and certification for technology development and adoption |
| PI032 : Guidance, regulation and incentives | Science and technology regulation and soft law | |
| TG16 : Social groups especially emphasised | Civil society |
| Sr# | Country | Insufficient | Unidentified | Sufficient | # of Samples |
|---|---|---|---|---|---|
| 1 | Canada | 30% | 6% | 64% | 149 |
| 2 | Finland | 32% | 11% | 57% | 80 |
| 3 | Germany | 25% | 7% | 68% | 193 |
| 4 | Korea | 27% | 20% | 53% | 146 |
| 5 | Spain | 31% | 13% | 56% | 142 |
| 6 | Türkiye | 47% | 12% | 41% | 135 |
| Generated | PI | Unique PI | TG | Unique TG | TH | Unique TH |
|---|---|---|---|---|---|---|
| By humans | 1,281 | 27 | 3,828 | 33 | 1,895 | 51 |
| By LLM | 2,336 | 28 | 4,727 | 33 | 3,013 | 57 |
| Policy instruments | Target groups | Policy themes | ||
|---|---|---|---|---|
| Sr# | Country | (A) | (B) | (C) |
| 1 | Canada | 84% | 97% | 85% |
| 2 | Finland | 84% | 98% | 84% |
| 3 | Germany | 80% | 97% | 83% |
| 4 | Korea | 88% | 93% | 84% |
| 5 | Spain | 85% | 94% | 82% |
| Sr.# | Model | Precision | Recall | F1 |
|---|---|---|---|---|
| 1 | RoBERTa-large [ 16 ] | 86.70 | 66.61 | 75.33 |
| 2 | BigBird-RoBERTa-base [ 26 ] | 87.04 | 69.25 | 77.12 |
| 3 | BigBird-RoBERTa-large [ 26 ] | 88.43 | 76.01 | 81.74 |
| 4 | Llama-3.1-8B-Instruct [ 9 ] | 82.02 | 90.79 | 86.18 |
| 5 | Mistral-7B-Instruct-v0.3 [ 18 ] | 92.91 | 90.90 | 91.89 |
| 6 | GPT-OSS-20B [ 21 ] | 97.63 | 97.83 | 97.73 |
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
| Setting | Value |
|---|---|
| Max sequence length | 512 / 1600 |
| Batch size (train/eval) | 8 / 8 |
| Learning rate | 3e-5 |
| Epochs | 40 |
| Folds | 5 |
| Weight decay | 0.01 |
| Setting | Value |
|---|---|
| Max input length | 7500 |
| Max new tokens | 200 |
| Batch size (train/eval) | 2 / 2 |
| Gradient accumulation | 8 (effective batch 16) |
| Learning rate | 2e-5 |
| Epochs | 4 |
| Parameter | Value |
|---|---|
| (rank) | 8 |
| lora_alpha | 32 |
| lora_dropout | 0.05 |
| bias | none |
| task_type | CAUSAL_LM |
| Target modules | q_proj, k_proj, v_proj, o_proj |