Large Language Models (LLMs) have renewed interest in autoformalization. Yet, when First-Order Logic (FOL) is considered as the target formalism, the field still lacks a unified task formulation and a systematic survey. This paper addresses this gap: we first provide a principled definition for the FOL-autoformalization task by distinguishing Ontology Extraction from Logical Translation, showing how their conflation obscures (cross-study) evaluation; we review existing datasets, evaluation metrics, and LLM-based methods, including fine-tuning, prompting, and verification-based refinement; we identify open challenges in benchmarking, semantic evaluation, ontology-aware methods, and end-to-end applications.
Figures & tables
Figure 1: A single NL sentence p′ admits multiple formalizations φ1,φ2,φ3 , each committing to a different signature. The three formulas can be considered all correct but they express information differently.
Dataset
Generation Type
No. of Pairs
NL
FOL
FOLIO
Human
Human (*)
∼ 3800
GGC
Human
Human (*)
∼ 280
MALLS
LLM
LLM
∼ 28K
ProverQA
LLM (*)
Rule-based
∼ 17K
Willow
LLM
LLM
∼ 16K
Table 1: Overview of NL–FOL datasets. ‘Generation Type’ indicates who generates NL and FOL. Asterisks denote that the output is obtained by a translation (from the NL or the FOL part); no asterisk indicates joint generation.
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Dataset
NL
FOL
No. of Instances
ReClor Yu et al. (2020)
Human
Not available
∼ 6k
LogiQA Liu et al. (2020)
Human
Not available
∼ 8.7k
AR-LSAT Zhong et al. (2021)
Human
Not available
∼ 2k
Appendix
Table 2: Reasoning datasets that require external annotation to be usable for FOL-autoformalization evaluation.
Dataset
NL
FOL
No. of Instances
RuleTaker/ProofWriter Clark et al. (2020) ; Tafjord et al. (2021)
Rule-based
Recoverable
∼ 500k
PrOntoQA Saparov and He (2022)
Rule-based
Recoverable
Variable
LogicNLI Tian et al. (2021)
Rule-based
Recoverable
∼ 20k
FLD × 2 Morishita et al. (2024)
Rule-based
Recoverable
∼ 100k
Appendix
Table 3: Reasoning datasets whose FOL annotation is not shipped as a text field but it is recoverable, with some effort, by inspecting the dataset’s generation procedure, in order to be usable for FOL-autoformalization evaluation.
Department of Information Management National Sun Yat-Sen University Kaohsiung, Taiwan · Graduate Institute of Data Science & Information Computing National Chung Hsing University Taichung, Taiwan