cs.CLSep 27, 2026

Quizzing the Translation: A Prover-Grounded Evaluation Metric for NL→\rightarrowFOL

Authors: Pu Suo, Ali Emami

Organizations: Emory University

Abstract

A standard pipeline for symbolic reasoning over natural-language problems translates them into first-order logic and invokes a theorem prover. The translation step is the bottleneck: swap "every" for "some" and every inference that follows is corrupted. Yet today's metrics often score more broken translations higher than less broken ones, because BLEU, BERTScore, and Smatch++ reward surface overlap that the worst errors happen to preserve. We introduce SIV, which derives two kinds of probes from the target formula and uses a theorem prover to verify the candidate translation against each. Positive probes are statements the candidate must entail, which detect translations that drop content; contrastive probes are statements the candidate must not entail, which detect translations that assert more than the original. On a controlled pool of perturbed FOLIO translations, the severity of the error accounts for 80% of SIV's score variance, compared with at most 17% for any prior metric. Across six error classes on a disjoint pool, SIV scores the reference above the perturbed candidate in over 99% of pairs. Because each probe is labeled with what it tests, the failure pattern also supplies a labeled error trace, recovering the perturbation class at macro-F1 0.638, nearly double the score-only baseline. On 434 expert-audited real LLM translations, SIV attains the top AUC, uniquely detects and grades expert-labeled major errors, and abstains, rather than mis-scoring, on out-of-vocabulary translations.

Figures & tables

Appendix figures & tables4 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Beyond BLEU: A Case for Redefining Sign Language Translation Benchmarks

    Sep 3, 2026Oline Ranum, Edward Fish, Simon Hadfield +1Gloss-Free Sign Language TranslationSign Language Translation

  2. Last Translation Benchmark

    Sep 3, 2026Vilém Zouhar, Niyati Bafna, Mukund Choudhary +257Machine Translation QualityReproducibility