cs.AISep 30, 2026

Fyan: A Human--AI Harness with Semantic Auditing for Document-Level Formalization

Authors: Wei Zhao, Yangshuo Zou, Chengxiang Ding, Yifan Wu, Xuchuan Wang, Zimu Mao, Lei Zhang, Tao Luo

Organizations: School of Mathematical Sciences, Shanghai Jiao Tong University, Shanghai 200240, China · University of California, Berkeley, CA 94720, USA · Zhiyuan College, Shanghai Jiao Tong University, Shanghai 200240, China · School of Biomedical Engineering, Shanghai Jiao Tong University, Shanghai 200240, China · School of Materials Science and Engineering, Shanghai Jiao Tong University, Shanghai 200240, China · Institute of Natural Sciences, MOE-LSC, Shanghai Jiao Tong University, Shanghai 200240, China · CMA-Shanghai, Shanghai Jiao Tong University, Shanghai 200240, China

Abstract

We present FYAN, a human--AI harness for document-level mathematical formalization. Rather than treating theorems in isolation, FYAN coordinates an end-to-end workflow spanning specification, proof planning, logical review, Lean proof construction, knowledge curation, and validation, with support for independent supervision and human guidance. A central component is evidence-grounded semantic auditing, which assesses whether formal statements faithfully preserve their informal specifications. A language model constructs structured evidence over local correspondences, omissions, scope, and logical relations, while a deterministic validator checks this evidence and produces reproducible judgments. When a substantive but admissible deviation is accepted, FYAN requires an explicit proof-transfer obligation connecting the formal statement back to a source-facing interpretation. With the same model (DeepSeek-V4.1-Flash) in every stage, FYAN proves 86 of 143 FormalTCS theorems under a strict Lean check, against 69 for a general agent harness, and raises the natural-language proof score from 0.501 to 0.851. On ConsistencyCheck, its semantic audit catches more inconsistent statements than a direct LLM judge, both on labels verified against the source (recall 0.777 vs. 0.636) and on the original labels (0.873 vs. 0.820), and localizes each mismatch it reports to a specific hypothesis, conclusion, or scope. FYAN also built ODENumLib, a 9,355-line Lean library for the numerical analysis of ordinary differential equation.

Figures & tables

Appendix figures & tables10 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Beyond Compilation: Evaluating Faithful Natural-Language-to-Lean Statement Formalization

    Jun 30, 2026Ke Zhang, Patricio Gallardo Candela, Sudhir Murthy +3Theorem ProvingExplanation Faithfulness

  2. AIProver: Agentic Auto-Formalization of Mathematical Research via Certificate-Driven Evolving Harness

    Oct 4, 2026Prithwish Jana, Viet Bach Hoang, Logan Luna +9Theorem Proving

  3. Evaluation of LLMs for Mathematical Formalization in Lean

    Jun 4, 2026Tyson Klingner, Drew Bladek, Escher Crawford +6Theorem ProvingFormalization