Reference. HybridProver: Augmenting Theorem Proving with LLM-Driven Proof Synthesis and Refinement

Formal methods play a crucial role in ensuring the reliability of critical systems through rigorous mathematical verification. However, their adoption remains limited due to the labor-intensive nature of manual proof construction. Recent advances in large language models (LLMs) have opened new opportunities for automated theorem proving. Two main paradigms have emerged: stepwise tactic-based generation and whole-proof synthesis. While both approaches have complementary strengths, existing work largely treats them in isolation. In this work, we propose HybridProver, a unified framework that integrates whole-proof synthesis and tactic-based generation through proof sketches as an intermediate representation. This design enables the reuse of partially correct proof structures while effectively combining high-level planning with fine-grained reasoning. We implement HybridProver in Isabelle/HOL and post-train two 7B-scale LLMs on our optimized Isabelle datasets. Experiments on the miniF2F Isabelle benchmark achieved a 73.8% success rate and improved upon the previous state of the art (61.9%), demonstrating that lightweight models, when combined with our approach, can effectively generate Isabelle/HOL proofs without relying on very large LLMs. Ablation studies further analyze the impact of dataset quality, training configurations, and sampling strategies on proof generation.

Cite

Cite as @hu-2025-hybridprover (helia, typst) · \cite{hu-2025-hybridprover} (LaTeX)
BibTeX
bibtex · 8 lines
@misc{hu-2025-hybridprover,
  author = {Jilin Hu and Jianyu Zhang and Yongwang Zhao and Talia Ringer},
  title = {HybridProver: Augmenting Theorem Proving with LLM-Driven Proof Synthesis and Refinement},
  year = {2025},
  month = {5},
  eprint = {2505.15740},
  archiveprefix = {arXiv}
}
hayagriva YAML (typst)
yaml · 11 lines
hu-2025-hybridprover:
  type: misc
  title: 'HybridProver: Augmenting Theorem Proving with LLM-Driven Proof Synthesis and Refinement'
  author:
  - Hu, Jilin
  - Zhang, Jianyu
  - Zhao, Yongwang
  - Ringer, Talia
  date: 2025-05
  serial-number:
    arxiv: '2505.15740'
Cites 72 works (2 here)
With notes (2)

Baldur: Whole-Proof Generation and Repair with Large Language Models first-2023-baldur

PDF · DOI · pldb

Proof Repair Infrastructure for Supervised Models: Building a Large Proof Repair Dataset reichel-2023-proof

We report on our efforts building a new, large proof-repair dataset and benchmark suite for the Coq proof assistant. The dataset is made up of Git commits from open-source projects with old and new versions of definitions and proofs aligned across commits. Building this dataset has been a significant undertaking, highlighting a number of challenges and gaps in existing infrastructure. We discuss these challenges and gaps, and we provide recommendations for how the proof assistant community can address them. Our hope is to make it easier to build datasets and benchmark suites so that machine-learning tools for proofs will move to target the tasks that matter most and do so equitably across proof assistants.
DOI
External (70)
hu-2025-hybridprover reference entries/refs/hu-2025-hybridprover/hu-2025-hybridprover.hel