Reference. Getting More out of Large Language Models for Proofs
Large language models have the potential to simplify formal theorem proving and make it more accessible. But how to get the most out of these models is still an open question. To answer this question, we take a step back and explore the failure cases of these models using common prompting-based techniques. Our talk will discuss these failure cases and what they can teach us about how to get more out of these models.
Cite
Cited by (1)
Cobblestone: A Divide-and-Conquer Approach for Automating Formal Verification kasibatla-2026-cobblestone
Cites 13 works (2 here)
With notes (2)
Baldur: Whole-Proof Generation and Repair with Large Language Models first-2023-baldur
QED at Large: A Survey of Engineering of Formally Verified Software ringer-2019-qed
Development of formal proofs of correctness of programs can increase actual and perceived reliability and facilitate better understanding of program specifications and their underlying assumptions. Tools supporting such development have been available for over 40 years, but have only recently seen wide practical use. Projects based on construction of machine-checked formal proofs are now reaching an unprecedented scale, comparable to large software projects, which leads to new challenges in proof development and maintenance. Despite its increasing importance, the field of proof engineering is seldom considered in its own right; related theories, techniques, and tools span many fields and venues. This survey of the literature presents a holistic understanding of proof engineering for program correctness, covering impact in practice, foundations, proof automation, proof organization, and practical proof development.
External (11)
- Conversational Automated Program Repair (2023)
- HyperTree Proof Search for Neural Theorem Proving (2022)
- Diversity-Driven Automated Formal Verification (2022)
- Self-Consistency Improves Chain of Thought Reasoning in Language Models (2022)
- Compilable Neural Code Generation with Compiler Feedback (2022)
- Program Synthesis with Large Language Models (2021)
- TacticZero: Learning to Prove Theorems from Scratch with Deep Reinforcement Learning (2021)
- TacTok: semantics-aware proof synthesis (2020)
- Generative Language Modeling for Automated Theorem Proving (2020)
- MPNet: Masked and Permuted Pre-training for Language Understanding (2020)
- HOList: An Environment for Machine Learning of Higher Order Logic Theorem Proving (2019)