Reference. AlphaVerus: Bootstrapping Formally Verified Code Generation through Self-Improving Translation and Treefinement
Automated code generation with large language models has gained significant traction, but there remains no guarantee on the correctness of generated code. We aim to use formal verification to provide mathematical guarantees that the generated code is correct. However, generating formally verified code with LLMs is hindered by the scarcity of training data and the complexity of formal proofs. To tackle this challenge, we introduce AlphaVerus, a self-improving framework that bootstraps formally verified code generation by iteratively translating programs from a higher-resource language and leveraging feedback from a verifier. AlphaVerus operates in three phases: exploration of candidate translations, Treefinement – a novel tree search algorithm for program refinement using verifier feedback, and filtering misaligned specifications and programs to prevent reward hacking. Through this iterative process, AlphaVerus enables a LLaMA-3.1-70B model to generate verified code without human intervention or model finetuning. AlphaVerus shows an ability to generate formally verified solutions for HumanEval and MBPP, laying the groundwork for truly trustworthy code-generation agents.
Cite
Cites 61 works (2 here)
With notes (2)
Baldur: Whole-Proof Generation and Repair with Large Language Models first-2023-baldur
Verus: Verifying Rust Programs using Linear Ghost Types lattuada-2023-verus
The Rust programming language provides a powerful type system that checks linearity and borrowing, allowing code to safely manipulate memory without garbage collection and making Rust ideal for developing low-level, high-assurance systems. For such systems, formal verification can be useful to prove functional correctness properties beyond type safety. This paper presents Verus, an SMT-based tool for formally verifying Rust programs. With Verus, programmers express proofs and specifications using the Rust language, allowing proofs to take advantage of Rust’s linear types and borrow checking. We show how this allows proofs to manipulate linearly typed permissions that let Rust code safely manipulate memory, pointers, and concurrent resources. Verus organizes proofs and specifications using a novel mode system that distinguishes specifications, which are not checked for linearity and borrowing, from executable code and proofs, which are checked for linearity and borrowing. We formalize Verus’ linearity, borrowing, and modes in a small lambda calculus, for which we prove type safety and termination of specifications and proofs. We demonstrate Verus on a series of examples, including pointer-manipulating code (an xor-based doubly linked list), code with interior mutability, and concurrent code.
External (59)
- Automated Proof Generation for Rust Code via Self-Evolution (2024)
- AutoVerus: Automated Proof Generation for Rust Code (2024)
- Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters (2024)
- Lean-STaR: Learning to Interleave Thinking and Proving (2024)
- From Decoding to Meta-Generation: Inference-time Algorithms for Large Language Models (2024)
- miniCodeProps: a Minimal Benchmark for Proving Code Properties (2024)
- Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models (2024)
- When Can LLMs Actually Correct Their Own Mistakes? A Critical Survey of Self-Correction of LLMs (2024)
- Towards Neural Synthesis for SMT-Assisted Proof-Oriented Programming (2024)
- A Survey on Deep Learning for Theorem Proving (2024)
- Easy-to-Hard Generalization: Scalable Alignment Beyond Human Supervision (2024)
- Towards AI-Assisted Synthesis of Verified Dafny Methods (2024)
- Occasionally Secure: A Comparative Analysis of Code Generation Assistants (2024)
- An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models (2024)
- DafnyBench: A Benchmark for Formal Software Verification (2024)
- HumanEval-Verus: Hand-written examples of verified Verus code derived from HumanEval (2024)
- V-STaR: Training Verifiers for Self-Taught Reasoners (2024)
- Code Llama: Open Foundation Models for Code (2024)
- Qwen2.5: A Party of Foundation Models (blog) (2024)
- Self-Taught Evaluators (2024)
- SGLang: Efficient Execution of Structured Language Model Programs (2024)
- Clover: Closed-Loop Verifiable Code Generation (2023)
- SWE-bench: Can Language Models Resolve Real-World GitHub Issues? (2023)
- Can LLM Replace Stack Overflow? A Study on Robustness and Reliability of Large Language Model Code Generation (2023)
- Scaling Relationship on Learning Mathematical Reasoning with Large Language Models (2023)
- FacTool: Factuality Detection in Generative AI - A Tool Augmented Framework for Multi-Task and Multi-Domain Scenarios (2023)
- CRITIC: Large Language Models Can Self-Correct with Tool-Interactive Critiquing (2023)
- Let's Sample Step by Step: Adaptive-Consistency for Efficient Reasoning and Coding with LLMs (2023)
- Tree of Thoughts: Deliberate Problem Solving with Large Language Models (2023)
- RL4F: Generating Natural Language Feedback with Reinforcement Learning for Repairing Model Outputs (2023)
- Self-Edit: Fault-Aware Code Editor for Code Generation (2023)
- Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code Generation (2023)
- Teaching Large Language Models to Self-Debug (2023)
- Self-Refine: Iterative Refinement with Self-Feedback (2023)
- Large Language Models and Simple, Stupid Bugs (2023)
- Check Your Facts and Try Again: Improving Large Language Models with External Knowledge and Automated Feedback (2023)
- Understanding the limits of AI coding (2023)
- StarCoder: may the source be with you! (2023)
- Toolformer: Language Models Can Teach Themselves to Use Tools (2023)
- A Survey of Deep Learning for Mathematical Reasoning (2022)
- Do Users Write More Insecure Code with AI Assistants? (2022)
- Generating Sequences by Learning to Self-Correct (2022)
- Self-Consistency Improves Chain of Thought Reasoning in Language Models (2022)
- Competition-level code generation with AlphaCode (2022)
- Formal Mathematics Statement Curriculum Learning (2022)
- The Effects of Reward Misspecification: Mapping and Mitigating Misaligned Models (2022)
- STaR: Bootstrapping Reasoning With Reasoning (2022)
- Asleep at the Keyboard? Assessing the Security of GitHub Copilot’s Code Contributions (2021)
- Program Synthesis with Large Language Models (2021)
- Evaluating Large Language Models Trained on Code (2021)
- TacTok: semantics-aware proof synthesis (2020)
- Generative Language Modeling for Automated Theorem Proving (2020)
- The Coq Proof Assistant (2020)
- Reinforcement Learning of Theorem Proving (2018)
- Dependent types and multi-monadic effects in F* (2016)
- ProverBot 9000 : Neural Networks for Proof Assistance (2016)
- Dafny: An Automatic Program Verifier for Functional Correctness (2010)
- Crafting Papers on Machine Learning (2000)
- Lean theorem prover