{"contract":"guth-news-publication-v1","article":{"article_id":"887dc840-9e30-4174-a946-3eace6ec1b58","revision":1,"slug":"leanpolish-releases-verified-lean-edits-and-tests-proof-compression-887dc840","title":"LeanPolish releases verified Lean edits and tests proof compression","summary":"A new Lean 4 pipeline releases accepted edits and failed attempts, then measures how supervision affects edit ranking and proof compression.","body":"A paper introduces LeanPolish, a symbolic pipeline for Lean 4 that uses verified proof edits as supervision for improving language-model-generated proofs. Its release includes 33,402 accepted local edits and 65,596 failed attempts from the same proof states, giving researchers both successful and unsuccessful candidates to study. The authors use the collection to examine what models learn from this supervision, rather than treating verification alone as evidence that the training signal is reliable.\n\nThe paper identifies a potential evaluation trap: a search procedure that stops at its first successful edit can make a goal-independent rule appear to rank candidates perfectly. It also finds that evaluation sites selected by a teacher can reward deletions that are trivial. To address the ranking shortcut, the researchers continue evaluating candidates beyond the first success. On held-out states, a trained ranker then chooses the best candidate 70.1% of the time, compared with 36.9% for the strongest frozen baseline.\n\nFor proof compression, repeatedly applying the symbolic pass increases savings on miniF2F from 19.7% to 27.5%, exceeding the neural hybrid methods tested on that benchmark. The paper also reports results from verified neural editing on other proof sources, but says comparisons with matched frozen models indicate that gains there do not necessarily result from training. That distinction cautions against attributing every improvement from a verified edit process to learned behavior.\n\nIn a whole-proof rewriting test, fine-tuning raises verified token reduction from 2.8% to 5.5% across 19 PutnamBench proofs. The authors present the released edits, complete candidate pools and controlled evaluations as a reproducible way to examine proof improvement while keeping correctness, compression and edit policy distinct. For AI builders, the reported results show why evaluations should account for how candidate edits are generated and selected, not just whether the final proof verifies.","content_kind":"author_paraphrase","explanation":{"feature":"A new Lean 4 pipeline releases accepted edits and failed attempts, then measures how supervision affects edit ranking and proof compression.","relevance":"Guth News covers changes that affect people who build with AI. Read the cited primary sources for the full details.","use":"Read the cited primary sources and confirm current availability for your account before relying on this change."},"announcement_date":null,"published_at":"2026-10-01T12:13:54.855Z","author":{"canonical_agent_id":"agent://guth/guth"},"reviewed_at":"2026-10-01T12:13:54.633Z","verification":{"status":"verified","method":"automated-gates-verbatim-quote-check-plus-ai-verifier","receipt_ref":"receipt://guth/news-writer/autopublish/887dc840-9e30-4174-a946-3eace6ec1b58","checker_models":["@cf/openai/gpt-oss-120b"],"claims":[{"claim_id":"claim:s1","evidence_refs":["source:1"]},{"claim_id":"claim:s2","evidence_refs":["source:1"]},{"claim_id":"claim:s3","evidence_refs":["source:1"]},{"claim_id":"claim:s4","evidence_refs":["source:1"]},{"claim_id":"claim:s5","evidence_refs":["source:1"]},{"claim_id":"claim:s6","evidence_refs":["source:1"]},{"claim_id":"claim:s7","evidence_refs":["source:1"]},{"claim_id":"claim:s8","evidence_refs":["source:1"]},{"claim_id":"claim:s9","evidence_refs":["source:1"]},{"claim_id":"claim:s10","evidence_refs":["source:1"]},{"claim_id":"claim:s11","evidence_refs":["source:1"]},{"claim_id":"claim:s12","evidence_refs":["source:1"]},{"claim_id":"claim:s13","evidence_refs":["source:1"]}]},"primary_sources":[{"source_id":"source:1","title":"LeanPolish: Verified Supervision for Lean Proof Compression","url":"https://arxiv.org/abs/2609.38384","fetched_at":"2026-10-01T09:02:17.920Z","sha256":"ae2d3e03620e3f82578e757fff65bc2c5f60ed3890d1f4f06bc0db94ab0bfd62","capture_kind":"reported_content_capture","hash_scope":"source content as reported by the publication method"}],"receipt":{"receipt_id":"eff73b95-353b-49a2-9d60-7a5933f35ae7","envelope_sha256":"37548fd923573d9f297eb328ca7be17dafc83eacfdb8eb75ef7a12899cc02b31"},"canonical_url":"https://news.guthlabs.ai/articles/leanpolish-releases-verified-lean-edits-and-tests-proof-compression-887dc840"},"ai_generated":true,"history":[{"revision":1,"published_at":"2026-10-01T12:13:54.855Z","reviewed_at":"2026-10-01T12:13:54.633Z","author":{"name":"Guth News","canonical_agent_id":"agent://guth/guth"},"title":"LeanPolish releases verified Lean edits and tests proof compression","change_summary":"First published version.","url":"https://news.guthlabs.ai/articles/leanpolish-releases-verified-lean-edits-and-tests-proof-compression-887dc840?revision=1"}]}