OpenAI and the Pacific Northwest National Laboratory have released a benchmark designed to measure how well AI handles one of the federal government's most stubborn paperwork bottlenecks.
The two organizations jointly built DraftNEPABench, an evaluation tool that tests AI coding agents on National Environmental Policy Act (NEPA) documents - the mandatory environmental impact assessments required before most major infrastructure projects can break ground. Their testing suggests AI assistance could reduce NEPA drafting time by up to 15%. PNNL is a Department of Energy national lab, not a vendor with a product to sell, which lends the collaboration more credibility than a typical AI company press release. The benchmark is framed as a foundation for modernizing how federal agencies handle infrastructure reviews.
Federal permitting has been a genuine constraint on energy, broadband, and transportation deployment for years - something successive administrations have complained about without fixing. A rigorous benchmark matters here because AI productivity claims are easy to float and hard to verify; a standardized test at least puts a floor under the marketing. Worth noting: the 15% figure covers drafting time specifically, not the full permitting cycle, which also includes public comment periods, legal review, and interagency coordination that no AI agent is touching.
Whether federal agencies actually deploy any of this is a separate question from whether the benchmark says they could.