Citizen Portal
Sign In

Get Full Government Meeting Transcripts, Videos, & Alerts Forever!

Get email alerts on the Nepa Tech topic

No spam. Unsubscribe anytime.

PNNL’s NEPA Text Corpus released publicly; team highlights limits on scope, legal checks and agency rollout

Pacific Northwest National Laboratory webinar · June 3, 2026
AI-Generated Content: All content on this page was generated by AI to highlight key points from the meeting. For complete details and context, we recommend watching the full video. so we can fix them.

Summary

PNNL said its NEPA Text Corpus (NEPA Tech v2) is publicly available but currently limited to federal agency‑authored documents; presenters said the tools are retrieval‑based, do not track litigation, and are intended to assist—not replace—agency legal and policy analysis.

PNNL researchers said during a webinar that the NEPA Text Corpus (NEPA Tech v2) has been released publicly on Hugging Face but that the dataset and the Permit AI applications have important scope and legal limitations.

Samira Horwala Pitana, the Permit AI project principal investigator, said NEPA Tech presently includes federal agency‑authored documents harvested from sources such as the EPA’s EIS APIs, Department of Energy releases and Bureau of Land Management materials, and does not yet include state, local or project‑applicant documents. “Not at this moment for the state and local entities, and also not for the project applicants,” Samira said when asked whether metadata included state or consultant‑produced documents.

The team stressed that Permit AI’s current capabilities are focused on information retrieval rather than automated legal compliance checks. Asked whether the model could identify compliance with all relevant environmental statutes, Samira said the platform is limited to retrieving information that appears in historical NEPA documents and cannot yet make compliance determinations or automatically flag statutory gaps.

On litigation and data currency, Samira said the applications do not currently surface whether a document has been subject to litigation or automatically flag outdated or legally superseded material; the project team described adding litigation‑tracking as a potential future feature.

Regarding accuracy and hallucination risk, the presenters acknowledged that large language models can hallucinate if not grounded in domain data. To reduce hallucinations, PNNL uses retrieval‑augmented generation and metadata extraction pipelines, and they publish benchmarks such as the Draft NEPA Bench on PNNL’s Hugging Face account to evaluate model performance. Samira said these methods have produced lower hallucination rates in SearchNEPA and CheckNEPA relative to ungrounded frontier models, but she also said error rates vary by application and remain a focus of ongoing validation.

The presenters invited federal agencies to coordinate with DOE to become beta testers; they said most applications remain limited to federal users during the beta phase, though parts of the dataset and benchmarks are public. They also encouraged partners to consider technology transfer, sponsored research or collaborative agreements to support wider deployment.

The session ended with an offer to follow up by email for unanswered questions and to continue coordinating with agency partners on data hosting, additional agency coverage, and future public releases.