Citizen Portal

Get email alerts on the Ai Datasets Hugging Face topic

No spam. Unsubscribe anytime.

Library of Congress posts datasets and tools on Hugging Face to spur AI research

Library of Congress · September 9, 2026

Summary

The Library of Congress said it is publishing component datasets and combined data and developing tools on Hugging Face so developers and researchers can build AI applications using Library, Smithsonian and NARA collections.

Natalie Buda Smith described how the Library, Smithsonian and National Archives are posting component datasets and a combined dataset on the Hugging Face platform so developers can access and experiment with the data. She said developers can download data or build tools on top of it and the Library hopes a community will form around reuse.

Smith emphasized that the hardest work was preparing and redigitizing legacy collections so that named-entity recognition and other AI services could meaningfully link records. "A lot of the data behind AI is really where the work goes," she said, noting that data quality determines outcomes for downstream AI applications.

AI generated

The text on this page is AI generated. Summaries, highlights, analysis, and video transcripts are all produced from the original source material.

AI can make mistakes, so if you spot one, and we will fix it for everyone.

Note: the source content is unaltered by us. Any content source we link to, be it a video, an audio recording, or a document, is presented exactly as its publisher released it. That publisher is usually a government body, sometimes an individual official or another organisation.

Source