Citizen Portal
Sign In

Get Full Government Meeting Transcripts, Videos, & Alerts Forever!

Get email alerts on the Proteomics Data Infrastructure topic

No spam. Unsubscribe anytime.

NCI seminar highlights Proteomic Data Commons: centralized proteomics, APIs and cloud pipelines

National Cancer Institute (NCI) data science seminar series · June 25, 2026
AI-Generated Content: All content on this page was generated by AI to highlight key points from the meeting. For complete details and context, we recommend watching the full video. so we can fix them.

Summary

At an NCI data science seminar, PDC staff demonstrated the Proteomic Data Commons (PDC), a centralized repository and analysis platform for mass‑spectrometry proteomic datasets, its GraphQL APIs, embedded visualization tools and links to cloud compute platforms that let researchers analyze large cancer proteomics collections in place.

The National Cancer Institute presented the Proteomic Data Commons (PDC) and its associated tools during a data science seminar, outlining how the platform centralizes proteomic data and provides analysis capabilities for cancer researchers. Dr. Ratna Tangadoo, the PDC project manager and lead, described the repository’s goals: to reduce fragmentation of proteomic resources, harmonize processing and accelerate multi‑omic research.

Tangadoo said the PDC focuses on mass spectrometry–based proteomics, including whole‑proteome measurements and post‑translational modifications such as phosphorylation, acetylation, glycosylation and ubiquitination. She described common acquisition types (data‑dependent and data‑independent acquisition, label‑free and multiplexed labeled experiments like iTRAQ/TMT) and explained that the PDC applies a common data analysis pipeline (CDAP) to harmonize outputs from different labs and instruments.

The portal aggregates data from major programs, Tangadoo said, including CPTAC, the International Cancer Proteogenomic Consortium (ICPC), Department of Defense and VA studies, and the Children’s Brain Tumor Network. She reported current holdings of roughly 137 studies, about 37 terabytes of data across more than 100,000 files, covering approximately 19 cancer types from 12 primary sites. Monthly downloads were described as roughly 30–40 TB with higher spikes at times.

Users can browse studies via an Explore page, Tangadoo said, using faceted filters for study, case, clinical or experimental metadata and gene/protein identifiers. Study summary pages show protocols, experimental design and available files. The portal also provides case and gene summary pages to inspect a donor’s samples or a gene’s distribution across studies.

For interactive visualization, the PDC embeds the Broad Institute’s Morpheus heat map viewer to display processed protein quantitation overlaid with clinical metadata. Tangadoo emphasized programmatic access as well: "Every piece of information we collect and present on the portal is also available programmatically through the APIs," she said, describing GraphQL endpoints and an integrated GraphQL playground for testing queries.

To help researchers avoid transferring very large raw datasets, Tangadoo outlined integration with cloud platforms in the Cancer Research Data Commons ecosystem — Broad/Terra (FireCloud), Seven Bridges, and ISB’s Cancer Genomics Cloud — and noted that some processed results are available in Google BigQuery to facilitate cross‑omics correlation analyses.

She also reviewed the submission process: a menu‑driven submission workspace with validation aligned to NCI data standards and the NIH data‑sharing policy, plus documentation, tutorials and help‑desk support. The portal URL provided in the seminar was pdc.cancer.gov.

The session concluded with Tangadoo acknowledging partner teams at ICF, the University of Washington and others, NCI program leads, and funders. The presentation then moved to Dr. Bing Zang for a follow‑on discussion of proteogenomic integration and the LinkOmics KB resource.

Next steps: presenters encouraged interested researchers to use the portal, consult the documentation, and contact PDC support for submission or technical questions; presenters noted plans for additional export features for very large interactive tables based on user feedback.