2026-10-29 –, Naturalis 01
Connecting digital objects across the data repository landscape is critical for ensuring that open data are discovered and reused. Global adoption of common PIDs across repositories facilitates broad cross-repository discoverability, enables impact assessment and researcher credit, and supports compliance monitoring by funders and institutions. The NIH Generalist Repository Ecosystem Initiative brings seven generalist repositories together - working in a “coopetition” model - with a key objective to connect digital objects by enhancing dataset metadata quality, completeness, and consistency. Over the past 4 years GREI repositories have developed and implemented a common metadata recommendation that relies heavily on PIDs including DataCite DOIs, ORCID, and ROR. This presentation will showcase examples of GREI repository PID implementation, use cases, and repository best practices, as well as the impact of PID-based collaborations with DataCite, ROR, and NLM, and the value of common PIDs for author, affiliation, and funding metadata in better connecting data repositories.
Across the global research landscape, keeping knowledge connected requires robust, common standards across platforms. The US National Institutes of Health (NIH) Generalist Repository Ecosystem Initiative (GREI) represents a unique "coopetition" model, bringing together seven established generalist repositories (Dataverse, Dryad, Figshare, Mendeley Data, Open Science Framework, Vivli, and Zenodo) to enhance data repository infrastructure support for improved NIH data sharing and reuse. A key component of this initiative is the development and implementation of common metadata standards including the use of PIDs to improve interoperability, discovery, and long‑term stewardship of research outputs.
This presentation will detail GREI’s collaborative journey to implement common Persistent Identifiers (PIDs) including DataCite Digital Object Identifiers (DOIs) for datasets, ORCID iDs for researchers, and Research Organization Registry (ROR) IDs for author affiliations and funders across the GREI repositories. We will explore how establishing common metadata recommendations that include best practices for PID implementation across traditionally competitive repositories eliminates data silos and directly supports the FAIR (Findable, Accessible, Interoperable, and Reusable) principles.
Attendees will learn about the practical challenges involved in aligning diverse repository architectures to support standardized PID capture, validation, and display. This presentation will highlight specific use cases and examples of PID implementation at GREI repositories, including:
- How GREI established consistent ROR implementation practices across repositories in partnership with ROR to support tracking of research outputs by both funding source and author affiliation.
- How optimized ORCID integration for author and co-author identification reliably connects researchers to their datasets, strengthening attribution and enabling more complete researcher profiles.
- How DOIs function as the critical connection for data citation, linking datasets to corresponding publications and funding and allowing for robust assessment for both impact and compliance.
Furthermore, we will also highlight partnerships, outcomes, and future directions for PID adoption within the GREI repositories, including:
- Our partnership with the NLM dataset catalog, which allows NIH to harvest GREI repository metadata and provide a cross-repository catalog of NIH-funded datasets.
- DataCite’s GREI dashboards, which leverage ROR IDs in GREI repository metadata to track datasets by funding source within the NIH.
- Ongoing efforts to strengthen PID implementation at each repository, encourage PID use by researchers, and re-curate legacy metadata to improve PID quality and completeness.
In addition, we will share recent GREI resources including metadata recommendations and approaches for measuring metadata completeness and improving metadata quality. Audience poll questions throughout the session in addition to the Q&A discussion will allow participants to share experiences and perspectives and provide feedback on future GREI work. By sharing both our successes and the lessons learned in navigating repository infrastructure barriers, this session will provide research infrastructure providers, repository managers, researchers, and librarians with actionable strategies for embedding PIDs in their research workflows and building stronger, more connected data ecosystems.
Ana is the Head of Customer Engagement for Government, Funders & Nonprofits at Digital Science and manages the Government and Funder program at Figshare. She is a PI in the NIH Generalist Repository Ecosystem Initiative and currently co-chairs community engagement activities for the program. She holds a PhD in Cognitive Psychology and has worked in open data and open science for more than 10 years.
Kristi Holmes, PhD, is a leader in knowledge management, FAIR data, and research impact at Northwestern University. She is Associate Dean for Knowledge Management and Strategy and Director of the Galter Health Sciences Library and Learning Center. She directs Informatics and Data Science within NUCATS and co‑leads its strategic management, and she serves as Chief of Knowledge Management in the Institute for Artificial Intelligence in Medicine. Her research focuses on improving information access, supporting data‑driven discovery, and addressing emerging issues in AI and research. She is dedicated to shared knowledge ecosystems and evidence‑based approaches that strengthen transparency, collaboration, and research impact across the scientific community.
John Chodacki is Director of the University of California Curation Center (UC3) at California Digital Library (CDL). As UC3 Director, John works across the UC campuses and the broader community to ensure that CDL’s digital curation services meet the emerging needs of the scholarly community – including digital preservation, persistent identifiers, data management, and data publishing. Prior to CDL, John oversaw product development activities at publishing organizations such as O’Reilly Media, Safari Books Online, and PLOS. He has served on the board and/or steering committees of DataCite, Crossref, FORCE11, ROR (Research Organization Registry), COUNTER, Collaborative Knowledge (Coko) Foundation, Open Citations, Metadata 20/20, Make Data Count (MDC), Collaborative Metadata Enrichment (COMET), RDA-US, and The Carpentries.
Matt Buys is Executive Director of DataCite, a global non-profit organization providing persistent identifier (PID) infrastructure to support open and connected research. He leads the organization’s strategy across community governance, services, and operations, with a focus on strengthening trusted, sustainable, and community-driven scholarly infrastructure serving stakeholders in over 65 countries. 
With over 15 years of experience in the research and scholarly communications ecosystem, Matthew has worked closely with institutions, funders, and international partners to advance open science practices and collaborative infrastructure. 
He is also an active contributor to governance in the open research landscape, serving as Board Chair of the Center for Open Science, Vice Chair of the DOI Foundation, and a board member of FORCE11, supporting global efforts to promote openness, persistence, and trust in research.
Audrey has many years of experience developing websites and digital repositories related to libraries and to scientific publishing, with previous roles in software development at Europe PMC, PubMed Central, and the University of Delaware Library. She holds a Master's in Library and Information Science from Drexel University.