PIDfest 26

What Happens When You Connect Every PID You Can Find
2026-10-27 , Naturalis 02

PRSM is a knowledge graph connecting publications, repositories, people, and institutions into 28+ million relationships. It exists because persistent identifiers exist. DOIs, ORCIDs, ROR IDs, and package registries made it possible to link data across OpenAlex, GitHub, and ecosyste.ms into something no single system could produce alone: full research impact analysis across publications and code, ecosystem health scoring, champions discovery, and dependency chains traced to compiled binaries.

As AI agents begin generating research artifacts at machine speed, persistent identifiers become even more valuable. Every synthetic dataset, automated analysis, and machine-written codebase needs provenance, attribution, and traceability.

This talk utilizes a live portfolio analysis to demonstrate the power of today's PID infrastructure, explore why identifiers become even more essential as AI-generated artifacts enter the research ecosystem, and discuss what's needed for identifier infrastructure to scale to trillions of generative artifacts.


PRSM, a knowledge graph connecting 28+ million relationships across publications, repositories, people, and institutions, exists because persistent identifiers exist. This presentation demonstrates the value PIDs provide today and why that value grows as AI transforms research.

Let's start with 59 DOIs and 43 GitHub URLs. PRSM follows the identifiers outward. DOIs lead to authors with ORCIDs, authors lead to institutions with ROR IDs, repositories link to publications via CITATION.cff, dependency manifests connect to packages across 7 ecosystems, and a self-expanding discovery chain builds from there. From those 102 seeds, PRSM assembles a knowledge graph of 3,990 entities and 13,258 relationships: 90,700 citations, 2.18 billion downloads, 3,500 contributors across 144 institutions, and 461 hidden dependencies no manifest declares. It surfaces the full impact story of research: which fields use your work, which institutions depend on it, which software powers it, which people sustain it, and how it all connects. A researcher can show a funder their true footprint across publications and code. An institution can discover contributions it didn't know its people were making. A funder can trace where investment lands across papers, software, communities, and infrastructure. PRSM also identifies the specific people holding ecosystems together: longevity leaders with sustained maintenance activities, collaboration bridges connecting disparate research communities, and cross-project contributors working across multiple codebases and research elements. It calculates bus factors, flags at-risk projects, and models cascade failure paths, giving funders and institutions visibility into sustainability risks that no single data source can surface. None of this exists in any single system. All of it becomes visible because persistent identifiers make it possible to follow the connections.

That value is about to multiply. AI agents are generating synthetic data, automated analyses, and machine-written code at a scale human researchers never could. Every one of those artifacts needs to be traced, attributed, and verified. Without persistent identifiers, there is no provenance chain from an AI-generated result back to the data, models, and code that produced it. PIDs are how AI-assisted research stays trustworthy.

This presentation utilizes a live portfolio analysis to:

  • Demonstrate what becomes visible when you connect PIDs across systems at scale
  • Show how persistent identifiers enable full research impact measurement across publications, software, people, and institutions
  • Explore why PIDs become even more essential as AI-generated artifacts enter the research ecosystem
  • Discuss what identifier infrastructure looks like when it needs to scale to trillions of generated artifacts at zero marginal cost

Jonathan Starr is the Executive Director of the Open Source Endowment, a 501(c)(3) building a community-managed permanent endowment for critical open source infrastructure. He also directs SciOS and the Institute of Open Science Practices, where he coordinates researchers and technologists building sustainable infrastructure for open science. His work spans funding mechanisms, coordination systems, and the shared technical substrate connecting diverse scientific systems.