2026-10-27 –, LUMC02
As data volumes and complexity grow the question of how to scale Persistent Identifier (PID) systems becomes increasingly important. In practice, data is often organized in hierarchies - from collections to datasets to individual files - raising a key challenge: Which level of detail truly benefits users whilst keeping the systems operational?
This session invites participants to share real-world use cases, challenges, and approaches to PID assignment across different domains. Depending on the context, PIDs at the collection or publication level may be sufficient for discovery and citation, while more granular identifiers are required for reproducibility, automated workflows, and machine actionability. However, increased granularity introduces overhead and complexity, potentially leading to unsustainable proliferation of identifiers.
We aim to bring together diverse perspectives from different stakeholders to explore how decisions about PID granularity are made in practice, and how to balance the trade-off between too many and too few PIDs.
Decisions about PID granularity are rarely one-size-fits-all. Different communities, infrastructures, and use cases require different approaches to identifying and referencing data. In some contexts, assigning PIDs at the collection or dataset level is sufficient to support discovery and citation. In others, especially where large volumes of heterogeneous data are processed automatically, more fine-grained identifiers at the file or sub-file level may be necessary.
At the same time, increasing the number of PIDs introduces challenges: technical overhead, management complexity, and questions about long-term sustainability. Conversely, insufficient granularity may limit reproducibility, traceability, and machine actionability. Striking the right balance is therefore a key challenge for scalable PID systems.
This session aims to provide an open space for participants to exchange experiences from their own domains, including data infrastructures, research projects, and operational workflows. We explicitly invite contributions that highlight both practical solutions and ongoing challenges.
In this session, we will discuss the following - and your - questions
- Which - conceptual and technical - approaches you follow to aggregate and assign datasets and files with PIDs in your disciplines ?
- What level of data granularity in your domain currently receives PIDs—and why?
- Which use cases (e.g. citation, discovery, reuse, automated workflows) require more or less granular identifiers?
- Where do you see clear benefits (or drawbacks) of more fine-grained identifiers?
- What are the practical or technical limits to scaling PID systems in your context?
- How do different use cases (e.g. citation, discovery, automated processing) influence PID assignment?
- How do you ensure that both humans and machines can navigate to the “right” level of data?
- How do you balance “too many” versus “too few” PIDs?
The goal of this BoF is not to reach a single solution, but to surface common patterns, trade-offs, and open questions that can inform future developments in PID practices and infrastructure.