2026-10-27 –, LUMC02
Are you involved in or assisting researchers with either making datasets produced in HPC environments available for reuse or trying to find the data you need for your work and finding it utterly cumbersome? Have you ever wondered about how AI applications can best find the training data required for a specific task across infrastructural boundaries? Indeed, in many domains, data are created at large, distributed HPC centers but are not operationally assigned globally resolvable PIDs at the point of production, making it hard to identify, track, and reuse these datasets across infrastructures.
In this BoF, you will discuss how PIDs could be integrated directly into HPC workflows at simulation runtime, explore current technical ideas and challenges, and discuss what would be needed to make data practically immediately usable across federated HPC systems. The focus is on practical questions around data access, metadata, provenance, and interoperability.
Are you envisioning a data ecosystem in which data stemming from HPC applications are practically immediately findable, traceable, and reusable via their PID without having to wait for post-processing steps? International HPC centers, sometimes connected to AI infrastructures, are the go-to tools when it comes to producing cutting-edge scientific datasets, but PIDs are often not part of the workflows employed to run these HPC applications. As a result, you have probably found that data remain difficult to search, find and access if you are not part of the team having produced the data.
In this session, you will explore what it would take to assign and register PIDs directly during data production, for example at simulation runtime, to enable global findability - an aspect especially relevant for existing and upcoming AI applications and workflows. An important aspect of this is how such data could be made available in storage systems close to where they are created and accessed across federated environments, including interoperability with existing data spaces like EOSC, DestinE, or institutional infrastructures. We will also consider cases where data are not openly available, where PIDs can still support discovery, access information, and FAIR reuse.
By joining us in this BoF, you will:
- discuss technical approaches to integrate PID assignment into HPC workflows for rapid findability of data assets
- better understand which minimal metadata are needed to support FAIR use (e.g. provenance, access, authorship, licensing), especially for AI applications
- exchange experiences and challenges with others working on similar problems across infrastructures
The session is designed as an open discussion. You are invited to share your own experiences, ongoing work and open questions and to explore together how PID-enabled workflows could support large-scale and distributed research - both classical and AI-enabled.