PIDfest 26

Who Did What to Which Dinosaur? Provenance Tracking with RAiDs
2026-10-28 , Naturalis 02

At Naturalis Biodiversity Center, we deal with a wide variety of research activities: analyses of genomic sequences, studies of physical collection specimens, and images from biomonitoring cameras, to name a few. In turn, those activities produce all sorts of data and involve different people, equipment, software tools, possibly from different projects and institutions. We started exploring how Research Activity Identifiers (RAiDs) can help in linking all those disparate elements together in a pilot study, and we're taking those insights along as we continue our developments around PIDs. So... are you wondering how to implement RAiDs across complex, scary, real-world scenarios? Are you also confused about what qualifies as a "research activity" and what doesn't? And would you like to know how RAiDs come into a picture full of sample IDs, project numbers, and specimen labels flying around? Then this lighting talk is for you!


At Naturalis Biodiversity Center, research activities and the data they produce come in all shapes and sizes: soil samples from far-away lands are analysed in bubbling laboratory vials to extract genomic information; ancient dinosaur bones from the depths of our natural history collection go through X-ray machines to get scanned into 3D-models; all while image and audio sensors hidden in the forest capture wild creatures as they pass by, unaware that they might be featured in the front page of the next biodiversity report... And have you heard of AI? As you can imagine, data pipelines are complicated.

We started collaborating with SURF in a pilot study to implement Research Activity Identifiers (RAiDs) at Naturalis and tackle the challenge of recording "what thing was done by who in the context of which project". In particular, we are interested in how things scale, since many of our use cases involve automated computational workflows that generate or update existing research data (and therefore, we also need to automate the minting and updating of the RAiD records as our systems execute actions). Imagine this simplified scenario:

  1. A scientist registers an upcoming expedition in our project management system
  2. During the expedition, a sample is collected and brought into our collection
  3. The sample is taken to the laboratory, where it undergoes certain analyses
  4. Those analyses produce some interesting raw data, which ends up in our database
  5. Researchers access and refine that data, producing some graphs and other derived outputs
  6. The scientific insights, code and datasets are submitted to external repositories as publications

Each step in this research process involves different people, equipment, and processes, might overlap to some degree and possibly take place under different projects and at different institutions. All 6 steps together can be understood as a research activity—and each of them individually as well! In any case, there's a clear need for keeping track of things (provenance), and persistent identifiers are a key piece of that puzzle. After our initial pilot study, where we explored the potential of RAiDs as a link between different interacting systems, we are continuing developments around RAiDs and PIDs to achieve this vision.

In under 10 minutes, you'll learn about our journey and how we identified several use cases within our institution, automated RAiD API calls in some of our workflows, and even learnt the definition of research activity by heart!

See also: Diagram of one of the use cases for RAiDs at Naturalis. It displays the 6-step scenario showcased in the proposal description. (540.9 KB)

I work as a Data Specialist at Naturalis Biodiversity Center in Leiden. I work in a team devoted to open science, and in particular, I focus on the aspects regarding technical implementation: PIDs, computational reproducibility, digital sovereignty, etc. Next to that, I take part in providing training workshops for researchers about research data management and other similar topics. I'm also passionate about open source software and hardware, since my background is in Robotic Engineering.