PIDfest 26

Take your best (snap)shot - lessons from the ORION initiative
2026-10-27 , LUMC04

For data providers and data users there is increasing interest in using large scale snapshots of entire datasources. For data providers this can relieve the stress and cost of running APIs. For users it allows for new kinds of analysis, creating global benchmarks and identifying gaps. PIDs are at the centre of these analyses, allowing us to link large datasources together reliably. ORION-DBs is an effort to make data dumps more usable by exploiting the capability of Google BigQuery. Setting up this federated set of cloud data resources we have identified patterns in which datasets are easy to ingest, which are easy to use and how they can be combined. In this BoF session we focus on a facilitated discussion of what the optimum data snapshot looks like in format, structure, regularity, locations and how these can help (or hinder) the ability of users to fully exploit these datasets.


The benefits of open research information, with PIDs at the center, are only just beginning to be realised. This operates from the small scale, where data access is easier than it has ever been, to very large scale analysis of whole data systems. At the same time, data providers are under assault from a barrage of scrapers sometimes using APIs, sometimes not. For providers, data snapshots can provide a way of shifting large scale usage to more appropriate systems. Both providers and users can benefit from well designed and easily accessible snapshots of whole datasets.

Shared infrastructure, where open datasets are hosted and made publicly available on a platform anyone can access for their own analyses, has the potential of opening up capacity for large-scale data analysis. The ORION-DBs initiative is currently facilitating this through Google Big Query, where multiple open datasets are made available in a collaborative effort of various groups (including the MultiObs team - continuing the work of the InSySPo team at Campinas, SUB Göttingen, Sesame Open Science and CWTS amongst others).

Google Big Query has emerged as a powerful tool. In particular it provides a capacity for combining and working on these large datasets at scale. It allows for providers to share data without requiring them to shoulder the burden of compute provision and costs. It allows users to run queries (at their own expense) without the need for access to specific hard-or software.

However, there are legitimate concerns around depending on US ‘big tech’ for shared community systems. Alternatives are emerging but do not yet deliver the collective functionality that Google Big Query offers.

Providing access to open research information in shared infrastructures, whether through Big Query or future alternatives, requires access to the datasources. Many organizations provide access to their data though APIs and full data snapshots in various formats and with varying frequency. Providing a full data snapshot is also a requirement of the Principles of Open Scholarly Infrastructures which many of these organizations are committed to.

Data snapshots are therefore an opportunity to provide greater certainty about the future of datasources, reducing costs to providers, and enabling new applications. We argue that data snapshots should be a first class way of providing data, and for the provision of these to be accompanied by good documentation, including versioning, licensing and data provenance.

Doing this efficiently will require discussion of how best to provide snapshots, including file formats, regularity, documentation and more. The potential benefits will be achieved through making the provision of documented, reliable, and trustworthy data snapshots, optimised for usability and cost of production and transfer.

In this session, we will discuss amongst providers and users of open datasets how the design and delivery of data snapshots can be optimized, and help the development of shared infrastructure for opening up use of these datasets.

We will do this by surfacing real-world issues in working with these datasets and highlighting good examples of dataset provision and documentation across the ecosystem.

Advisor, research analyst and facilitator at Sesame Open Science
Executive Director, Barcelona Declaration on Open Research Information

This speaker also appears in: