Cameron Neylon
Sessions
For data providers and data users there is increasing interest in using large scale snapshots of entire datasources. For data providers this can relieve the stress and cost of running APIs. For users it allows for new kinds of analysis, creating global benchmarks and identifying gaps. PIDs are at the centre of these analyses, allowing us to link large datasources together reliably. ORION-DBs is an effort to make data dumps more usable by exploiting the capability of Google BigQuery. Setting up this federated set of cloud data resources we have identified patterns in which datasets are easy to ingest, which are easy to use and how they can be combined. In this BoF session we focus on a facilitated discussion of what the optimum data snapshot looks like in format, structure, regularity, locations and how these can help (or hinder) the ability of users to fully exploit these datasets.
In short the storyline would focus on what challenges lie ahead for PIDs to keep knowledge connected. In a world where geopolitical turmoil puts strain on knowledge to be shared openly and neutral, PIDs could ba a backbone to trust. But in order to do so (and save the world ) the PID community would need to reflect, adapt and rethink the ways of working.