2026-10-28 –, Naturalis 02
While the share of digital only objects in library holdings and scientific output is growing rapidly, the comparative short life span of both digital resources and the tools needed to access them make them vulnerable to damage and loss. Geopolitical developments have been adding to this vulnerability, putting political and financial pressure on institutions providing important public data bases. Digital preservation counters these inherent and external tendencies to dissolution by securely archiving resources, keeping their contents both intact and readable. Yet to retrieve resources from possibly dark archives when they are lost from public data bases we need findable metadata telling us where to look. This talk gives a short oversight on which kind of PID metadata could be especially useful for documenting digital preservation and aid recovery, and how this information can be or is already stored with different PID metadata setups.
Around May 2025, media reports picked up on the loss of knowledge from scientific data bases situated in the U.S. Resources were reported to be disappearing or being altered, apparently in avoidance of research topics deemed objectionable by the current administration. Political and financial pressure also calls the maintenance and independence of important data bases like PubMed into question. In response, grassroot initiatives and major institutions quickly joined in an international effort to duplicate and archive endangered data pools and secure them for scientific use.
Though geopolitical developments have been exacerbating the situation, the vulnerability of digital records is nothing new. The history of storing knowledge is also the history of the physical life span of the media involved becoming shorter, of the tools needed for access becoming more complex, and of modes of recording and reading becoming obsolete ever faster – compare engraved stone, parchment, paper, magnetic tape, and a Visicalc file on a floppy disk. Yet the share of digital only objects in library holdings and scientific output is growing rapidly. For all the advantages of digitalization, digitalized knowledge needs digital preservation, as a dedicated and ongoing process and effort.
But for preservation services to be of all the help they can be, we need to know which archive to go to when looking for a resource. Digital preservation can be provided by a different entity than the database the DOI points to and/or in dark archives, which are not searchable from external applications. Once scientific resources disappear from the internet, search indexes, etc., you need a stubbornly findable clue where to look for them instead, in order to access them and maybe instigate their re-publication. The obvious place for this kind of clue is in the PID metadata of the resources.
A2 of the FAIR principles sketches this use case: „Metadata are accessible, even when the data are no longer available“. While this aims at securing minimum information on resources which have indeed vanished, it also predestines FAIR metadata (and the PIDs included in them, according to principle F1) as the permanent signpost for resources not altogether gone, but gone from sight, or having been damaged or corrupted, to the places where they are securely archived. Better still to be lost and found than to be lost and persistently remembered.
But which sort of information should we ideally note down in the metadata to document the archiving and aid recovery, and how well fitted are different types of PIDs and their metadata schemata to accommodate this? Are there gaps in PID metadata design and usage we might want to fill? This talk will give you a short oversight and input on the topic, as an invitation to address it from your own specific PID point of view.
Information Manager for PID and Metadata Services at TIB Hannover