2026-10-28 –, LUMC05
Many publishers turn to the ROR API to match text-based affiliations to persistent identifiers (PIDs). This tool typically yields a matching percentage of just 52% to 57%, if one restricts oneself to reliable outcomes. In this session, you will discover further possibilities for matching, such as Human Intelligence Tasks, Machine Learning, and Vector Search. This is based on concrete experiences from a project carried out at a medium-sized German science and medical publisher. An intriguing question to which we ourselves do not yet have an answer: will artificial intelligence ultimately remove the necessity for a human in the loop?
Matching researchers’ affiliations to a persistent identifier (PID) is an ongoing challenge for publishers. The text-based data that authors provide at submission, is often messy and ambiguous, making it difficult to standardise. To enhance data quality for analysis and reporting downstream, it is vital to match this raw affiliation data with accurate, machine-readable institutional information.
Many publishers turn to the ROR (Research Organization Registry) API for this purpose. However, relying solely on this tool typically yields a matching percentage of just 52% to 57%, if one restricts oneself to reliable outcomes. Whilst we greatly appreciate and make extensive use of the ROR API, its limitations are evident.
In this session, you will discover further possibilities for matching text-based affiliation to ROR IDs. This is based on concrete experiences from a project carried out at a medium-sized German science and medical publisher. We shall share our analysis of the dataset as the initial state, intermediate steps with considerations, lessons learnt, and, of course, the final state. Among other things, we shall demonstrate why we do not surpass 85% in this specific case.
As further possibilities for matching, we discuss:
• Human Intelligence Tasks (HITs): When automated software cannot find a definitive match, human experts manually review and match affiliation texts to suggested ROR IDs. This process also generates high-quality training data, which is essential for continually improving machine learning models for future automated matching.
• Machine Learning (ML): Employing techniques such as Machine Translation (MT) for language conversion and Natural Language Processing (NLP) to extract institution names from affiliation texts, in combination with ML algorithms learning from HITs, will continually improve matching rates over time.
• Vector Search: Moving beyond basic ROR API searches, we embed the ROR dataset as vectors and use advanced vector search applications to efficiently compare and match new affiliation texts. This approach builds on our extensive collection of previously matched affiliation texts, resulting in superior speed, stability, and accuracy.
An intriguing question to which we ourselves do not yet have an answer: will artificial intelligence ultimately remove the necessity for a human in the loop? This question could provide direction for a subsequent Q&A.
Strategic technology services partner in STM Publishing