Chemical Identity Lost in Regulation: A Study of Semantic Interoperability in European Chemical Substance Data
Cite thisCitation
Van Haute, G., Goedertier, S., Fannes, P., & Van de Wynckel, M. (2026). Chemical Identity Lost in Regulation: A Study of Semantic Interoperability in European Chemical Substance Data. Posters, Demos, Blue Sky, and Tutorials at SEMANTiCS 2026.
@inproceedings{vanhaute2026,
author = {Van Haute, Geert and Goedertier, Stijn and Fannes, Pieter and Van de Wynckel, Maxim},
booktitle = {Posters, {Demos}, {Blue} {Sky}, and {Tutorials} at {SEMANTiCS} 2026},
year = {2026},
organization = {CEUR-WS.org},
title = {Chemical {Identity} {Lost} in {Regulation}: A {Study} of {Semantic} {Interoperability} in {European} {Chemical} {Substance} {Data}},
}
Authors
Abstract
European regulatory datasets represent chemical substances heterogeneously: CAS numbers, EC numbers, and textual names that may refer to single molecules, mixtures, substance groups, or analytical parameters. We analyse 18 ECHA-based datasets (41,813 records) using a seven-tier linkability taxonomy. Only 62.5% of entries link to a defined molecular structure; 37.5% are non-structure-based. CAS numbers prove unreliable as unique identifiers. Four sentence-embedding models achieve ChemOnt Hit@1 ≤ 5.75%; Claude Sonnet 4.6 reaches 39.2% yet still misclassifies the majority. Structure-defined substances are integrated into a knowledge graph with ChemOnt classification and ChEBI biological roles (~48,400 cross-domain links). For non-structure entries, the LLM abstains from direct classification far more often than for structure-defined substances, appropriately reflecting their lack of a defined structure; a separate LLM-based scope-mapping assessment yields 304 validated SKOS triples linking regulatory group entries to ChemOnt nodes. Semantic interoperability requires structural changes at the level of regulatory data modelling and legislation.