A full-day Barstuhl workshop · ISWC 2026

Computational cHAllenges fRom hiGhly divErse Data

CHARGED brings together computer scientists, digital humanities researchers, and cultural heritage professionals to tackle the integration of heterogeneous, asynchronous data streams — from historical archives to contemporary sensor feeds. This Dagstuhl-like workshop will take place in conjunction with the 25th International Semantic Web Conference (ISWC 2026) in Bari, Italy on 25 or 26 October.

Introduction

Data is the key to understanding and solving challenges such as the loss of biodiversity, predicting recurring weather phenomena, or the fragmentation and loss of context of our cultural heritage. For this, vast amounts of data from many different sources need to be interpreted. Due to advances in computer power and the increasing availability of data spanning longer periods of time, historical data is becoming more important to include.

Unfortunately, many AI tools are not well-suited to deal with this type of data, and integrating contemporary and historical data streams is a complex problem. In this workshop, we bring together computer scientists, digital humanities researchers, and cultural heritage professionals to create a whitepaper detailing the challenges for AI-assisted integration of heterogeneous and asynchronous data streams from contemporary and historical sources.

Motivation and Topics

Libraries, archives, and museums (LAMs) have digitized — and are still digitizing — vast collections of scientific materials, ranging from botanical records to early climate observations, opening new opportunities for a more nuanced understanding of environmental change. Europe's natural history museums alone house data on more than 1.5 billion plants, animals, and minerals, collected across the globe over the last four hundred years.

The computational study of such materials — created by Carl Linnaeus and other naturalists — can help enrich present-day understanding of scientific knowledge as it is created across environmental sciences. With advances in AI methods such as language and computer vision technologies, historians and other researchers can now analyze large, heterogeneous datasets to uncover long-term patterns and changes in environmental and other forms of knowledge production. However, these domains pose challenges to AI methods since the data is generally more prone to noise, and to organisational and interpretation challenges.

At the same time, new data streams are becoming available, such as drone footage and sensor and tag data for monitoring biodiversity. Integrating historical and contemporary data, across different modalities and fidelities, is a major challenge for knowledge representation — and one that must be solved to build world models that support resilient responses to natural hazard problems.

The workshop is organised around two core challenges — making historical and contemporary data speak to each other, and doing so honestly. We identify three intersecting topic areas:

Fidelity mismatch

  • Integrating datasets created in different time periods or from different perspectives, with different levels of granularity and standardisation — stemming from divergent tools, instruments, and archival forms or nomenclature used to produce, capture, and document data over time (e.g. differing sensor calibrations, precision, cataloguing software, or metadata schemas).
  • Multimodality and heterogeneity: historical data can come as tables, text, drawings, and maps; contemporary data can come in these modalities or additional ones such as video and sensor data. At times these modalities are mixed within the same data stream, and untangling where information complements — or contradicts — is challenging.

Spatio-temporal context

  • Tracking and modelling long-term concept drift, from both natural language and knowledge representation perspectives.
  • Contemporary data is more standardised and calibrated, and recording tools are documented. Interpreting historical data for which the creators and domain experts have perished makes it often not directly usable, introducing a level of uncertainty about interpretation that must be accounted for in models and analyses.
  • Data is further shaped by the positionality of those who produced it — institutional priorities, disciplinary conventions, and colonial or national framings — introducing interpretive bias that needs to be made explicit rather than flattened during integration.

Positionality & bias

  • Making explicit the institutional, disciplinary, and (post)colonial framings embedded in both historical and contemporary datasets.
  • Developing practices for surfacing — rather than flattening — interpretive bias when datasets are combined into shared models.
  • Data ethics considerations for AI-assisted integration across asynchronous sources.

Use Cases

We bring three use cases to the workshop, chosen for their strong connection to current ecological and societal debates, the availability of datasets, and the organisers' expertise and networks. In the run-up to the workshop we will liaise with prospective participants to introduce these use cases and provide access to datasets and core readings, so the group can get started right away.

Natural history

Botanical and zoological collections spanning centuries — from Linnaean records to present-day biodiversity monitoring — and the challenge of reconciling historical taxonomies with modern data.

Meteorological & maritime history

Long historical weather and marine records set against contemporary climate data, requiring models that reason across centuries of instrumentation change.

Digital cultural heritage

Large-scale digitised archival and museum collections, and the knowledge representation challenges of connecting them to present-day interpretation and use.

Chairs

Photo of Marieke van Erp

Marieke van Erp

KNAW Humanities Cluster, Netherlands

Homepage

marieke.van.erp@dh.huc.knaw.nl
Photo of Catherine Faron

Catherine Faron

Université Côte d'Azur, France

Homepage

catherine.faron@univ-cotedazur.fr
Photo of Andreas Weber

Andreas Weber

University of Twente, Netherlands

Homepage

a.weber@utwente.nl
Photo of Célian Ringwald

Célian Ringwald

University of Bologna, Italy

Homepage

celian.ringwald@unibo.it

Expression of Interest

This workshop seeks bring together a diverse group of researchers and practitioners working on the integration of historical and contemporary data for environmental and climate research. We specifically aim to involve semantic web researchers, digital humanities researchers with expertise in historical data, and cultural heritage professionals responsible for large-scale digital collections.

As this is the first time we are organising this workshop, we'd like to get to know our potential participants a bit. Please fill out the expression of interest form to help us tailor the workshop.