Cultural Memory in the Age of Data Infrastructures

When we talk about repositories for storing datasets produced in scientific research, we usually come across a few familiar names: Zenodo, Open Science Framework, and Figshare. They are important, robust, and, in the open science environment, virtually indispensable. However, (not only) for research in the humanities, they are more of a starting point than an end goal.

27 May 2026 Natálie Čornyjová

In the context of digital humanities, data are not understood purely as technical objects stored in a database, but as carriers of interpretation, context, and cultural memory. Datasets contain traces of manuscripts, textual fragments, annotations, maps, or linguistic layers, and this is precisely why a distinct ecosystem of repositories, archives, and infrastructural platforms has emerged in the humanities, one that operates differently and follows discipline-specific logics compared to repositories in the natural sciences.

From General-Purpose Repositories to Disciplinary Ecosystems

General-purpose repositories such as Zenodo offer a crucial advantage: stability. They assign DOIs, ensure long-term preservation, and provide relatively clear citation practices. They are therefore well suited for datasets, articles, software, and supplementary materials, in short, for everything research may produce. 

However, humanities data tend to be more complex. What, for instance, constitutes a dataset in a critical edition of a medieval text? Are they transcriptions? TEI-encoded XML files? Variant readings? Images? Metadata layers and editorial commentary? And what if interpretation is part of the data itself?

At this point, the universality of general repositories begins to fall short, and infrastructures such as CLARIN or DARIAH become increasingly important.

CLARIN 

The CLARIN infrastructure is an example of how a repository can simultaneously function as a research environment. It is not only about storing linguistic data, but about an entire ecosystem of corpora, lexical databases, natural language processing (NLP) tools, interoperability standards, authentication services, and methodological know-how.

Humanities data within CLARIN are not neutral objects, but results of annotation and interpretation. Each tag in a corpus is grounded in linguistic theory, and metadata thus become an epistemological layer in their own right rather than a mere “technical supplement.”

TEI and the Question of Interpretation

The fact that metadata become a separate layer for interpreting the world is particularly evident in the community around the Text Encoding Initiative (TEI). TEI is not a repository in the classical sense, but a standard for text representation. Nevertheless, it has given rise to an entire archival and publishing ecosystem. Encoding a text in TEI means deciding what counts as a paragraph, a correction, an authorial intervention, an uncertain reading, or where commentary begins and ends. This demonstrates that humanities data are rarely “raw.” They are curated. As already suggested, this is perhaps one of the key differences compared to the understanding of data in certain natural and technical disciplines.

There are many guides to getting started with TEI, but if you want to learn more, we recommend checking out, for example Introduction to Encoding Texts in TEI (Part 1) or A beginner’s guide to XML and TEI.

Example of a TEI/XML markup in the Visual Studio Code editor: the text is represented here as a structured layer of semantic tags. Source: https://www.pmoran.ie/posts/guide-to-xml/

Digital Libraries as Cultural Infrastructures

Alongside research repositories, there is also the world of digital libraries and cultural archives. Examples include Europeana which connects millions of digitised objects from European institutions, or Gallica the French digital library, which demonstrates how a national library can develop a long-term digital infrastructure. A specific category is represented by web archives, such as the Czech Webarchiv which focuses on preserving the national web space and digitally emerging culture.

Because such infrastructures effectively function as models of cultural memory, they raise questions such as: What gets digitised and what remains undigitised? Which metadata standards are used? Which languages and regions are represented? And what happens to materials once funding ends?

Digital humanities often operate within these invisible decisions.

The Most Interesting Cases Are Often Small, Specialised Databases

Perhaps the greatest value lies in small, discipline-specific projects emerging from concrete research needs. Such repositories or digital archives often experiment with different ways of modelling data relations.

ORBIS for instance, does not treat maps in a traditional sense but as a dynamic model of mobility in the Roman world, where space is understood as a network of temporal and logistical costs. Mapping Gothic France translates architecture into spatial analytics, enabling us to understand Gothic buildings as networks of relations between places, styles, and historical contexts. In another mode, TalkBank stores spoken language as interactional events in time rather than as simple textual transcripts. Similarly, Nomisma, is a database of coins that simultaneously models economic, political, and symbolic relations in the ancient world.

These and many other projects in the digital humanities demonstrate that repositories can become experimental environments in which new ways of translating cultural reality into structured, machine-readable forms are tested, along with the very definition of what counts as a knowable object.

Screenshot of the ORBIS database interface. Source: https://orbis.stanford.edu
Screenshot of the Mapping Gothic France database interface. Source: https://mcid.mcah.columbia.edu/mapping-gothic

The Fragility of Projects

We must also not overlook the darker side. Many of these projects are grant-funded: three years of funding, a team of doctoral researchers, a web application, a database, publications, and then? Silence. Domains expire, APIs stop working, documentation disappears, and datasets remain without maintenance. This results in the production of new digital ruins. For this reason, long-term infrastructures such as CLARIN, DARIAH or national repository networks are becoming increasingly important. Beyond data storage, they ensure cultural continuity.

As this article has shown, repositories do not only preserve data, but also ways of interpreting, connecting, and re-reading the past. And perhaps this is where their most significant role lies: they become maps of how contemporary culture understands and engages with its own heritage.


More articles

All articles

You are running an old browser version. We recommend updating your browser to its latest version.

More info