DARIAH Annual Event 2026
Rome, Italy. May 26–29, 2026
Conference Agenda
Overview and details of the sessions of this conference. Please select a date or location to show only sessions at that day or location. Please select a single session for detailed view (with abstracts and downloads if available).
Please note that all times are shown in the time zone of the conference. The current conference time is: 11th Sept 2026, 11:31:58am CEST
|
Daily Overview |
| Session | |
|
Poster and Demo Session Location: Foyer | |
| Presentation 21 | |
A National Archive Database as an Infrastructure of Engagement Swedish National Archives While the digitization of cultural heritage has long promised to democratize access to history, a significant "analog barrier" persists. Although millions of documents are available as digital images, they remain machine-unreadable – effectively locking their content away from search and analysis. Consequently, vast historical narratives remain hidden within unstructured data, to a large extent accessible only to experts with specialized paleographic skills. This paper presents the ongoing transformation of the Swedish National Archive Database (NAD) into a dynamic e-infrastructure for research (2026–2030, funded by the Swedish Research Council). This infrastructure is designed to fulfill three primary objectives. First, it enhances interoperability and retrieval through streamlined API integration and standardized frameworks. By leveraging IIIF and OAI-PMH, the initiative expands access to research datasets, full-text, AI search and reference data, providing researchers with a more robust digital toolkit. Second, the platform enables the development of domain-specific AI models fine-tuned on archival holdings. Through extensive collaboration with READ COOP and the Transkribus community, the National Archives has developed the Swedish Lion model, trained on 50,000 manually verified pages. Unlike general large language models (LLMs) – which often suffer from "hallucinations" and lack historical nuance – our model possesses the contextual vocabulary necessary to navigate the dialectal and orthographic complexities of historical Swedish language. Future iterations will extend these capabilities to non-linear formats, such as maps and tables. Third, the infrastructure establishes a bidirectional data ecosystem, allowing researchers and the public to contribute transcriptions and metadata corrections. While NAD remains the authoritative system for core archival information, this platform creates a mechanism for harnessing external expertise. By integrating user-generated contributions with AI-driven processing, the National Archives can significantly enhance the quality and connectivity of its digital collections. Finally, the presentation explores the project’s commitment to ethical AI and transparency. Moving away from opaque "black box" models, NAD prioritizes scientific explainability through validation and our own open-source pipeline for historical text, HTRflow (http://github.com/AI-Riksarkivet/htrflow). In doing so, the National Archives retains control, guaranteeing that the "ground truth" of historical inquiry remains rooted in authentic archival sources. In sum, we argue that this new infrastructure represents a transformative shift: from digital archives as static repositories to a collaborative e-infrastructure that fosters co-creation and new research. | |
