DARIAH Annual Event 2026
Rome, Italy. May 26–29, 2026
Conference Agenda
Overview and details of the sessions of this conference. Please select a date or location to show only sessions at that day or location. Please select a single session for detailed view (with abstracts and downloads if available).
Please note that all times are shown in the time zone of the conference. The current conference time is: 11th Sept 2026, 11:47:12am CEST
|
Daily Overview |
| Session | |
|
Poster and Demo Session Location: Foyer | |
| Presentation 3 | |
Corpusense: A Lightweight Infrastructure for Producing Structured Data from Digitised Heritage Collections 1: LASTIG/ENSG/IGN, France; 2: Bibliothèque nationale de France, France; 3: Laboratoire de recherche de l'EPITA, Le Kremlin-Bicêtre, France; 4: CRH/EHESS, France; 5: Centre-Jean-Mabillon/ENC, Paris Mezanno is a program jointly led by the Bibliothèque nationale de France (BnF), IGN, EHESS, and EPITA, which aims to facilitate the production of structured data from digitized serial historical sources (such as directories, registers, and administrative records). In this demonstration, we will present Corpusense (https://mezanno.xyz/corpusense ), the web interface of the Mezanno project, designed to enable humanities and social sciences (HSS) teams and heritage institutions to manage, without heavy infrastructure, a complete workflow from images to structured data: importing sources via IIIF, building working collections, running automated processes (layout analysis/segmentation, OCR, and structuring), editing and correcting outputs within the interface, and exporting data in standard formats (CSV/JSON/Excel). Corpusense is a lightweight platform (deployed as static files, with projects stored locally) that lowers technical and financial barriers while fostering the autonomy of research communities. Its architecture implements a clear separation of responsibilities and expertise: heritage institutions ensure sustainable access to content (notably via IIIF); technical experts deploy remote services accessible through APIs (transcription, structuring, enrichment); and researchers and heritage practitioners retain control over source selection, modeling choices, and the critical validation of results. This approach is situated with respect to existing platforms--collaborative transcription environments and widely used transcription tools as well as open environments such as eScriptorium--while focusing specifically on producing reusable tabular structured data directly usable in downstream HSS workflows. The audience will follow a guided and reproducible workflow, illustrated using a complex heritage corpus. We will demonstrate: (1) the selection and import of images from an IIIF manifest; (2) the creation of a working corpus and its organization into collections; (3) the configuration of an extraction task (definition of fields, formats, constraints, and normalization rules); (4) the execution of automated processes and the display of results within the interface; (5) the review and editing of results by the user prior to export, supported by simple indicators of errors and cases requiring verification; and (6) the options for exporting and sharing the resulting data. We will also discuss two issues that are central to responsible, usable infrastructures for digital humanities. First, trust in automated results: how to support rigorous evaluation campaigns despite their cost and protocol complexity, how to define metrics that make sense for end users, and how to address alignment problems between predictions and reference data. Second, robustness and interoperability: the modular integration of evolving OCR/HTR components and the stabilization of APIs to support future extensions and richer outputs. | |
