DARIAH Annual Event 2026
Rome, Italy. May 26–29, 2026
Conference Agenda
Overview and details of the sessions of this conference. Please select a date or location to show only sessions at that day or location. Please select a single session for detailed view (with abstracts and downloads if available).
Please note that all times are shown in the time zone of the conference. The current conference time is: 11th Sept 2026, 11:47:12am CEST
|
Daily Overview |
| Session | |
|
Poster and Demo Session Location: Foyer | |
| Presentation 15 | |
Beyond Archives: Designing a Segmentation and ATR Workflow for Mid-Twentieth-Century Typescripts in a Contested-Memory Case Study University of Naples "L'Orientale", Italy This poster presents results from Beyond Archives, focusing on layout segmentation and ATR for mid-twentieth-century Italian typescripts in a contested-memory case study. Beyond Archives bridges archival science and digital humanities to examine how interfaces shape visibility, mediation, and reuse in contexts marked by divergent institutional and lived memories (Hedstrom 2002; Gilliland 2017). Using the post-war Julian-Dalmatian exodus to Naples as its framework, the project is developing a documented workflow for comparable collections, attentive to memory institutions and affected communities (McKemmish and Piggott 2013). FAIR-oriented, standards-based methods record technical decisions, so models, error profiles, and processing choices remain inspectable and reusable (Tomasi and Buzzetti 2012; Michetti 2023). The case study examines a polarized memorial landscape where administrative documents, testimonies, and memories intersect (Lazzarich 2021). The infrastructure must keep these divergences readable to support research, stewardship, and public dialogue (Hedstrom 2010). The project works across two corpora: Prefecture of Naples records (1946-1958), administrative typescripts marked by stamps and custodial handling; and semi-structured interviews with former refugees, collected under ethical protocols and transcribed with speech-to-text tools. Keeping both streams in view supports exploration and dialogue across institutions and communities (Tomasi 2022). The workflow follows three stages: digitization, formalization, and publication. Digitization produces high-quality images with IIIF references. Publication is being developed in a TEI- and IIIF-based environment, building on a TEI baseline and preserving contextual layers and provenance traces. This poster focuses on formalization for the written corpus, where segmentation and ATR produce PAGE/ALTO outputs integrated into TEI-XML, with transformations documented for scrutiny (Chiffoleau 2025). Segmentation is performed in eScriptorium with the Kraken engine (Kiessling 2019; Kiessling et al. 2019). A SegmOnto-aligned vocabulary distinguishes text from paratextual zones, including stamps, signatures, and notes (Gabay et al. 2021). Ground-truth is expanded iteratively to cover heterogeneous layouts, treating these zones as evidence of custodial intervention. Longitudinal comparison across models shows stronger line-level performance, especially for body and head, while block-level over-segmentation remains the main limit; stamps are the weakest class, and signatures remain less stable than the main text zones. ATR fine-tuning starts from CATMuS-Print (Gabay and Clérice 2024) and proceeds in staged runs on Prefecture pages. Metrics are evaluated in fitness-for-purpose terms, taking into account small ground-truth samples (Bubula et al., 2025). The current reference checkpoint is ATRModel-5, achieving 96.81% character accuracy and 87.83% word accuracy in its consolidation run, and outperforming the later model tested on the extended corpus. These results support TEI structuring and entity-oriented access, while keeping residual noise visible for controlled reuse. For reuse and accountability, Beyond Archives treats ground-truth and paradata as research outputs. Its purpose is not preservation alone, but to make complexity visible and open to critical negotiation across institutional records and lived testimony. Once benchmarks stabilize, trained segmentation and ATR models, with curated training artifacts, will be deposited on Zenodo. The release serves a dual aim: sharing trained models for comparable corpora and providing a documented workflow and three-stage pipeline that others can adapt to similar contested-memory collections while keeping their tensions legible over time. | |
