DARIAH Annual Event 2026
Rome, Italy. May 26–29, 2026
Conference Agenda
Overview and details of the sessions of this conference. Please select a date or location to show only sessions at that day or location. Please select a single session for detailed view (with abstracts and downloads if available).
Please note that all times are shown in the time zone of the conference. The current conference time is: 11th Sept 2026, 11:47:07am CEST
|
Daily Overview |
| Session | |
|
Poster and Demo Session Location: Foyer | |
| Presentation 5 | |
Making Folklore Collections Searchable and Reusable: A Tool for Transparent Tale Categorisation (Poster + Demo) University of Bologna, Italy Folklore research often starts from difficult material conditions: collections are scattered across institutions, described with inconsistent metadata, and frequently available only as scans. In multilingual settings, this fragmentation is amplified by mixed cataloguing traditions and uneven documentation. As a result, researchers and heritage professionals spend substantial time on basic tasks—finding relevant items, assessing text usability, and manually assigning categories—before they can ask interpretive questions. This poster and demo present a locally deployable tool that helps turn fragmented folklore material into a research-ready corpus while keeping the decision process transparent. The project is demonstrated on a corpus of Russian-language magic folktales preserved in an Estonian archive. Working across languages and institutional traditions, it provides a transnational, intercultural case study of how public digital humanities can be supported by pragmatic, resilient infrastructure. The workflow has three goals: make text available from scans in a way that can be inspected, consolidate metadata into a stable index suitable for citation and analysis, and enable automatic, verifiable type prediction for new tales using an existing, expert-assigned typology. First, the pipeline produces machine-readable text from scanned pages and generates a simple quality log that documents processing status and known issues for each item. Second, it builds a corpus index with stable identifiers for texts and volumes and a consolidated set of descriptive fields that researchers typically need (e.g., provenance signals and archive-facing descriptions). This index serves as the backbone for searching, sampling, and reporting, and supports stewardship through consistent identifiers and documented processing outcomes. Third, the corpus already contains reference tale-type labels (assigned according to the Aarne–Thompson–Uther, ATU, index). The tool leverages these labels as training data to build a model that can predict the most likely type(s) for an unseen tale. For each input text, it outputs a shortlist of three candidate types. For a curated subset, the system also provides short textual evidence (“quotes from the tale”) that supports each candidate, enabling rapid expert verification. The intention is not to re-classify the archive, but to demonstrate a reusable baseline for extending typological access to additional, external, or newly digitised tales. To support reuse and collaboration, the project exports corpus metadata in widely used linked-data formats (RDF/Turtle and JSON-LD) and provides example queries and basic validation rules to ensure that key fields are present and consistent. Overall, the contribution aims to lower the entry barrier to working with under-documented folklore materials and to create a trustworthy bridge between digitisation outcomes, archival description, and research workflows. State of the system: work in progress, master thesis project; the poster will be accompanied by an in-person software demonstration. project repo https://github.com/eugeniavd/magic_tagger | |
