DARIAH Annual Event 2026
Rome, Italy. May 26–29, 2026
Conference Agenda
Overview and details of the sessions of this conference. Please select a date or location to show only sessions at that day or location. Please select a single session for detailed view (with abstracts and downloads if available).
Please note that all times are shown in the time zone of the conference. The current conference time is: 11th Sept 2026, 11:47:11am CEST
|
Daily Overview |
| Session | |
|
Poster and Demo Session Location: Foyer | |
| Presentation 16 | |
Benefits of Research Data Repositories: Showcasing the Parallel Bible Corpus and the TextGrid Repository 1: Georg-August-University Göttingen, SUB Göttingen; 2: Information Commissioner's Office This poster shows the benefits of integrating resources previously available in GitHub repositories into research data repositories. More specifically, we will show the different ways in which the quality of the data of the Multilingual Parallel Bible Corpus could be improved through its integration in the TextGrid Repository (TGR). The corpus contains biblical texts spanning over 100 languages, encoded originally in Corpus Encoding Standard (CES) as XML files (Christodouloupoulos & Steedman 2015). The project aimed to create a corpus aligned at the verse level for comparing methods in highly multilingual contexts, including those involving low-resource languages. In 2025, the resource has been integrated into the TGR, being part of Text+’s portfolio, the consortium for text- and language-based research data within the German National Research Data Infrastructure (NFDI). The TGR aims to integrate existing resources improving the quality of the data in line with the FAIR principles, and providing long-term archiving. As a TEI-specific repository, the TGR has been equipped with new features, such as project-specific options, the use of library classification and authority file systems (Calvo Tello et al. 2023), a Python library for accessing data (Hynek et al. 2024), and a new publication workflow (Veentjer et al. 2025). Other project with XML-TEI data can benefit similarly from these features when publishing in the TGR. During its integration, several aspects of the data were updated following the FAIR-principles (Wilkinson et al. 2016). This proposal highlights only a few. Considering the text, the original CES files have been encoded into TEI. Considering the metadata, some information has been enriched:
Thus, this corpus is one of those that adhere most closely to the FAIR principle. The Bible offers a unique chance for analysis across languages. To demonstrate the previously mentioned characteristics of the resource and the TGR, we are currently creating Jupyter Notebooks applying Named Entity Recognition algorithms in various languages. The poster will highlight different aspects: First, it will provide a quantitative overview of the corpus, presenting data on various measures such as documents, words, languages, works, and authors. Secondly, a visual representation of the connections across different levels will demonstrate the various types of resources to which the corpus is currently linked. Thirdly, the benefits of the corpus within the TGR will be displayed, with several QR-codes linking to different functionalities. Finally, the poster will give access and show the main results of the above-mentioned Jupyter Notebooks. Furthermore, this contribution serves as an invitation for other projects with TEI files to import their data into the TGR and enjoy similar advantages. | |
