DARIAH Annual Event 2026
Rome, Italy. May 26–29, 2026
Conference Agenda
Overview and details of the sessions of this conference. Please select a date or location to show only sessions at that day or location. Please select a single session for detailed view (with abstracts and downloads if available).
Please note that all times are shown in the time zone of the conference. The current conference time is: 11th Sept 2026, 11:46:59am CEST
|
Daily Overview |
| Session | |
|
Topic: From Speech to Symbols: AI Systems Transforming Cultural Knowledge Location: Aula Bisconti Session Chair: German Rigau, University of the Basque Country | |
| Presentation 4 | |
5:15pm - 5:30pm
FrWhisper – An Open-Source, Verbatim Automatic Speech Recognition System for Older French Speakers 1: Hasso Plattner Institute, Germany; 2: University Potsdam, Germany Despite their usefulness, AI-supported automatic speech recognition (ASR) systems such as Whisper (Radford et al., 2022) often do not meet crucial academic standards of the humanities. First, Whisper-based transcripts reflect an idealised language, adapted to the standard written language through additions (e.g., of grammatically mandatory words) and omissions (e.g., of repetitions, interjections, and discourse markers, Lea et al., 2023). This increases the need for ASR systems producing non-idealised transcripts. Such ASR systems are currently available only for English (Wagner et al., 2024), reinforcing a language bias in the digital humanities. Second, many ASR systems exhibit a pronounced age bias (Fukuda et al., 2023), reflected by less accurate transcriptions for older speakers. Idealised language transcription and age bias impose severe limitations for the analysis of the procedural nature of spoken languages in the humanities and cultural heritage preservation, frequently relying on data from older speakers. Against this background, this paper presents FrWhisper, an open-source ASR model for French that produces non-idealised transcripts and is specifically optimised for older speakers. FrWhisper is based on Whisper and was fine-tuned using two datasets from fieldwork corpora featuring spontaneous speech. Both datasets dominantly comprise speech from older speakers, from Orléans and its surroundings: The LangAge corpus (www.langage-corpora.org) consists of biographical interviews with mostly retired speakers (mean age = 80.1, SD = 9.3); the Enquête Sociolinguistique sur Orléans (eslo.huma-num.fr) is a large-scale interview corpus (mean age = 51.0, SD = 15.6). Empirical evaluation shows that FrWhisper substantially outperforms Whisper, reducing Word Error Rate from 98.98% to 84.64% on a held-out validation set. Qualitative analyses further demonstrate that FrWhisper more reliably preserves linguistically meaningful features of spoken French, including interjections (e.g., euh), ne-deletion (e.g., c’est pas vs. ce n’est pas), and repetitions. FrWhisper enables new forms of collaboration between academia, archives, libraries, and the public. It supports participatory research practices such as community-driven oral history projects, citizen-science annotation workflows, and public humanities initiatives that seek to make spoken cultural heritage accessible without erasing linguistic diversity. FrWhisper is openly available under the GNU General Public License v3.0, enabling reuse by researchers, memory institutions, and citizen-science initiatives. Its open documentation (Müller & Gerstenberg, 2025) further serves as a blueprint for developing verbatim ASR systems for other languages, aligning with DARIAH’s commitment to sustainable, ethical, and resilient research infrastructures. References Fukuda, M. et al. (2023). A new speech corpus of super-elderly Japanese for acoustic modeling. Computer Speech & Language, 77, 101424. https://doi.org/10.1016/j.csl.2022.101424 Lea, C. et al. (2023). From user perceptions to technical improvement: Enabling people who stutter to better use speech recognition. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems (pp. 1–16). New York, NY: ACM. Müller, Hanno & Gerstenberg, Annette. (2025): Transkription mit Fine-Tuned Whisper-Modellen. AI Service Centre. https://github.com/aihpi/pilotproject-FrWhisper Radford, A. et al. (2022). Robust speech recognition via large-scale weak supervision. In Proceedings of the 39th International Conference on Machine Learning. https://doi.org/10.48550/arXiv.2212.04356 Wagner, L., Thallinger, B., & Zusag, M. (2024). CrisperWhisper: Accurate timestamps on verbatim speech transcriptions. arXiv preprint arXiv:2408.16589. | |
