DARIAH Annual Event 2026
Rome, Italy. May 26–29, 2026
Conference Agenda
Overview and details of the sessions of this conference. Please select a date or location to show only sessions at that day or location. Please select a single session for detailed view (with abstracts and downloads if available).
Please note that all times are shown in the time zone of the conference. The current conference time is: 11th Sept 2026, 11:47:07am CEST
|
Daily Overview |
| Session | |
|
Topic: Applied AI and Reproducible Workflows: Sustainable Infrastructures for Public Knowledge Location: Aula Bisconti Session Chair: Tugce Karatas, University of Luxembourg | |
| Presentation 3 | |
12:00pm - 12:15pm
Open, On-Premise AI for Public Knowledge in German Administrations Hasso Plattner Institute, Germany Recent advances in language technologies have renewed interest in how public knowledge can be made more accessible, navigable, and understandable. Yet in public administrations, the deployment of AI systems raises specific infrastructural, ethical, and political questions that experience similarity with those in academic contexts. This practice-based, reflective infrastructure paper presents three AI-based knowledge infrastructures developed with German public administrations and reflects on their implications for public engagement, digital sovereignty, and long-term stewardship of public knowledge. The contribution is authored institutionally by the AI Service Centre Berlin-Brandenburg and it is positioned from the perspective of knowledge infrastructure. The paper discusses three concrete deployments that are architecturally distinct. First, a retrieval-augmented generation (RAG) system for the German Bundestag, enabling complex semantic queries over the archive of the Bundestag’s Scientific Service, supporting in-depth exploration of scientific studies on politically relevant research questions. The system is based on a modular RAG backbone that can be instantiated for different authoritative document collections and extended with heterogeneous sources. Second, a translation system that renders administrative texts, which are difficult to process, into Leichte Sprache (easy language) and is used by municipal administrations in Brandenburg. Third, a protocolling system that transcribes committee meetings and drafts minutes from audio recordings. The system, which is based on automatic speech recognition and a fine-tuned large language model, is deployed in the Landtag Brandenburg and Brandenburg municipalities. Across these heterogeneous applications, a set of shared infrastructural design principles emerges. All systems are fully open source, optimized for local, on-premise deployment on low-power GPUs, and deliberately avoid dependencies on external cloud APIs. They are designed to comply with European privacy and data protection regulations and to process sensitive political and administrative data. Local deployment is motivated not primarily by performance, but by digital sovereignty, cost and accessibility, operational robustness within administrations, and the ability to treat AI systems as long-term public infrastructure rather than transient services. Conceptually, the paper frames these systems as knowledge infrastructures rather than decision-making systems. They structure access, retrieval, translation, and summarization of public knowledge without automating political judgment or administrative discretion. This distinction is central to democratic accountability and trust: the systems aim to improve access, comprehensibility, and navigability of politically and socially relevant information, thereby supporting transparency and participation, while leaving normative decisions firmly with human actors. The paper explicitly addresses safeguards that are treated as preconditions for trust in public-sector AI, including mechanisms to mitigate hallucination, systematic traceability of sources in RAG outputs, and human-in-the-loop verification workflows. It reflects on the opportunities and limitations of local language models, the infrastructural requirements of privacy-compliant AI, and the societal implications of embedding such systems in everyday administrative practice. Finally, it highlights the reusability of the three systems for external researchers, both as a basis for domain-specific knowledge infrastructures and as a research instrument for exploratory, qualitative, and mixed-methods work on large document collections. In doing so, the paper aligns with DARIAH’s mission to build ethical, resilient infrastructures of engagement between academia, public institutions, and society. | |
