Open Repositories 2026
Online | 8 - 11 June 2026
Conference Agenda
Overview and details of the sessions of this conference. Please select a date or location to show only sessions at that day or location. Please select a single session for detailed view (with abstracts and downloads if available).
Please note that all times are shown in the time zone of the conference. The current conference time is: 25th Aug 2026, 02:20:07pm UTC
|
Daily Overview |
| Date: Wednesday, 10/June/2026 | |
| 11:00 - 11:50 | Developer Track: DSpace 2 (Automated Metadata and Full Text Population) Location: Parallel Session 2 Session Chair A: Iryna Kuchma, Electronic Information for Libraries (EIFL) Session Chair B: Erin Jerome, University of Massachusetts Amherst Recording available at https://doi.org/10.5281/zenodo.20778607 |
|
|
Fill The Gap – Automated retrieval of full text from emerging open APIs 1: University of Galway, Ireland; 2: Atmire Populating an open repository with high-quality, consistent metadata is a substantial task. The challenge becomes even harder when records also need to be enriched with the corresponding full text at scale, particularly when authors are not involved in deposit workflows. In 2025, the University of Galway and Atmire developed and deployed a DOI-driven workflow to enrich metadata-only repository records with open full text links. The tool queries multiple open services using the DOI, selects the most credible full text candidate, and records both provenance and outcomes to support review and reporting. In production, this approach identified and attached thousands of full text PDFs with minimal manual intervention, while surfacing cases that require follow-up due to redirects, inconsistent landing pages, or unclear licensing signals. The implementation is designed to be extensible, with additional sources and local policy rules added as needed. The session will demonstrate the Google Apps Script and Google Sheets version, describe key design trade-offs (accuracy, coverage, validation, and rate limiting), and share an approach that other repository teams can adapt to their own infrastructure. Currently supported sources include OpenAIRE, Unpaywall, CORE, and OpenAlex. Reducing Barriers: Automating Metadata Extraction in Submission Forms for DSpace Repositories KEEP Solutions, Portugal As digital repositories evolve at the intersection of people, practice, and emerging technologies, the burden of manual metadata entry remains a significant barrier to the timely dissemination of open research. This paper presents a novel integration for the DSpace platform designed to streamline the submission process through automated metadata extraction. The proposed functionality leverages an external API powered by Artificial Intelligence (AI) to analyze uploaded documents in real-time. By identifying and mapping key bibliographic data directly from the file content, the system automatically populates submission forms, reducing human error and cognitive load for depositors. Central to this development are two critical considerations: interoperability and privacy. The architecture utilizes a flexible API framework that allows the repository to request services from various external providers, ensuring the system remains adaptable to future technological shifts. Furthermore, the integration is built with a "privacy-by-design" approach, ensuring that sensitive file data is handled securely during the AI analysis phase. By automating the "practice" of data entry, this feature moves us closer to an "Open to All" ecosystem where researchers can focus on dissemination rather than administration, ultimately fostering a more efficient and inclusive repository environment. |
| 11:00 - 11:50 | Presentations: Presenting Metadata Differently Location: Parallel Session 1 Session Chair A: Ianthe Sutherland, University of Edinburgh Session Chair B: Richard William Fyson, CoSector, University of London Recording available at https://doi.org/10.5281/zenodo.20778609 |
|
|
Beyond FAIR: Designing Cognitive Accessibility in Data-Rich Repositories Cottage Labs, United Kingdom FAIR principles make data technically accessible but do not guarantee cognitively accessible experiences, especially in data-rich repositories. Cognitive accessibility is often overlooked, yet interfaces with high cognitive load affect all users, not just those with diagnosed disabilities. Neurodivergent researchers, likely overrepresented in academia, face additional challenges due to differences in working memory, executive function, and processing speed, making repository UX a critical equity issue. This presentation addresses three common cognitive pain points: scanning large volumes of metadata to find relevant information; switching between broad and highly specific searches while maintaining context; and navigating repository structures that do not reflect the varied priorities of different user groups. Drawing on cognitive load theory and W3C COGA guidelines, the presentation demonstrates practical UX patterns that reduce extraneous cognitive load: grouping related metadata into clear, scannable sections; eliminating redundant information; and providing persistent mechanisms to carry context across search and discovery modes. Attendees will learn how to identify cognitive anti-patterns in their own repositories and apply these design solutions making FAIR-compliant metadata more usable, inclusive, and equitable for a wide range of researchers. Public Health Data in Context: Publishing the RKI research graph 1: Robert Koch Institute, Germany; 2: Cottage Labs, United Kingdom MEx (Metadata Exchange) is a digital platform that maps metadata on research activities and research data at the Robert Koch Institute (RKI) in a transparent way, aiming to facilitate the reuse of this data. MEx is one of many measures being taken to implement the German government's open data strategy. The core data is managed in the MEx management platform, and this is maintained, kept up-to-date, and evolves as the requirements of the research environment change. The objective of the repository project is to ensure that as that data evolves and is added to, that a stable view can be provided to researchers which maintains the version history of changes, and to enable researchers to discover and view research important to them. This presentation will introduce MEx as a project and data model/vocabulary, and talk about the motivation for publishing the dataset in a repository. The platform chosen for this is InvenioRDM, and we will talk about the reasons for that, and then look at the technical changes that were required to support the model. |
| 12:00 - 12:50 | Panel: Connecting Research Assets: RAiD (Research Activity Identifier) and Repositories Location: Parallel Session 1 Session Chair A: Isabel Bernal, CSIC (Spanish National Research Council) Session Chair B: Nick Sheppard, University of Leeds Recording available at https://doi.org/10.5281/zenodo.20789045 |
|
|
Connecting Research Assets: RAiD (Research Activity Identifier) and Repositories 1: Lyrasis, United States of America; 2: Australian Research Data Commons, Australia; 3: University of Texas at Austin, United States of America; 4: Louisiana State University, United States of America Many repositories already use persistent identifiers (PIDs) such as ORCID iDs for individuals and DOIs for research outputs to better understand research activity and to make research information more Findable, Accessible, Interoperable, and Reusable (FAIR). However, until recently, there has been no standardized, interoperable way to connect these elements from the perspective of an overarching research project. RAiD (Research Activity Identifier), an international standard and system for identifying research projects (ISO 23527:2022), was established in 2022 by the Australian Research Data Commons (ARDC) to address this gap. RAiD can be used alongside existing PIDs to more effectively identify, track, and share information about research projects and their associated outputs. RAiD is already established in Australia, where research organizations have been investigating the use of RAiD in their repositories to meet funder requirements. As part of the United States (US) RAiD Pilot, multiple universities are testing the integration of RAiD into their repository workflows. This session aims to provide an introduction to RAiD and share initial findings from recent RAiD adoption efforts to offer recommendations and insights for repositories interested in adopting RAiD. |
| 12:00 - 12:50 | Presentations: Openness & AI Location: Parallel Session 2 Session Chair A: Leigh Stork, University of Strathclyde Session Chair B: Adrian Ho, American University of Sharjah Recording available at https://doi.org/10.5281/zenodo.20778611 |
|
|
When Openness Meets a Breaking Point: Perspectives on capacity, responsibility and stewardship under the threat of AI-driven harvesting. 1: Fedora; 2: Metropolitan New York Library Council The growing prevalence of artificial intelligence has renewed attention on the role of data in training large language models (LLMs). For decades, digital libraries and repositories have focused on providing well-structured, searchable, and openly accessible information to the public. As a result, these systems have become major targets for large-scale AI data harvesting. The volume and intensity of automated access now place significant strain on technical infrastructure and on the people who maintain it, often exceeding the capacity intended to serve human users. In response, some institutions have limited access or taken systems offline, raising challenges to long-standing commitments to openness and public service. This panel addresses the operational, ethical, and strategic questions emerging from this reality. Drawing on the work of a cross-institutional working group, the session brings together diverse perspectives from roles involved in repository stewardship. Panelists will discuss how AI-driven harvesting affects daily operations, planning, and decision-making, and how responsibilities and constraints vary across roles, institutions, and legal contexts. By creating space for cross-role dialogue, the panel aims to advance discussion around mitigation, responsibility, and sustaining public mandates in an evolving, AI-driven internet. Repository Practices for the Dark Arts: When Openness Conflicts with Safety Queen's University Belfast, United Kingdom Open science principles increasingly demand that research data be findable, accessible, interoperable, and reusable (FAIR), yet cybersecurity research, particularly malware analysis and threat intelligence, exists in a paradox. The artifacts that underpin reproducible research are the very materials that could enable harm if openly shared. This proposal presents an initial open question at the intersection of open science and cybersecurity research. The main aim is to promote dialogue within the community on developing repository practices, ethical frameworks, and technical architectures that balance legitimate research needs against security risks, asking: how can repositories be "open to all" when the data itself can be weaponised? |
| 13:00 - 13:50 | Developer Track: DSpace 3 (Performance and AI Enhancement) Location: Parallel Session 2 Session Chair A: Iryna Kuchma, Electronic Information for Libraries (EIFL) Session Chair B: Ianthe Sutherland, University of Edinburgh Recording available at https://doi.org/10.5281/zenodo.20778615 |
|
|
DSpace Reimagined: AI-Powered Search and Accessibility PCG Academia, Poland This presentation showcases a next-generation, AI-powered enhancement layer for DSpace repositories, focused on two core areas: intelligent search and improved accessibility. Moving beyond traditional keyword-based discovery, the solution introduces AI-driven retrieval based on Retrieval-Augmented Generation (RAG), combining structured metadata, full-text content, and vector search to deliver accurate, context-aware answers grounded strictly in repository holdings. This approach improves relevance, supports multilingual and cross-language discovery, and scales to very large collections without compromising trustworthiness. In parallel, the presentation demonstrates AI-powered PDF processing designed to improve accessibility and reuse of repository content. Advanced OCR and document conversion pipelines transform complex or scanned PDFs into fully searchable, screen-reader-friendly, WCAG-aligned formats, significantly broadening access for users with disabilities and enabling downstream AI processing. Through short live demonstrations, architectural overviews, and real-world use cases, this presentation illustrates how DSpace repositories can evolve from passive storage systems into active, inclusive research discovery platforms – supporting FAIR principles, lowering barriers for new users, and responding pragmatically to the opportunities of emerging AI technologies. Challenges and solutions for reliable and performant DSpace repositories 4Science, Italy This presentation aims to share the strategies and solutions adopted at 4Science to provide a reliable, high-performance hosting service for DSpace repositories. In the cloud era, repository platforms must support the adoption of cloud-native paradigms to deliver reliable, performant, and cost-effective services. In 2025, 4Science transitioned its hosting infrastructure from a traditional VM-based architecture to a modern, containerized deployment powered by Kubernetes and AWS cloud-native services. This transition has required changes and fine-tuning to the DSpace application codebase, a review of the development life cycle, and the adoption of new operational tools. We will explain the reasons behind the changes and the benefits obtained. Based on the operational data, we have identified several areas of improvements across the different application layers, from the frontend to the backend. Lack of support for horizontal scalability has been addressed to provide HA and consistent performance under heavy load. Many improvements have already been contributed to the DSpace codebase directly or via the ongoing merger of DSpace-CRIS; others will be discussed with the community and offered for inclusion in future versions. This approach keeps DSpace service sustainable, stopping institutions from over-provisioning of costly resources and therefore reducing the environmental impact. |
| 13:00 - 13:50 | Presentations: Cultural Heritage Location: Parallel Session 1 Session Chair A: Nick Sheppard, University of Leeds Session Chair B: William J. Nixon, Research Libraries UK (RLUK) Recording available at https://doi.org/10.5281/zenodo.20778617 |
|
|
Historica: Exploring Cultural Heritage Through Space and Time with DSpace-GLAM 1: 4Science, Italy; 2: ALMA MATER STUDIORUM - Università di Bologna, Italy Digital Cultural Heritage management has evolved from passive content consumption to a demand for sophisticated analysis, exploration tools, and narratives. This paper introduces Historica, the digital library of the University of Bologna. Historica is built on top of DSpace GLAM, the Digital Library Management System based on DSpace, developed by 4Science. DSpace-GLAM implements a complex data model aimed towards a deep interrelation among digital objects and contextual information. While traditional digital libraries often manage digital objects as isolated files, Historica leverages DSpace-GLAM to structure relationships among digital objects and entities such as persons, places, events, etc. This architecture supports advanced features, enabling users to explore complex historical scenarios rather than isolated items, and to visualise digital objects through timelines and maps. Moreover, the platform integrates IIIF based services for high quality image analysis, delivery, annotation and storytelling and aligns with national and international standards for interoperability and preservation. The paper illustrates the dialogue between DSpace-GLAM and the needs of this academic digital library to support research, teaching, and public engagement. By combining digital library services with tools for analysis and exploration, Historica became a laboratory for new forms of historical inquiry and narrative, turning digital collections into navigable cultural landscapes. Open to All? Repository Workflows for Ethical Access, FAIRness, and Cultural Heritage Data University of Cape Town, South Africa Open repositories are critical infrastructures for sustaining open knowledge exchange, advancing FAIR principles, and enabling long-term preservation. At the same time, they increasingly operate under pressure from ethical obligations, community authority, and emerging technologies such as automation and AI-driven reuse. These pressures raise a central question for repositories today: what does it mean to be “open to all” in practice? This paper examines how repository workflows, infrastructure choices, and staff practices operationalise FAIR principles while negotiating the limits of openness when stewarding sensitive cultural and linguistic heritage data. Drawing on two contrasting case studies from the University of Cape Town Libraries in South Africa, it demonstrates how openness is produced through socio-technical decision-making rather than assumed as a default. By comparing these repositories, the paper highlights how platform affordances, governance models, and staff expertise mediate FAIRness, machine reuse, and ethical stewardship. It argues that integrating CARE principles alongside FAIR is essential for building repositories that are sustainable, accountable, and genuinely open in an AI-intensive research landscape. |
| 14:00 - 14:50 | Panel: Repository Responses on the Frontline of the Artificial Intelligence Revolution Location: Parallel Session 1 Session Chair A: Richard Jones, Cottage Labs Session Chair B: Leigh Stork, University of Strathclyde Recording available at https://doi.org/10.5281/zenodo.20789106 |
|
|
Repository Responses on the Frontline of the Artificial Intelligence Revolution 1: The Ohio State University, United States of America; 2: CoSector, University of London, United Kingdom; 3: arXiv, United States; 4: University of Edinburgh, United Kingdom Artificial Intelligence (AI) is rapidly transforming the scholarly communications landscape and reshaping how repositories are defining “Open to All”. Faced with challenges to their core missions, repositories are re-examining their fundamental commitment to “open”. Grappling with affronts to research integrity, eroding trust in the ecosystem, and evolving questions of authorship and ethics, repositories must understand their role in this new frontier and how to muster their responses to it. The moderated panel highlights the perspectives of three repository stewards tasked with managing this whirlwind of change. The first “open repository”, arXiv, is mounting new defenses and new offensive maneuvers as they deal with a ballooning onslaught of AI-generated content, authorship rings, and citation scams. The University of Edinburgh Library manages in-house Archipelago and DSpace repositories and hosts DSpace repositories across Scotland. They are proactively addressing AI to maintain the uptime of their repositories and are currently working on new licensing policies. CoSector, which develops for and hosts Samvera and EPrints repositories, is navigating the myriad of client AI perspectives in its support of multiple repositories. The audience is invited to share their perspectives on standing firm on the commitment to openness while combating the forces that threaten that ideal. |
| 14:00 - 14:50 | Presentations: National Repositories Location: Parallel Session 2 Session Chair A: Emily Bongiovanni, Carnegie Mellon University Session Chair B: Nick Sheppard, University of Leeds Recording available at https://doi.org/10.5281/zenodo.20778619 |
|
|
Developing a national aggregator of open access repositories in Algeria: project proposal University of Tamanghasset, Algeria In Algeria, the digital repository ecosystem has diversified significantly, with a current total of 71 institutional repositories. It is beneficial for the research community in Algeria to discuss the roles and connections of various repositories and to explore opportunities for improving national coordination. Despite the widespread adoption of repositories in Algeria, many encounter challenges like low visibility within the research community, outdated software platforms, and insufficient staffing. The importance of a national aggregator project is obvious in its ability to showcase the scientific and academic contributions of Algerian universities on national and international stages through a centralized national portal, which promotes access to open digital resources and improves the interoperability of Algerian digital repositories. The study aims to suggest a project to establish a national aggregator for open access repositories to enhance the discoverability, accessibility, and visibility of Algerian scholarly output. Reducing regional asymmetries in Brazil through digital repositories: the RBRD experience 1: Brazilian Institute of Information in Science and Technology (Ibict), Brazil; 2: State University of Santa Catarina (UDESC); Brazilian Institute of Information in Science and Technology (Ibict), Brazil This presentation challenges the premise that technical openness alone guarantees universal access, arguing that meaningful openness requires active policies to reduce inequalities. It analyzes the case of the Brazilian Network of Digital Repositories (RBRD), coordinated by Ibict, which adopts a decentralized governance model structured around five regional sub-networks to address historical asymmetries between Brazil’s South–Southeast axis and the North, Northeast, and Central-West regions. Drawing on the concept of “Informational Justice,” the analysis demonstrates how Brazilian repositories operate not only as digital platforms but as “citizenship infrastructures,” promoting informational sovereignty and returning publicly funded knowledge to society under an ethic of care. |
| 15:00 - 15:50 | Developer Track: DSpace 4 and Hyku (Upgrades and Customizations) Location: Parallel Session 2 Session Chair A: Paul Walk, Antleaf Ltd. Session Chair B: William J. Nixon, Research Libraries UK (RLUK) Recording available at https://doi.org/10.5281/zenodo.20778621 |
|
|
Why harvest your own DSpace? Benefits of separating repository and search University of Jyväskylä, Finland JYX is the institutional repository of University of Jyväskylä (JYU). The repository has been in use since 2008 containing over 90000 items of various types including publications, theses, datasets, historical maps, and audiovisual material. JYX utilizes DSpace software with customizations and REST-based integrations such as retrieving self-archived materials from CRIS and other publication workflows using external tools. In addition, JYX provides metadata via OAI-PMH for other services such OpenAIRE and Finna2, the national search service for archives, libraries, and museums. Because of DSpace-related customizations, upgrades have turned out to be complex efforts, especially preserving UI-based customizations such as special formatting for restricted items. When the support for DSpace 6 ended in 2023, we faced a choice: to work within the native DSpace UI or to build something more flexible. We chose to decouple repository and user interface. This presentation explores the technical and functional benefits of using an external search engine (VuFind) to "harvest ourselves." Unpacking Hyku Knapsack: Sustainable Customization With Less Upgrade Pain Notch8, United States of America Repository programs always need local customization (metadata defaults, workflows, branding, integrations). The trouble starts when those changes land in core code: upgrades become archaeology, diffs sprawl, and “we’ll upgrade later” turns into years. Hyku Knapsack is how we avoid that in the Samvera Hyku ecosystem. It’s a wrapper repository that keeps upstream Hyku as a Git submodule and keeps institution-specific code in the wrapper. Rails loads the wrapper first, so local overrides win without forking upstream. In this Developer Track session I’ll do a quick architecture overview, then a live demo: adding a custom override the knapsack way (mirror upstream paths, use _decorator.rb, prefer super, and use Module#prepend when needed). I’ll close with upgrade/migration tips and a simple rubric for what should live in your wrapper versus what belongs upstream. |
| 15:00 - 15:50 | Presentations: Scaling Agile Practices & APTrust Roadmap Location: Parallel Session 1 Session Chair A: Joseph Kraus, Colorado School of Mines Session Chair B: Erin Jerome, University of Massachusetts Amherst Recording available at https://doi.org/10.5281/zenodo.20778623 |
|
|
TigerData’s Big Year: Scaling Agile Practices Across Diverse Teams to Support Research Activity Princeton University, United States of America In May 2025, Princeton University launched TigerData, a research data management service designed to support storage, access, and evaluation of research data projects. This included the launch of the TigerData web portal, which supports service delivery and provides researchers with a UI to manage their projects. This paper examines the collaborative, multi-year process behind the portal's development, highlighting the sociotechnical challenges and organizational strategies required to deliver a production-ready service that spans across institutional boundaries. The project required sustained coordination among multiple departments at Princeton University as well as a cross-cutting governance group, each bringing distinct cultures, expertise, and operational constraints. We discuss early communication and coordination challenges, the impact of adopting a novel storage platform with a smaller preexisting community up front, and the need to integrate legacy projects that predated the mature service framework. We will also describe how intentional investments in trust-building, role clarification, and sharing experiences with agile practices at scale enabled progress. Key outcomes include a fully automated deployment pipeline and a workflow-driven project request workflow that supports efficient project provisioning. This case study offers practical lessons for library technologists designing complex, collaborative research infrastructure services at their institutions and beyond. New Horizons: How APTrust Developed a Community-Centered Technical Roadmap Process for Digital Preservation APTrust, United States of America This presentation will discuss APTrust’s journey in creating a new technical roadmap planning process. APTrust is a consortium based at the University of Virginia, dedicated to digital preservation and providing preservation storage across multiple geolocations to a variety of academic and non-academic member institutions. In 2025, our aim was to develop a new technical roadmap planning process driven by direct feedback from our members. Our new process included a comprehensive survey, focus groups, and data analysis. We were able to use the analysis from this new, feedback-driven paradigm to produce a robust technical roadmap organized into software goals, infrastructure goals, and security and risk management goals. As Lead Developer at APTrust, I am excited to share our process and answer questions from the audience about developing a technical roadmap for a digital preservation organization. |
| 16:00 - 16:50 | Panel: Balancing Act: Achieving Accessible and Sustainable Repositories Location: Parallel Session 1 Session Chair A: Heather Greer Klein, Samvera Session Chair B: Kimberly Chapman, University of Arizona Libraries Recording available at https://doi.org/10.5281/zenodo.20789004 |
|
|
Balancing Act: Achieving Accessible and Sustainable Repositories 1: University of Nevada, Reno, United States of America; 2: Montana State University, United States of America; 3: Indiana University Indianapolis, United States of America With the Spring 2026 deadline in the United States for online material to meet federal accessibility standards, ensuring that material in institutional repositories meets these standards has become a top priority for IR managers. At the same time, making IR content accessible can feel like an overwhelming task, especially in terms of remediation. This panel will discuss how three U.S. university libraries have addressed accessibility of materials in their IRs and plans for future deposits. Panelists will speak to the history of accessibility and their IR, how their university’s policy on accessibility has affected their work, and how they’ve approached working with authors and creators to ensure their works are accessible. How has accessibility requirements affected deposit rates? What does it mean for increased labor of IR staff? How is preservation affected? They will also discuss what roadblocks they continue to face, resources that have helped, and what more is still needed, especially from publishers and other third parties that benefit from scholarship. |
| 16:00 - 16:50 | Presentations: Metadata Workflows & Service Design Location: Parallel Session 2 Session Chair A: Richard William Fyson, CoSector, University of London Session Chair B: Allison Sherrick, Metropolitan New York Library Council Recording available at https://doi.org/10.5281/zenodo.20778625 |
|
|
Open research information in repositories - modelling repository metadata workflows 1: Barcelona Declaration on Open Research Information; 2: University of Maribor Library Institutional repositories (IRs) have many variations in their set up and operation, resulting in metadata workflows that are often complex, heterogeneous and insufficiently documented. Differences include technical providers and platforms, internal and external sources of metadata and how these are captured (including the relationship to CRIS systems), and the extent to which metadata are exposed for download and harvesting. To demonstrate current variability in approaches, and provide a framework to identify where improvements could be made in the use and availability of open research information in repositories, a working group of the Barcelona Declaration on Open Research Information designed a survey to identify a set of typical repository workflows. The survey aims to gather comparable, high-level information about how IRs collect, manage, and expose research information, from a workflow perspective.It asks about the repository’s scope and content, how and from which sources information is collected, and how metadata are made available to others (including interfaces provided, data formats and licensing). We will discuss Initial survey results (representing a variety of repositories, institutions and countries) and demonstrate how visualizing repository metadata workflows can be valuable for identifying shared patterns, recurring pain points, and approaches to address these. Institutional repository service design at NYU Libraries New York University, United States of America NYU Libraries provides its researchers with a range of individual repositories and related services; in the coming years we plan to merge and align these services as much as possible. To fully understand the challenges, and refine our requirements and solutions, we’ve engaged in a robust service design project. This presentation will describe the application of service design principles to the migration and relaunch of our university’s repository services. For the past two years, a group of diverse specialists has been meeting weekly. For the first year, our charge was discovery – charting the multidimensional journeys of current workflows and identifying opportunities for improvement. Currently, we are in the design phase – solutioning for a future suite of services, through collaborative ideation and convergence work. We will share advice for peer institutions interested in enhancing their own repository services through the use of this powerful approach. |