2026 CSDH/SCHN
Annual Conference
June 3rd to 5th, 2026
University of Montreal
Conference Agenda
Overview and details of the sessions of this conference. Please select a date or location to show only sessions at that day or location. Please select a single session for detailed view (with abstracts and downloads if available).
|
Daily Overview |
| Session | ||
Session 2.1
| ||
| Presentations | ||
Digitizing Materiality: Affordances and Limitations of DH for Tactile Books University of Alberta, Canada If a book or bookish object is primarily tactile in nature, how can that be digitized? Why should it be digitized? What affordances or limitations do digitization tools have for working with materials that are meant to be touched? As part of my postdoctoral fellowship project looking at the history of tactile writing and reading, I am working with optical character recognition (OCR), 3D scanning, and 3D printing to digitize and remediate tactile materials. The OCR project is a continuation of work from my doctoral dissertation to train an OCR model to read and reproduce braille characters for the purpose of making rare braille books available digitally and physically to braille readers. This model will also benefit researchers who want to include braille materials in corpus-based projects but cannot read braille themselves. In preparation for an exhibit, I am also working with 3D scanning and 3D printing. Many older tactile codes are no longer in circulation since braille won the code wars, so remaining materials that use or produced those codes are rare and, in most cases, extremely fragile. But these materials were made to be touched. In order to allow—and encourage—people to interact with these materials tactilely, I am using 3D scanning and 3D printing to produce replicas. At the time of writing this abstract, I am working on my first 3D scan -> 3D print prototype with a bible embossed with the tactile code called moon, housed by the Bruce Peel Special Collections Library. I will repeat this process for materials related to various tactile codes (including braille, moon, and raised type) on an archives and museums trip in February–March 2026. At the time of this conference, I will be able to show several reproductions as examples of how these tools and methods can work for rare, fragile, tactile books and bookish objects. Although these digitizations are useful for reproducing the materials or making braille texts available digitally (to be read with digital braille devices), the digitizations themselves are facsimiles of the original tactile materials. OCR requires flat, scanned images of books and reproduces visual representations of braille text. The process of 3D scanning maintains some of that texture and tactile nature, but any instance of working with tactile materials on a screen is a simplification or a facsimile for the original textured materiality of that object. This paper will explore the affordances and limitations of using this technology and these digitization processes for tactile materials. This paper speaks to the conference theme of “untranslatable” by examining how digital humanities tools and methods can be used to digitize tactile materials in a way and with results that benefit average users of tactile materials. This work is at the intersection of book history, disability studies, and digital humanities, and it prioritizes making material more accessible for disabled readers. Reading the Unreadable: Optical Character Recognition of Early Modern Print University of Saskatchewan, Canada The quality of machine transcription –also known as Optical Character Recognition (OCR)– has posed a significant issue in the field of history. While digital historians are uniquely affected by this due to our use of machine transcription in large-scale text-mining projects, the issue is evident to every scholar in the field. Almost everyone who has used a digital archive will recall how poorly most search functions work because of transcription errors. These errors typically worsen with the age of the document, with early modern documents commonly achieving OCR accuracy of less than 50%. This narrative is changing with the help of Large Language Models (LLMs). Newly developed open-source tools using Visual Language Models (VLMs) have demonstrated promising results, achieving near-perfect transcription of modern documents. My research leverages these industry tools to demonstrate their effectiveness in transcribing early modern English documents dated 1600-1800. Testing seven of these new tools against one another and three commonly used OCR tools of the past has revealed that new VLM-based OCR models outperform traditional models in every metric, achieving over 90% accuracy when transcribing early modern English documents. The testing methodology involved hand-transcribing 100 gold-standard pages, chosen at random from a selection of over 20,000 pages. I then tested 10 different OCR models on these same 100 pages, comparing their results to the created gold standard. This comparison examined several metrics, including Word Error Rate, Significant Word Accuracy, BLEU score, and Hallucination rate. Throughout testing, olmOCRv2, an open-source, trainable model, came out on top across almost every metric, showing an over 80% increase in quality compared to the models used 10 years ago. Previous and ongoing research by historians such as Mark Humphries has demonstrated the value of using General-Purpose LLMs like Gemini for transcribing historical documents. However, these generalized models can only go so far in transcribing older historical documents, as they are primarily trained on modern works. The models I present are specifically designed and trained for OCR, yielding more accurate and reliable results than Gemini. Additionally, models like olmOCRv2 are small enough to run on personal computers –where this testing was done– and are trainable, meaning new historical documents can be fed into the models to improve their accuracy further. These new VLM-based models are important to the future of digital humanities and historical studies for multiple reasons. OCR transcription with over 90% accuracy enables us to conduct text mining and linguistic studies at a larger scale with more accurate results. Simple keyword searching through documents is also made more effective, with a guarantee that more of the searched words appear in the results. Notably, the early modern age of these documents posed additional challenges in terms of document quality, unique character forms, and obsolete spellings. In solving these challenges, this technology has proven itself not only effective for researchers of the modern era but also a valuable tool for almost all scholars. Semantics of Empire: Machine Translation, Artificial Intelligence, and the Case of Ottoman Turkish Stanford University, United States of America Recent developments in computational technologies, particularly Large Language Models (LLMs), have transformed how we engage with text in academia and beyond. However, these developments are not distributed equally across languages. Low-resourced and historical languages remain underrepresented, shaping whose languages appear in computational research and education during the digital turn. This presentation focuses on Ottoman Turkish and the specific task of machine translation for this language. I discuss experiments with neural machine translation model training, including finding suitable datasets and models, as well as performing sentence alignment to create more training data. I also examine how LLMs (ChatGPT, Gemini, Claude) handle Ottoman translation. By testing these models on real historical texts covering diverse topics, rather than standard LLM benchmarks, I reveal insights into model capabilities and limitations that conventional test data cannot capture. This approach demonstrates how historical research can serve as a unique lens for understanding and critiquing contemporary technology. Finally, I situate this work within the entwined histories of machine translation and Middle Eastern studies in the U.S., tracing patterns of information extraction from the Cold War through the post-9/11 era to argue that we, as historians and scholars in Ottoman and Middle Eastern studies, can actively intervene in these technological developments to imagine and build applications that foster mutual understanding rather than surveillance and extraction. | ||
