Computational analysis of early printed book descriptions (PhD Placement)
Computational analysis of early printed book descriptions (PhD Placement)
Computational analysis of early printed book descriptions (PhD Placement)
The British Library’s collections of early printed books published in the 15th century (known as ‘incunabula’) comprises ca 23,000 volumes and is one of the most important and frequently used resources by researchers. The incunabula collection was catalogued in the ‘Catalogue of books printed in the 15th century now at the British Museum [or Library]’ (known as the BMC) over a period of 100 years between 1908 – 2007, and the detailed descriptions are a valuable source of information about the collection content, acquisition, transmission, previous ownership, and materiality. As part of a previous digital scholarship project (https://app.transkribus.org/sites/BL-Incunabula) we used computational methods with the incunabula descriptions from BMC volumes 1-10 as means of improving our understanding of the collection’s cataloguing history and curatorial practice. The PhD placement will help us to progress this work for two further volumes of the BMC and prepare more incunabula descriptions data for publication and analysis. The Library’s digital scholarship projects have demonstrated the value of providing access to our collections catalogues as data for research, public engagement, innovation and creativity. This work makes an important contribution to the Library’s mission and especially to the Race Equality Action Plan as it creates and enhances metadata to remove barriers to discovery and improves the accessibility of our collection.
Collection Details
Now showing 1 - 8 of 8
- Some of the metrics are blocked by yourconsent settings
Item type:Dataset, BMC 11 Catalogue Entry Text (English only)(2026)Lloyd, HarryThis dataset contains dervied text data for Jeanette Croen's 25/26 PhD placement. The XMLs containing transciptions of images of volume XI of the 'Catalogue of books printed in the 15th century now at the British Museum' (BMC) were parsed using the code in 'End of PhD Placement Release - Python Entry Extraction Code' code entry in this collection. The parsing extracts all catalogue entries from BMC XI which are then run through a language recognition algorithm to extract only the English text, before being combined into a txt file with one catalogue entry per line.4 4 - Some of the metrics are blocked by yourconsent settings
Item type:Dataset, Incunabula Printed Catalogue Dataset: Volume 11(2026)Croen, JeanetteThis dataset contains the source data for Jeanette Croen's 25/26 PhD placement. Digital images of volume XI of the 'Catalogue of books printed in the 15th century now at the British Museum' (BMC) were uploaded to Optical Character Recognition platform Transkribus. The images were transcribed using two separate layout models, one where the page contains two columns, and one where it contains four columns. Outputs are grouped into corresponding BMC_11_2 and BMC_11_4 subfolders, with the following naming convention XXXX_LD_31_b_`730_YYYY. The YYYY number can be used to intercalate the 2 column and 4 column images. The XXXX number is a four digit number indicating an incorrect reading order and should be disregarded. Two files are exported for each page, the original image and an xml containing the transcribed text.2 - Some of the metrics are blocked by yourconsent settings
Item type:Journal article, Ethical AI at the British Library, A PhD Placement Project(2026-04-15)Croen, JeanetteThis article is a scholarly reflection on using the AI tools Transkribus and AntConc as part of a digital humanities PhD Placement within the British Library to extract metadata from printed catalogues for the online catalogue. This project focuses on BMC XI, the catalogue of English Incunabula at the British Library published in 2007. Transkribus is a “comprehensive platform for the digitization, AI-powered text recognition, transcription, and searching of historical documents” while AntConc is a “freeware corpus analysis toolkit for concordancing and text analysis”. Together, these tools can be used to extract information en masse to be uploaded to the specialist databases MEI (Material Evidence Incunabula) and the ISTC (Incunabula Short Title Catalogue) as well as pick out patterns and trends within incunabula descriptions. This project followed FRAIM (Framing responsible AI implementation and management) principles.2 2 - Some of the metrics are blocked by yourconsent settings
Item type:Presentation, Computational Analysis of Early Printed Book Descriptions(2026)Croen, JeanetteThis presentation was delivered at the British Library’s Digital Scholarship meeting about Jeanette Croen's 2025/26 PhD placement at the British Library. The research used digital methods to extract data from volume XI of the 'Catalogue of books printed in the 15th century now at the British Museum' (BMC). The presentation describes the background to the research, project structure, transcription work using Transkribus, and analysis using AntConc and Python.1 2 - Some of the metrics are blocked by yourconsent settings
Item type:Blog post, Computational Analysis of Book Descriptions A Placement Project(2026-04-13)Croen, JeanetteThis blog post was written for the British Library’s Digital Research blog about Jeanette Croen's 2025/26 PhD placement at the British Library. The research used digital methods to extract data from volume XI of the 'Catalogue of books printed in the 15th century now at the British Museum' (BMC). This post describes the background to the research, project structure, transcription work using Transkribus, and analysis using AntConc and Python.3 3 - Some of the metrics are blocked by yourconsent settings
Item type:ConferenceItem Conference poster (unpublished), Computational Analysis of Early Printed Book Descriptions(2026)Croen, JeanetteThis presentation was presented at the Collections and Curation Open Day on 02/02/2026 on Jeanette Croen's 2025/26 PhD placement at the British Library. The research used digital methods to extract data from volume XI of the 'Catalogue of books printed in the 15th century now at the British Museum' (BMC). The poster describes the background to the research, project structure, transcription work using Transkribus, and analysis using Python. - Some of the metrics are blocked by yourconsent settings
Item type:Software, End of PhD Placement Release - Python Entry Extraction Code(2026)Croen, JeanetteThis repository release contains a notebook produced by Jeanette Croen for her 2025/26 PhD placement project. The zip file contains and End of Project Release of the Catalogue Entry Extraction Python code, a v1.1.0 release targeting the 'field-model' branch that contains a notebooks\check_field_assignment.ipynb Jupyter notebook Jeanette created to identify whether fields labels assigned by a field model on Transkribus are correct. The output of the code isn't exported from the notebook.5 - Some of the metrics are blocked by yourconsent settings
Item type:Dataset, Incunabula Printed Catalogue Dataset Metadata: Volume 11(2026)Lloyd, HarryThis dataset contains dervied text data for Jeanette Croen's 25/26 PhD placement. The XMLs containing transciptions of images of volume XI of the 'Catalogue of books printed in the 15th century now at the British Museum' (BMC) were parsed using the code in 'End of PhD Placement Release - Python Entry Extraction Code' code entry in this collection. The parsing extracts all catalogue entries from BMC XI which are then exported in csv format with an entry_text column containing the entry text and the remaining columns containing metadata about the entry, including which XMLs in the raw data that entry was extracted from.3 1