Repository logo
Home
Research Outputs
Collections
Statistics
Shared Repository Homepage
  1. Home
  2. Cultural Heritage Shared Repository Service
  3. British Library
  4. Dataset
  5. Incunabula Printed Catalogue Dataset: Volume 11

Incunabula Printed Catalogue Dataset: Volume 11

Thumbnail Image
Download
Name

BMC_11.zip

Description
visibility:open
Size

552.59 MB

Format

Unknown

Checksum (CRC64NVME)

B55MPpUYkYE=

Resource type
Dataset
Creator (Person)
Croen, Jeanette
ORCIDORCID logo
Date published
2026
Abstract
This dataset contains the source data for Jeanette Croen's 25/26 PhD placement. Digital images of volume XI of the 'Catalogue of books printed in the 15th century now at the British Museum' (BMC) were uploaded to Optical Character Recognition platform Transkribus. The images were transcribed using two separate layout models, one where the page contains two columns, and one where it contains four columns. Outputs are grouped into corresponding BMC_11_2 and BMC_11_4 subfolders, with the following naming convention XXXX_LD_31_b_`730_YYYY. The YYYY number can be used to intercalate the 2 column and 4 column images. The XXXX number is a four digit number indicating an incorrect reading order and should be disregarded. Two files are exported for each page, the original image and an xml containing the transcribed text.
Contributor (person)
Atanassova, Rossitza
ORCIDORCID logo
Lloyd, Harry
ORCIDORCID logo
Steiner, Alyssa
ORCIDORCID logo
Project(s)
Computational Analysis of Early Printed Book Descriptions
Publisher
British Library
Place of publication
UK
Official URL
https://doi.org/10.23636/2xte-t302
Licence
https://creativecommons.org/publicdomain/mark/1.0/
https://creativecommons.org/licenses/by/4.0/
DOI
10.23636/2xte-t302
Keywords
AI
book history
optical character recognition
Metadata
catalogues
Incunabula
early printed books
Additional information
The dataset was created using automated layout analysis, text recognition and catalogue entry detection. Transkribus was used to train a structure model for the pages with catalogue entries and the code used for the entry detection was developed by Isaac Dunford and improved by Harry Lloyd. More information is available at https://github.com/britishlibrary/Incunabula-Catalogue-Entry-Detection . The dataset includes the raw and processed data outputs as described in the readme file. The text is released under CC0 1.0 Universal Public Domain and the images are released under CC-BY 4.0 license. Contact us at digitalresearch@bl.uk if you have questions, or to let us know how you used our data.
Collection
Computational analysis of early printed book descriptions (PhD Placement)
Managed by the British Library and supported by the AHRC

Built with DSpace-CRIS software - Extension maintained and optimized by 4Science

  • Cookie settings
  • End User Agreement
  • About
  • Contact
  • Help
Repository logo COAR Notify