British Library Datasets
British Library Datasets
British Library Datasets
As part of our work to open our data to wider use, we make copies of some datasets available for research and creative purposes. We aim to describe collections in terms of their data format (images, full text, metadata, etc), licences, temporal and geographic scope, originating purpose (e.g. specific digitisation projects or exhibitions) and collection, and related subjects or themes.
We'd love to hear what you've done or made with the data.
34 results
Now showing 1 - 10 of 34
- Some of the metrics are blocked by yourconsent settings
Item type:Dataset, OCR text derived from digitised books published 1800 - 1809 in ALTO XML(2014) ;British LibraryBritish Library LabsThis set consists 1502 volumes, published between 1800-1809. The dataset comprises text from the collection of digitised books created using Optical Character Recognition (OCR) technology. The books cover a wide range of subject areas including philosophy, history, poetry and literature. The dataset is in Analysed Layout and Text Object (ALTO) Extensible Markup Language (XML) format. - Some of the metrics are blocked by yourconsent settings
Item type:Dataset, Books related to the Industrial Revolution derived from the Digitised 19th Century books dataset(2020) ;British LibraryBritish Library LabsA dataset which is a subset of the Digitised 19th Century Books dataset comprising books related to the Industrial Revolution in Britain. The subset of 354 items was refined by using keywords associated with placenames and the topic of industrialism. This dataset was curated by the Aepyi student group at University College London for the module BASC003, Information through the ages in Autumn 2019.3 1 - Some of the metrics are blocked by yourconsent settings
Item type:Dataset, OCR text derived from digitised books published 1870 - 1879 in ALTO XML(2014) ;British LibraryBritish Library LabsThis set consists 8630 volumes, published between 1870-1879. The dataset comprises text from the collection of digitised books created using Optical Character Recognition (OCR) technology. The books cover a wide range of subject areas including philosophy, history, poetry and literature. The dataset is in Analysed Layout and Text Object (ALTO) Extensible Markup Language (XML) format.11 2 - Some of the metrics are blocked by yourconsent settings
Item type:Dataset, OCR text derived from digitised books published 1890 - 1899 in ALTO XML(2014) ;British LibraryBritish Library LabsThis set consists 14847 volumes, published between 1890-1899. The dataset comprises text from the collection of digitised books created using Optical Character Recognition (OCR) technology. The books cover a wide range of subject areas including philosophy, history, poetry and literature. The dataset is in Analysed Layout and Text Object (ALTO) Extensible Markup Language (XML) format.2 2 - Some of the metrics are blocked by yourconsent settings
Item type:Dataset, OCR text derived from digitised books (unknown precise publication dates) in ALTO XML(2014) ;British LibraryBritish Library LabsThis set consists 284 volumes. The dataset comprises text from the collection of digitised books created using Optical Character Recognition (OCR) technology. The books cover a wide range of subject areas including philosophy, history, poetry and literature. The dataset is in Analysed Layout and Text Object (ALTO) Extensible Markup Language (XML) format.1 - Some of the metrics are blocked by yourconsent settings
Item type:Dataset, OCR text derived from digitised books published 1810 - 1819 in ALTO XML(2014) ;British LibraryBritish Library LabsThis set consists 2338 volumes, published between 1810-1819. The dataset comprises text from the collection of digitised books created using Optical Character Recognition (OCR) technology. The books cover a wide range of subject areas including philosophy, history, poetry and literature. The dataset is in Analysed Layout and Text Object (ALTO) Extensible Markup Language (XML) format.15 2 - Some of the metrics are blocked by yourconsent settings
Item type:Dataset, Digitised Quarterly Lists PDFs and Metadata(2016)Derrick, TomThe files in this dataset are derived from the British Library’s collection of bound volume Quarterly Lists: printed catalogue records of Indian books published quarterly and by province of British India between 1867 and 1947. The dataset comprises full-text searchable PDFs of 215 volumes as well as the associated metadata for each volume and represents a rich source for researchers interested in the publishing industry and book history in India. The catalogues are predominantly in English language with some Indian scripts and mostly arranged in table format, capturing descriptive metadata about the books, including the name and addresses of printers and publishers, the number of copies printed and often the price, as well as much more. The catalogues have been made available through the British Library's Two Centuries of Indian Print project, which is also digitising rare Bengali books dating from 1713-1914, the datasets of which will also be made available through this website.43 53 - Some of the metrics are blocked by yourconsent settings
Item type:Dataset, OCR text derived from digitised books published 1850 - 1859 in ALTO XML(2014) ;British LibraryBritish Library LabsThis set consists 5818 volumes, published between 1850-1859. The dataset comprises text from the collection of digitised books created using Optical Character Recognition (OCR) technology. The books cover a wide range of subject areas including philosophy, history, poetry and literature. The dataset is in Analysed Layout and Text Object (ALTO) Extensible Markup Language (XML) format. - Some of the metrics are blocked by yourconsent settings
Item type:Dataset, Books related to theatre derived from the Digitised 19th Century Books dataset(2020) ;British LibraryBritish Library LabsA dataset derived from the Digitised 19th Century Books dataset which contains books pertaining to theatre written in English. The dataset of 841 items was created by filtering by keywords which are related to different genre of play including Drama, Act, Scene, Play, Comedy, Farce, Pantomime, Tragedy and Shakespeare and manual curation to identify works which were written to be performed. This dataset was curated by students at University College London for the module BASC003, Information through the ages in Autumn 2019. - Some of the metrics are blocked by yourconsent settings
Item type:Dataset, Pelagios Project: Liber insularum Cycladum. Arundel MS 93.art.7(2017)British LibraryThis dataset comprises 45 images from the Liber insularum Cycladum produced by Christophori Bondelmonti around 1422. The digitisation was sponsored by A. W. Mellon Foundation through the Pelagios Project. Due to UK copyright law they are technically in copyright until 2039. However, given the age of the manuscripts and their place of production the Library believes it highly unlikely a public domain release will offend anyone; more information can be found at: https://www.bl.uk/catalogues/illuminatedmanuscripts/reuse.asp1