British Library Datasets

Now showing 1 - 10 of 34
  • Thumbnail Image
    Some of the metrics are blocked by your 
    Item type:Dataset,
    OCR text derived from digitised books published 1800 - 1809 in ALTO XML
    (2014)
    British Library
    ;
    British Library Labs
    This set consists 1502 volumes, published between 1800-1809. The dataset comprises text from the collection of digitised books created using Optical Character Recognition (OCR) technology. The books cover a wide range of subject areas including philosophy, history, poetry and literature. The dataset is in Analysed Layout and Text Object (ALTO) Extensible Markup Language (XML) format.
  • Thumbnail Image
    Some of the metrics are blocked by your 
    Item type:Dataset,
    Books related to the Industrial Revolution derived from the Digitised 19th Century books dataset
    (2020)
    British Library
    ;
    British Library Labs
    A dataset which is a subset of the Digitised 19th Century Books dataset comprising books related to the Industrial Revolution in Britain. The subset of 354 items was refined by using keywords associated with placenames and the topic of industrialism. This dataset was curated by the Aepyi student group at University College London for the module BASC003, Information through the ages in Autumn 2019.
      3  1
  • Thumbnail Image
    Some of the metrics are blocked by your 
    Item type:Dataset,
    OCR text derived from digitised books published 1870 - 1879 in ALTO XML
    (2014)
    British Library
    ;
    British Library Labs
    This set consists 8630 volumes, published between 1870-1879. The dataset comprises text from the collection of digitised books created using Optical Character Recognition (OCR) technology. The books cover a wide range of subject areas including philosophy, history, poetry and literature. The dataset is in Analysed Layout and Text Object (ALTO) Extensible Markup Language (XML) format.
      11  2
  • Thumbnail Image
    Some of the metrics are blocked by your 
    Item type:Dataset,
    OCR text derived from digitised books published 1890 - 1899 in ALTO XML
    (2014)
    British Library
    ;
    British Library Labs
    This set consists 14847 volumes, published between 1890-1899. The dataset comprises text from the collection of digitised books created using Optical Character Recognition (OCR) technology. The books cover a wide range of subject areas including philosophy, history, poetry and literature. The dataset is in Analysed Layout and Text Object (ALTO) Extensible Markup Language (XML) format.
      2  2
  • Thumbnail Image
    Some of the metrics are blocked by your 
    Item type:Dataset,
    OCR text derived from digitised books (unknown precise publication dates) in ALTO XML
    (2014)
    British Library
    ;
    British Library Labs
    This set consists 284 volumes. The dataset comprises text from the collection of digitised books created using Optical Character Recognition (OCR) technology. The books cover a wide range of subject areas including philosophy, history, poetry and literature. The dataset is in Analysed Layout and Text Object (ALTO) Extensible Markup Language (XML) format.
      1
  • Thumbnail Image
    Some of the metrics are blocked by your 
    Item type:Dataset,
    OCR text derived from digitised books published 1810 - 1819 in ALTO XML
    (2014)
    British Library
    ;
    British Library Labs
    This set consists 2338 volumes, published between 1810-1819. The dataset comprises text from the collection of digitised books created using Optical Character Recognition (OCR) technology. The books cover a wide range of subject areas including philosophy, history, poetry and literature. The dataset is in Analysed Layout and Text Object (ALTO) Extensible Markup Language (XML) format.
      15  2
  • Thumbnail Image
    Some of the metrics are blocked by your 
    Item type:Dataset,
    Digitised Quarterly Lists PDFs and Metadata
    (2016)
    Derrick, Tom
    The files in this dataset are derived from the British Library’s collection of bound volume Quarterly Lists: printed catalogue records of Indian books published quarterly and by province of British India between 1867 and 1947. The dataset comprises full-text searchable PDFs of 215 volumes as well as the associated metadata for each volume and represents a rich source for researchers interested in the publishing industry and book history in India. The catalogues are predominantly in English language with some Indian scripts and mostly arranged in table format, capturing descriptive metadata about the books, including the name and addresses of printers and publishers, the number of copies printed and often the price, as well as much more. The catalogues have been made available through the British Library's Two Centuries of Indian Print project, which is also digitising rare Bengali books dating from 1713-1914, the datasets of which will also be made available through this website.
      43  53
  • Thumbnail Image
    Some of the metrics are blocked by your 
    Item type:Dataset,
    OCR text derived from digitised books published 1850 - 1859 in ALTO XML
    (2014)
    British Library
    ;
    British Library Labs
    This set consists 5818 volumes, published between 1850-1859. The dataset comprises text from the collection of digitised books created using Optical Character Recognition (OCR) technology. The books cover a wide range of subject areas including philosophy, history, poetry and literature. The dataset is in Analysed Layout and Text Object (ALTO) Extensible Markup Language (XML) format.
  • Thumbnail Image
    Some of the metrics are blocked by your 
    Item type:Dataset,
    Books related to theatre derived from the Digitised 19th Century Books dataset
    (2020)
    British Library
    ;
    British Library Labs
    A dataset derived from the Digitised 19th Century Books dataset which contains books pertaining to theatre written in English. The dataset of 841 items was created by filtering by keywords which are related to different genre of play including Drama, Act, Scene, Play, Comedy, Farce, Pantomime, Tragedy and Shakespeare and manual curation to identify works which were written to be performed. This dataset was curated by students at University College London for the module BASC003, Information through the ages in Autumn 2019.
  • Thumbnail Image
    Some of the metrics are blocked by your 
    Item type:Dataset,
    Pelagios Project: Liber insularum Cycladum. Arundel MS 93.art.7
    (2017)
    British Library
    This dataset comprises 45 images from the Liber insularum Cycladum produced by Christophori Bondelmonti around 1422. The digitisation was sponsored by A. W. Mellon Foundation through the Pelagios Project. Due to UK copyright law they are technically in copyright until 2039. However, given the age of the manuscripts and their place of production the Library believes it highly unlikely a public domain release will offend anyone; more information can be found at: https://www.bl.uk/catalogues/illuminatedmanuscripts/reuse.asp
      1