Repository logo
Home
Research Outputs
Collections
Statistics
Shared Repository Homepage
  1. Home
  2. Cultural Heritage Shared Repository Service
  3. British Library
  4. Dataset
  5. 19th Century United States Newspaper Advert images with 'illustrated' or 'non illustrated' labels

19th Century United States Newspaper Advert images with 'illustrated' or 'non illustrated' labels

Thumbnail Image
Download
Name

sample.csv

Description
visibility:open
Size

879.5 KB

Format

CSV

Checksum (CRC64NVME)

EWU9G/ARVR4=

Thumbnail Image
Download
Name

newspaper-navigator-sample-metadata.csv

Description
visibility:open
Size

1014.07 KB

Format

CSV

Checksum (CRC64NVME)

nPIaWQWLCsc=

Resource type
Dataset
Creator (Person)
van Strien, Daniel
ORCIDORCID logo
Date published
2021
Abstract
The Dataset contains images derived from the Newspaper Navigator (news-navigator.labs.loc.gov/), a dataset of images drawn from the Library of Congress Chronicling America collection (chroniclingamerica.loc.gov/). [The Newspaper Navigator dataset] consists of extracted visual content for 16,358,041 historic newspaper pages in Chronicling America. The visual content was identified using an object detection model trained on annotations of World War 1-era Chronicling America pages, including annotations made by volunteers as part of the Beyond Words crowdsourcing project. source: https://news-navigator.labs.loc.gov/ One of these categories is 'advertisements. This dataset contains a sample of these images with additional labels indicating if the advert is 'illustrated' or 'not illustrated'. The data is organised as follows: The images themselves can be found in `images.zip` `newspaper-navigator-sample-metadata.csv` contains metadata about each image drawn from the Newspaper Navigator Dataset. `ads.csv` contains the labels for the images as a CSV file `sample.csv` contains additional metadata about the images (based on the newspapers those images came from). This dataset was created for use in an under-review Programming Historian tutorial (http://programminghistorian.github.io/ph-submissions/lessons/computer-vision-deep-learning-pt1) The primary aim of the data was to provide a realistic example dataset for teaching computer vision for working with digitised heritage material. The data is shared here since it may be useful for others. This data documentation is a work in progress and will be updated when the Programming Historian tutorial is released publicly. The metadata CSV file contains the following columns: - filepath
- pub_date
- page_seq_num
- edition_seq_num
- batch
- lccn
- box
- score
- ocr
- place_of_publication
- geographic_coverage
- name
- publisher
- url
- page_url
- month
- year
- iiif_url
Contributor (organisation)
Living with Machines
Project(s)
Living with Machines
Funder
Funder nameAwards
Arts and Humanities Research Council (AHRC)
AH/S01179X/1
Version
0.0.1
Publisher
Zenodo
Official URL
https://doi.org/10.5281/zenodo.5838410
Related URL
https://zenodo.org/record/5838410
DOI
10.5281/zenodo.5838410
Related identifier
IdentifierTypeRelation
https://arxiv.org/abs/2005.01583
URL
isderivedfrom
10.5281/zenodo.5537185
DOI
requires
10.5281/zenodo.4075210
DOI
isversionof
Keywords
machine learning
GLAM
historic newspapers
computer vision
Collection
Living with Machines
Managed by the British Library and supported by the AHRC

Built with DSpace-CRIS software - Extension maintained and optimized by 4Science

  • Cookie settings
  • End User Agreement
  • About
  • Contact
  • Help
Repository logo COAR Notify