Decade-level Word2Vec models from automatically transcribed 19th-century newspapers digitised by the British Library (1800-1919)
Resource type
Dataset
Creator (Person)
Pedrazzini, Nilo
Date published
May 2, 2023
Abstract
Word embeddings trained on a 4.2-billion-word corpus of 19th-century British newspapers using Word2Vec and specific parameters.
The embeddings are divided into periods of ten years each. Unlike those in this repository, these were not aligned and OCR errors skimmed from the vocabulary.
See related GitHub repository for the full documentation: https://github.com/Living-with-machines/DiachronicEmb-BigHistData.
Project website (Living with Machines): https://livingwithmachines.ac.uk/
Contributor (person)
Ridge, Mia
Contributor (organisation)
British Library
Living with Machines
Project(s)
Living with Machines
Funder
| Funder name | Awards |
UK Research and Innovation | AH/S01179X/1 |
Arts and Humanities Research Council | AH/S01179X/1 |
Version
1
Publisher
Zenodo
Official URL
Related identifier
| Identifier | Type | Relation |
10.5281/zenodo.7181682 | DOI | iscontinuedby |
10.5281/zenodo.7887304 | DOI | isversionof |
Collection