DeezyMatch training set for OCR
Resource type
Dataset
Date published
July 2023
Abstract
Optical character recognition (OCR) is the process of automatically transcribing text from images. The presence of OCR-induced errors in digitised text is a common problem in the digital humanities. OCR errors are usually due to the misrecognition of characters, such as "h" recognised as "b", or "c" recognised as "o". The DeezyMatch library was built to address this issue through fuzzy string matching, using a deep neural network approach. In order to train a DeezyMatch model, a training set consisting of positive and negative string pairs is needed. We present a new dataset of positive and negative OCR variations, which can be used to train a DeezyMatch model, which can then be used for fuzzy string matching for the downstream task of entity linking. This dataset has been automatically generated from word2vec embeddings trained on digitised historical news texts, and has been expanded with toponym alternate names extracted from Wikipedia.
Project(s)
Living with Machines
Funder
| Funder name | Awards |
Arts and Humanities Research Council | AH/S01179X/1 |
Alan Turing Institute | EP/N510129/1 |
Publisher
British Library
Place of publication
UK
Related identifier
| Identifier | Type | Relation |
10.5281/zenodo.7887305 | DOI | |
Additional information
If you use this dataset, please cite: "Mariona Coll Ardanuy, Federico Nanni and Nilo Pedrazzini. 2023. 'DeezyMatch training set for OCR'. British Library Research Repository."