Repository logo
Home
Research Outputs
Collections
Statistics
Shared Repository Homepage
  1. Home
  2. Cultural Heritage Shared Repository Service
  3. British Library
  4. Dataset
  5. DeezyMatch training set for OCR

DeezyMatch training set for OCR

Thumbnail Image
Download
Name

README.txt

Description
visibility:open
Size

4.08 KB

Format

Text

Checksum (CRC64NVME)

FRg+MGcrgFc=

Thumbnail Image
Download
Name

w2v_ocr_pairs.txt

Description
visibility:open
Size

24.28 MB

Format

Text

Checksum (CRC64NVME)

pmd3/SlBgyk=

Resource type
Dataset
Creator (Person)
Coll Ardanuy, Mariona
ORCIDORCID logo
Nanni, Federico
ORCIDORCID logo
Pedrazzini, Nilo
ORCIDORCID logo
Date published
July 2023
Abstract
Optical character recognition (OCR) is the process of automatically transcribing text from images. The presence of OCR-induced errors in digitised text is a common problem in the digital humanities. OCR errors are usually due to the misrecognition of characters, such as "h" recognised as "b", or "c" recognised as "o". The DeezyMatch library was built to address this issue through fuzzy string matching, using a deep neural network approach. In order to train a DeezyMatch model, a training set consisting of positive and negative string pairs is needed. We present a new dataset of positive and negative OCR variations, which can be used to train a DeezyMatch model, which can then be used for fuzzy string matching for the downstream task of entity linking. This dataset has been automatically generated from word2vec embeddings trained on digitised historical news texts, and has been expanded with toponym alternate names extracted from Wikipedia.
Project(s)
Living with Machines
Funder
Funder nameAwards
Arts and Humanities Research Council
AH/S01179X/1
Alan Turing Institute
EP/N510129/1
Publisher
British Library
Place of publication
UK
Licence
https://creativecommons.org/licenses/by/4.0/
Related identifier
IdentifierTypeRelation
10.5281/zenodo.7887305
DOI
Keywords
natural language processing
OCR
Living with Machines
DeezyMatch
string variation
fuzzy string matching
newspapers
digital humanities
Additional information
If you use this dataset, please cite: "Mariona Coll Ardanuy, Federico Nanni and Nilo Pedrazzini. 2023. 'DeezyMatch training set for OCR'. British Library Research Repository."
Managed by the British Library and supported by the AHRC

Built with DSpace-CRIS software - Extension maintained and optimized by 4Science

  • Cookie settings
  • End User Agreement
  • About
  • Contact
  • Help
Repository logo COAR Notify