Repository logo
Home
Research Outputs
Collections
Statistics
Shared Repository Homepage
  1. Home
  2. Cultural Heritage Shared Repository Service
  3. British Library
  4. Article
  5. Automated Language Identification of Bibliographic Resources

Automated Language Identification of Bibliographic Resources

Thumbnail Image
Download
Name

Morris_Automated_language_identification.docx

Description
visibility:open
Size

3.84 MB

Format

Microsoft Word XML

Checksum (CRC64NVME)

YMHslBDRQeQ=

Resource type
Journal article
Creator (person)
Morris, Victoria
ORCIDORCID logo
Date published
December 20, 2019
Abstract
This article describes experiments in the use of machine learning techniques at the British Library to assign language codes to catalog records, in order to provide information about the language of content of the resources described. In the first phase of the project, language codes were assigned to 1.15 million records with 99.7% confidence. The automated language identification tools developed will be used to contribute to future enhancement of over 4 million legacy records.
Journal title
Cataloging & Classification Quarterly
Publisher
Taylor & Francis
Place of publication
UK
ISSN
0163-9374
eISSN
1544-4554
Date accepted
November 29, 2019
Official URL
https://doi.org/10.1080/01639374.2019.1700201
Rights statement
In Copyright
DOI
10.1080/01639374.2019.1700201
Keywords
machine learning
automatic metadata generation
language identification
legacy record enhancement
metadata
Additional information
This is an Accepted Manuscript of an article published by Taylor & Francis Group in Cataloging & Classification Quarterly.
Managed by the British Library and supported by the AHRC

Built with DSpace-CRIS software - Extension maintained and optimized by 4Science

  • Cookie settings
  • End User Agreement
  • About
  • Contact
  • Help
Repository logo COAR Notify