Automated Language Identification of Bibliographic Resources
Name
Morris_Automated_language_identification.docx
Description
visibility:open
Size
3.84 MB
Format
Microsoft Word XML
Checksum (CRC64NVME)
YMHslBDRQeQ=
Resource type
Journal article
Creator (person)
Morris, Victoria
Date published
December 20, 2019
Abstract
This article describes experiments in the use of machine learning techniques at the British Library to assign language codes to catalog records, in order to provide information about the language of content of the resources described. In the first phase of the project, language codes were assigned to 1.15 million records with 99.7% confidence. The automated language identification tools developed will be used to contribute to future enhancement of over 4 million legacy records.
Journal title
Cataloging & Classification Quarterly
Publisher
Taylor & Francis
Place of publication
UK
ISSN
0163-9374
eISSN
1544-4554
Date accepted
November 29, 2019
Official URL
Rights statement
In Copyright
Additional information
This is an Accepted Manuscript of an article published by Taylor & Francis Group in Cataloging & Classification Quarterly.