Repository logo
Home
Research Outputs
Collections
Statistics
Shared Repository Homepage
  1. Home
  2. Cultural Heritage Shared Repository Service
  3. British Library
  4. Conference Item
  5. Cross-disciplinary Collaborations to Enrich Access to Non-Western Language Material in the Cultural Heritage Sector

Cross-disciplinary Collaborations to Enrich Access to Non-Western Language Material in the Cultural Heritage Sector

Thumbnail Image
Download
Name

McGregor_Derrick_DATeCH2019.pdf

Description
visibility:open
Size

454.47 KB

Format

Adobe PDF

Checksum (CRC64NVME)

wYgnOCwKQ64=

Resource type
Conference paper (published)
Creator (person)
Derrick, Tom
McGregor, Nora
Date published
2019
Abstract
The British Library is home to millions of items representing every age of written civilisation, including books, manuscripts and newspapers in all written languages. Large digitisation programmes currently underway are opening up access to this rich and unique historical content on an ever increasing scale. However, particularly for historical material written in non-Latin scripts, enabling enriched full-text discovery and analysis across the digitised output, something which would truly transform access and scholarship, is still out of reach. This is due in part to commercial text recognition solutions currently on the market today having largely been optimised for modern documents and Latin scripts. This paper will report on a series of initiatives undertaken by the British Library to investigate, evaluate and support new research into enhancing text recognition capabilities for two major digitised collections of non-Western language collections: printed Bangla and handwritten Arabic. It seeks to present lessons learned and opportunities gained from cross-disciplinary collaboration between the cultural heritage sector and researchers working at the cutting edge of text recognition, with a view towards informing and encouraging future such partnerships.
Event title
DATeCH2019: 3rd International Conference on Digital Access to Cultural Textual Heritage
Publisher
ACM Press
Official URL
https://doi.org/10.1145/3322905.3322907
Rights statement
In Copyright
DOI
10.1145/3322905.3322907
Keywords
datasets
page analysis
OCR
Arabic script
HTR
recognition
layout analysis
Bangla script
Additional information
The attached file is the authors' accepted manuscript of this paper.
Managed by the British Library and supported by the AHRC

Built with DSpace-CRIS software - Extension maintained and optimized by 4Science

  • Cookie settings
  • End User Agreement
  • About
  • Contact
  • Help
Repository logo COAR Notify