Repository logo
Home
Research Outputs
Collections
Statistics
Shared Repository Homepage
  1. Home
  2. Cultural Heritage Shared Repository Service
  3. British Library
  4. Conference Item
  5. Arabic dialect identification in the context of bivalency and code-switching

Arabic dialect identification in the context of bivalency and code-switching

Thumbnail Image
Download
Name

2018_paper.pdf

Description
visibility:open
Size

147.82 KB

Format

Adobe PDF

Checksum (CRC64NVME)

JFacb7wt33o=

Resource type
Conference paper (published)
Creator (person)
El-Haj, Mahmoud
Rayson, Paul
Aboelezz, Mariam
Date published
2018
Abstract
In this paper we use a novel approach towards Arabic dialect identification using language bivalency and written code-switching. Bivalency between languages or dialects is where a word or element is treated by language users as having a fundamentally similar semantic content in more than one language or dialect. Arabic dialect identification in writing is a difficult task even for humans due to the fact that words are used interchangeably between dialects. The task of automatically identifying dialect is harder and classifiers trained using only n-grams will perform poorly when tested on unseen data. Such approaches require significant amounts of annotated training data which is costly and time consuming to produce. Currently available Arabic dialect datasets do not exceed a few hundred thousand sentences, thus we need to extract features other than word and character n-grams. In our work we present experimental results from automatically identifying dialects from the four main Arabic dialect regions (Egypt, North Africa, Gulf and Levant) in addition to Standard Arabic. We extend previous work by incorporating additional grammatical and stylistic features and define a subtractive bivalency profiling approach to address issues of bivalent words across the examined Arabic dialects. The results show that our new methods classification accuracy can reach more than 76% and score well (66%) when tested on completely unseen data.
Editor
Calzolari, Nicoletta
Choukri, Khalid
Cieri, Christopher
Declerck, Thierry
Goggi, Sara
Hasida, Kà´iti
Isahara, Hitoshi
Maegaard, Bente
Mariani, Joseph
Mazo, Hélène
Moreno, Asuncion
Odijk, Jan
Piperidis, Stelios
Tokunaga, Takenobu
Event title
LREC 2018, Eleventh International Conference on Language Resources and Evaluation
Publisher
European Language Resources Association
Official URL
http://www.lrec-conf.org/proceedings/lrec2018/pdf/237.pdf
Rights statement
In Copyright
Licence
https://creativecommons.org/licenses/by-nc/4.0/
Keywords
machine learning
NLP
dialects
language identification
Arabic
bivalency
Managed by the British Library and supported by the AHRC

Built with DSpace-CRIS software - Extension maintained and optimized by 4Science

  • Cookie settings
  • End User Agreement
  • About
  • Contact
  • Help
Repository logo COAR Notify