Arabic Corpus Enhancement using a New Lexicon/Stemming Algorithm

Ashraf AbdelRaouf, Colin A. Higgins, Tony Pridmore, Mahmoud I. Khalil

2013

Abstract

Optical Character Recognition (OCR) is an important technology and has many advantages in storing information for both old and new documents. The Arabic language lacks both the variety of OCR systems and the depth of research relative to Roman scripts. An authoritative corpus is beneficial in the design and construction of any OCR system. Lexicon and stemming tools are essential in enhancing corpus retrieval and performance in an OCR context. A new lexicon/stemming algorithm is presented based on the Viterbi path method which uses a light stemmer approach. Lexicon and stemming lookup is combined to obtain a list of alternatives for uncertain words. This list removes affixes (prefixes or suffices) if there are any; otherwise affixes are added. Finally, every word in the list of alternatives is verified by searching the original corpus. The lexicon/stemming algorithm also assures the continuous updating of the contents of the corpus presented by (AbdelRaouf et al., 2010), which copes with the innovative needs of Arabic OCR research.

Download


Paper Citation


in Harvard Style

AbdelRaouf A., A. Higgins C., Pridmore T. and I. Khalil M. (2013). Arabic Corpus Enhancement using a New Lexicon/Stemming Algorithm . In Proceedings of the 2nd International Conference on Pattern Recognition Applications and Methods - Volume 1: ICPRAM, ISBN 978-989-8565-41-9, pages 435-440. DOI: 10.5220/0004260704350440

in Bibtex Style

@conference{icpram13,
author={Ashraf AbdelRaouf and Colin A. Higgins and Tony Pridmore and Mahmoud I. Khalil},
title={Arabic Corpus Enhancement using a New Lexicon/Stemming Algorithm},
booktitle={Proceedings of the 2nd International Conference on Pattern Recognition Applications and Methods - Volume 1: ICPRAM,},
year={2013},
pages={435-440},
publisher={SciTePress},
organization={INSTICC},
doi={10.5220/0004260704350440},
isbn={978-989-8565-41-9},
}


in EndNote Style

TY - CONF
JO - Proceedings of the 2nd International Conference on Pattern Recognition Applications and Methods - Volume 1: ICPRAM,
TI - Arabic Corpus Enhancement using a New Lexicon/Stemming Algorithm
SN - 978-989-8565-41-9
AU - AbdelRaouf A.
AU - A. Higgins C.
AU - Pridmore T.
AU - I. Khalil M.
PY - 2013
SP - 435
EP - 440
DO - 10.5220/0004260704350440