Open AccessProceedings Article
Annotating an Arabic learner corpus for error
Ghazi Abuhakema,Reem Faraj,Anna Feldman,Eileen Fitzpatrick +3 more
- 01 Jan 2008
- pp 1347-1350
TL;DR: This paper describes an ongoing project in which a learner corpus of Arabic is collected, developing a tagset for error annotation and performing Computer-aided Error Analysis (CEA) on the data.
read more
Abstract: This paper describes an ongoing project in which we are collecting a learner corpus of Arabic, developing a tagset for error annotation and performing Computer-aided Error Analysis (CEA) on the data. We adapted the French Interlanguage Database FRIDA tagset (Granger, 2003a) to the data. We chose FRIDA in order to follow a known standard and to see whether the changes needed to move from a French to an Arabic tagset would give us a measure of the distance between the two languages with respect to learner difficulty. The current collection of texts, which is constantly growing, contains intermediate and advanced-level student writings. We describe the need for such corpora, the learner data we have collected and the tagset we have developed. We also describe the error frequency distribution of both proficiency levels and the ongoing work.
read more
Chat with Paper
AI Agents for this Paper
Find similar papers on Google Scholar, PubMed and Arxiv
Write a critical review of this paper
Analyze citations of this paper to find unaddressed research gaps
Citations
Multilingual native language identification
Shervin Malmasi,Mark Dras +1 more
TL;DR: The first comprehensive study of Native Language Identification applied to text written in languages other than English, using data from six languages is presented, finding that most differences between first language groups lie in the ordering of the most basic word categories.
A survey on author profiling, deception, and irony detection for the Arabic language
Paolo Rosso,Francisco Rangel,Irazu Hernández Farías,Leticia Cagnina,Wajdi Zaghouani,Anis Charfi +5 more
TL;DR: The state of the art about some of the main author profiling problems, as well as deception and irony detection, especially focusing on the Arabic language are reviewed.
Arabic learner corpus (ALC) v2: a new written and spoken corpus of Arabic learners
Abdullah Alfaifi,Eric Atwell,Hedaya Ibraheem +2 more
- 01 Jun 2014
TL;DR: The Arabic Learner Corpus is introduced, being developed at Leeds University, and comprises of 282,732 words, collected from learners of Arabic in Saudi Arabia, to provide an open-source of data for some linguistic research areas related to Arabic language learning and teaching.
57
•Proceedings Article
Processing Spontaneous Orthography
Ramy Eskander,Nizar Habash,Owen Rambow,Nadi Tomeh +3 more
- 01 Jun 2013
TL;DR: It is shown that a two-stage process can reduce divergences from this standard by 69%, making subsequent processing of Egyptian Arabic easier.
Correction Annotation for Non-Native Arabic Texts: Guidelines and Corpus
Wajdi Zaghouani,Nizar Habash,Houda Bouamor,Alla Rozovskaya,Behrang Mohit,Abeer Heider,Kemal Oflazer +6 more
- 01 Jun 2015
TL;DR: The overarching goal is to use the annotated corpus to develop components for automatic detection and correction of language errors that can be used to help Standard Arabic learners improve the quality of the Arabic text they produce.
References
The montclair electronic language learner database
Eileen Fitzpatrick,Steve Seegmiller +1 more
- 01 Aug 2001
7
Related Papers (5)
Abdullah Alfaifi,Eric Atwell,Ghazi Abuhakema +2 more
- 19 Sep 2013
Alla Rozovskaya,Dan Roth +1 more
- 05 Jun 2010