Open AccessProceedings Article
Multiword Expression Identification with Tree Substitution Grammars: A Parsing tour de force with French
Spence Green,Marie-Catherine de Marneffe,John Bauer,Christopher D. Manning +3 more
- 27 Jul 2011
- pp 725-735
TL;DR: It is shown that even the simplest parsing models can effectively identify MWEs of arbitrary length, and that Tree Substitution Grammars achieve the best results.
read more
Abstract: Multiword expressions (MWE), a known nuisance for both linguistics and NLP, blur the lines between syntax and semantics. Previous work on MWE identification has relied primarily on surface statistics, which perform poorly for longer MWEs and cannot model discontinuous expressions. To address these problems, we show that even the simplest parsing models can effectively identify MWEs of arbitrary length, and that Tree Substitution Grammars achieve the best results. Our experiments show a 36.4% F1 absolute improvement for French over an n-gram surface statistics baseline, currently the predominant method for MWE identification. Our models are useful for several NLP tasks in which MWE pre-grouping has improved accuracy.
read more
Chat with Paper
AI Agents for this Paper
Find similar papers on Google Scholar, PubMed and Arxiv
Write a critical review of this paper
Analyze citations of this paper to find unaddressed research gaps
Citations
Multiword expression processing: A survey
Mathieu Constant,Gülşen Eryiğit,Johanna Monti,Lonneke van der Plas,Carlos Ramisch,Mike Rosner,Amalia Todirascu +6 more
TL;DR: A shared understanding of what is meant by “MWE processing” is offered, distinguishing the subtasks of MWE discovery and identification, and the interactions between MWE processing and two use cases: Parsing and machine translation are elucidated.
The PARSEME Shared Task on Automatic Identification of Verbal Multiword Expressions
Agata Savary,Carlos Ramisch,Silvio Ricardo Cordeiro,Federico Sangati,Veronika Vincze,Behrang QasemiZadeh,Marie Candito,Fabienne Cap,Voula Giouli,Ivelina Stoyanova,Antoine Doucet +10 more
- 04 Apr 2017
TL;DR: An initiative meant to bring about substantial progress in understanding, modelling and processing VMWEs is described, to elaborate universal terminologies and annotation guidelines for 18 languages and its main outcome is a multilingual 5-million-word annotated corpus which underlies a shared task on automatic identification of VMwes.
Discriminative Lexical Semantic Segmentation with Gaps: Running the MWE Gamut
TL;DR: A novel representation, evaluation measure, and supervised models are presented for the task of identifying the multiword expressions (MWEs) in a sentence, resulting in a lexical semantic segmentation, enabling efficient sequence tagging algorithms for feature-rich discriminative models.
•Proceedings Article
Comprehensive Annotation of Multiword Expressions in a Social Web Corpus
Nathan Schneider,Spencer Onuffer,Nora Kazour,Emily Danchik,Michael T. Mordowanec,Henrietta Conrad,Noah A. Smith +6 more
- 01 May 2014
TL;DR: This work advocates for a comprehensive annotation approach: proceeding sentence by sentence, the annotators manually group tokens into MWEs according to guidelines that cover a broad range of multiword phenomena.
Edition 1.1 of the PARSEME shared taskon automatic identification of verbal multiword expressions
Carla Parra Escartín,Abigail Walsh +1 more
- 01 Aug 2018
TL;DR: This paper describes the PARSEME Shared Task 1.1 on automatic identification of verbal multiword expressions, and presents the annotation methodology, focusing on changes from last year's shared task.
70
References
The WEKA data mining software: an update
TL;DR: This paper provides an introduction to the WEKA workbench, reviews the history of the project, and, in light of the recent 3.6 stable release, briefly discusses what has been added since the last stable version (Weka 3.4) released in 2003.
Accurate Unlexicalized Parsing
Dan Klein,Christopher D. Manning +1 more
- 07 Jul 2003
TL;DR: It is demonstrated that an unlexicalized PCFG can parse much more accurately than previously shown, by making use of simple, linguistically motivated state splits, which break down false independence assumptions latent in a vanilla treebank grammar.
Mixtures of Dirichlet Processes with Applications to Bayesian Nonparametric Problems
TL;DR: In this article, the conditional distribution of the random measure, given the observations, is no longer that of a simple Dirichlet process, but can be described as being a mixture of DirICHlet processes.
Multiword Expressions: A Pain in the Neck for NLP
Ivan A. Sag,Timothy Baldwin,Francis Bond,Ann Copestake,Dan Flickinger +4 more
- 17 Feb 2002
TL;DR: The various kinds of multiword expressions should be analyzed in distinct ways, including listing "words with spaces", hierarchically organized lexicons, restricted combinatoric rules, lexical selection, "idiomatic constructions" and simple statistical affinity.
Building a Treebank for French
Anne Abeillé,Lionel Clément,François Toussenel +2 more
- 01 Jan 2003
TL;DR: A treebank project for French has annotated a newspaper corpus of 1 Million words with part of speech, inflection, compounds, lemmas and constituency and presents some uses of the corpus.
Related Papers (5)
Anne Abeillé,Lionel Clément,François Toussenel +2 more
- 01 Jan 2003
Christopher D. Manning,Hinrich Schütze +1 more
- 28 May 1999