login

Guest Editors' Introduction: Machine Learning and Natural Language

Machine LearningPublished 1 February 1999Open access
Claire Cardie, Raymond J. Mooney
Citations37
SJR quartileQ1
SJR score1.15
SNIP2.14
View PDF

TL;DR

This special issue attempts to bridge the divide between machine learning and empirical NLP by assembling an interesting variety of recent research papers on various aspects of natural language learning and presenting them to the readers of Machine Learning.

Abstract

The application of machine learning techniques to natural language processing (NLP) has increased dramatically in recent years under the name of "corpus-based," "statistical," or "empirical" methods.However, most of this research has been conducted outside the traditional machine learning research community.This special issue attempts to bridge this divide by assembling an interesting variety of recent research papers on various aspects of natural language learning -many from authors who do not generally publish in the traditional machine learning literature -and presenting them to the readers of Machine Learning.In the last five to ten years there has been a dramatic shift in computational linguistics from manually constructing grammars and knowledge bases to partially or totally automating this process by using statistical learning methods trained on large annotated or unannotated natural language corpora.The success of statistical methods in speech recognition (Stolcke, 1997;Jelinek, 1998) has been particularly influential in motivating the application of similar methods to other aspects of natural language processing.There is now a variety of work on applying learning methods to almost all other aspects of language processing as well (Brill & Mooney, 1997), including morphological and syntactic analysis (Charniak, 1997), semantic disambiguation and interpretation (Ng & Zelle, 1997), discourse processing and information extraction (Cardie, 1997), and machine translation (Knight, 1997).Some concrete publication statistics clearly illustrate the extent of the revolution in natural language research.According to data recently collected by Hirschberg (1998), a full 63.5% of the papers in the Proceedings of the Annual Meeting of the Association for Computational Linguistics and 47.4% of the papers in the journal Computational Linguistics concerned corpus-based research in 1997.For comparison, 1983 was the last year in which there were no such papers and the percentages in 1990 were still only 12.8% and 15.4%.Nevertheless, traditional machine learning research in artificial intelligence has had limited influence on recent research in computational linguistics.This is unfortunate since we believe that machine learning and empirical NLP have much to offer each other and that increased interaction and exchange of ideas would greatly benefit both areas.Most current learning research in NLP employs particular statistical techniques inspired by research in speech recognition, such as hidden Markov models (HMMs) and probabilistic context-free grammars (PCFGs).A variety of other learning methods including decision tree and rule induction, neural networks, instance-based methods, Bayesian network learning, inductive logic programming, explanation-based learning, and genetic algorithms can also be applied to natural language problems and can have significant advantages in particular applica-

Keywords

Computer Science