Automatic Detection of Text Genre
arXiv (Cornell University)Published 8 July 1997Open access
Brett Kessler, Geoffrey Nunberg, Hinrich Schuetze
Citations206
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
Abstract
As the text databases available to users become larger and more heterogeneous, genre becomes increasingly important for computational linguistics as a complement to topical and structural principles of classification. We propose a theory of genres as bundles of facets, which correlate with various surface cues, and argue that genre detection based on surface cues is as successful as detection based on deeper structural properties.
Keywords
Computer Science
Cambridge University Press eBooksVariation across Speech and Writing
5,128 Citations1988Douglas Biber
In Variation Across Speech and Writing, six dimensions of variation are identified through a factor analysis, on the basis of linguistic co-occurence patterns, and the resulting model of variation provides for the description of the distinctive linguistic characteristic of any spoken or written text and enables reconciliation of the contradictory conclusions reached in previous research.
Language<b>Dimensions of register variation:</b> A cross-linguistic comparison. By Douglas Biber. Cambridge: Cambridge University Press, 1995. Pp. xvi, 428.
721 Citations1999Thomas E. Nunnally
LanguageSpoken and Written Textual Dimensions in English: Resolving the Contradictory Findings
548 Citations1986Douglas Biber
Backpropagation: the basic theory
298 Citations1995David E. Rumelhart, Richard Durbin +2 more
Since the publication of the PDP volumes in 1986, learning by backpropagation has become the most popular method of training neural networks because of the underlying simplicity and relative power of the algorithm.
Recognizing text genres with simple metrics using discriminant analysis
293 Citations1994Jussi Karlgren, Douglass R. Cutting
A simple method for categorizing texts into pre-determined text genre categories using the statistical standard technique of discriminant analysis is demonstrated with application to the Brown corpus.
Americanae (AECID Library)The linguistics of punctuation
213 Citations1990Geoffrey Nunberg
This chapter discusses the grammar and functions of text-categories in English, and the limitations of contrastive approaches to grammar and syntax.
Computers and the HumanitiesThe multi-dimensional approach to linguistic analyses of genre variation: An overview of methodology and findings
129 Citations1992Douglas Douglas
The present paper summarizes the major methods and results of the multi-dimensional approach to genre variation, which combines the resources of computational tools, large text corpora, and multivariate statistical tools (such as factor analysis and cluster analysis).
