Stylistic Experiments in Information Retrieval
Text, speech and language technologyPublished 1 January 1999
Jussi Karlgren
Citations151
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
A discussion on various experiments to utilize stylistic variation among texts for information retrieval purposes and how this affects decision-making in the field of retrieval.
Abstract
A discussion on various experiments to utilize stylistic variation among texts for information retrieval purposes.
Keywords
Computer Science
C4.5: Programs for Machine Learning
23,665 Citations1992J. R. Quinlan
A complete guide to the C4.5 system as implemented in C for the UNIX environment, which starts from simple core learning methods and shows how they can be elaborated and extended to deal with typical problems such as missing data and over hitting.
Cambridge University Press eBooksVariation across Speech and Writing
5,128 Citations1988Douglas Biber
In Variation Across Speech and Writing, six dimensions of variation are identified through a factor analysis, on the basis of linguistic co-occurence patterns, and the resulting model of variation provides for the description of the distinctive linguistic characteristic of any spoken or written text and enables reconciliation of the contradictory conclusions reached in previous research.
TextTiling: segmenting text into multi-paragraph subtopic passages
1,253 Citations1997Marti A. Hearst
The algorithm is fully implemented and is shown to produce segmentation that corresponds well to human judgments of the subtopic boundaries of 12 texts, which should be useful for many text analysis tasks, including information retrieval and summarization.
Subtopic structuring for full-length document access
330 Citations1993Marti A. Hearst, Christian Plaunt
It is argued that the advent of large volumes of full-length text, as opposed to short texts like abstracts and newswire, should be accompanied by corresponding new approaches to information access and a partition of the text into coherent multi-paragraph units that represent the pattern of subtopics that comprise the text.
Recognizing text genres with simple metrics using discriminant analysis
293 Citations1994Jussi Karlgren, Douglass R. Cutting
A simple method for categorizing texts into pre-determined text genre categories using the statistical standard technique of discriminant analysis is demonstrated with application to the Brown corpus.
The Collection Fusion Problem.
166 Citations1994Ellen M. Voorhees, N. K. Gupta +1 more
This paper examines two collection fusion techniques that use the results of past queries to compute the number of documents to retrieve from each of a set of subcollections such that the total number of retrieved documents is equal to N, the number to be returned to the user.
Munich Personal RePEc Archive (Ludwig Maximilian University of Munich)Natural language information retrieval: TREC-5 report
82 Citations1996Tomek Strzalkowski, Louise Guthrie +7 more
This paper reports on the joint GE/Lockheed Martin/Rutgers/NYU natural language information retrieval project as related to the 5th Text Retrieval Conference (TREC-5), which uses natural language processing techniques to enhance the effectiveness of full-text document retrieval.
Robust text processing in automated information retrieval
32 Citations1994Tomek Strzalkowski
It is demonstrated that the use of syntactic compounds in the representation of database documents as well as in the user queries, coupled with an appropriate term weighting strategy, can considerably improve the effectiveness of retrospective search.
Information Processing & ManagementText windows and phrases differing by discipline, location in document, and syntactic structure
16 Citations1996Robert M. Losee
Different syntactic structures in sublanguages are examined, and their use is considered for discriminating between specific academic disciplines and, more generally, between theory vs practice or knowledge vs applications-oriented documents.
Visualizing stylistic variation
8 Citations2002Jussi Karlgren, T. Straszheim
This work consists of an implementation of a visualization tool for document databases that uses use principal components analysis to combine a quite large number of stylistic items into two most significant dimensions of variation and plot the document space under consideration into a plane.
Cambridge University Press eBooksIntroduction: textual dimensions and relations
2 Citations1988Douglas Biber
