Cost-Sensitive Feature Extraction and Selection in Genre Classification
LDV-Forum/Journal for language technology and computational linguisticsPublished 1 July 2009Open access
Ryan Levering, Michal Cutler
Citations2
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
This paper presents a framework for the extraction of various Web-specific feature groups from distinct data models based on a tree of potentials models and the transformations that create them and provides an algorithm for forming wrapper-based feature selection on this tree.
Abstract
Cost-Sensitive Feature Extraction and Selection in Genre Classifica-
Keywords
Computer Science
Elsevier eBooksData Mining: Practical Machine Learning Tools and Techniques
25,718 Citations2011Ian H. Witten, Eibe Frank
Communications of the ACMMapReduce
18,538 Citations2008Jay B. Dean, Sanjay Ghemawat
This presentation explains how the underlying runtime system automatically parallelizes the computation across large-scale clusters of machines, handles machine failures, and schedules inter-machine communication to make efficient use of the network and disks.
Text genre detection using common word frequencies
166 Citations2000Efstathios Stamatatos, Nikos Fakotakis +1 more
It is shown that the most frequent words of the British National Corpus, representing the most Frequence of the written English language, are more reliable discriminators of text genre in comparison to the most frequently spoken words in a training corpus.
ePrints Soton (University of Southampton)Composite Kernels for Hypertext Categorisation
145 Citations2001Thorsten Joachims, Nello Cristianini +1 more
Recognition of common areas in a Web page using visual information: a possible application in a page classification
106 Citations2003Miloš Kovačević, Michelangelo Diligenti +2 more
A new, hierarchical representation that includes browser screen coordinates for every HTML object in a page is proposed that shows that a Naive Bayes classifier clearly outperforms the same classifier using only information about the content of documents.
Automatic Identification of Genre in Web Pages
55 Citations2011Marina Santini
It is argued that automatic identification of genre in web pages needs more flexible genre classification schemes, and experiments are described that support this claim.
Genre as interface metaphor: exploiting form and function in digital environments
48 Citations2003Elaine G. Toms, D. Grant Campbell
The findings indicate that the form attributes of a genre play a significant role in the identification of corresponding documents, and suggest that genre can potentially serve as an interface metaphor.
Genre Classification of Web Pages: User Study and Feasibility Analysis
45 Citations2004Sven Meyer zu Eissen, Benno Stein
Lecture notes in computer scienceOn Feature Selection with Measurement Cost and Grouped Features
38 Citations2002Pavel Paclı́k, Robert P. W. Duin +2 more
It is shown, that employing grouping improves the performance significantly for low measurement costs and an application where limiting the computation time is a very important topic: the segmentation of backscatter images in product analysis is discussed.
Iterative Information Retrieval Using Fast Clustering and Usage-Specific Genres
27 Citations1999Jussi Karlgren, Ivan Bretan +3 more
This paper describes how collection specific empirically defined stylistics based genre prediction can be brought together together with rapid topical clustering to build an interactive information retrieval interface with multi-dimensional presentation of search results.
Using Visual Features for Fine-Grained Genre Classification of Web Pages
22 Citations2008Ryan Levering, Michal Cutler +1 more
It is confirmed that using HTML features and particularly URL address features can improve classification beyond using textual features alone, and it is shown that adding visual features can be useful for further improving fine-grained genre classification.
