A Survey of Text Mining Techniques and Applications
Journal of Emerging Technologies in Web IntelligencePublished 1 August 2009
Vishal Gupta, Gurpreet Singh Lehal
Citations686
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
In this paper, a Survey of Text Mining techniques and applications have been presented.
Abstract
Text Mining has become an important research area. Text Mining is the discovery by computer of new, previously unknown information, by automatically extracting information from different written resources. In this paper, a Survey of Text Mining techniques and applications have been s presented.
Keywords
Computer Science
Journal of the ACMAuthoritative sources in a hyperlinked environment
9,060 Citations1999Jon Kleinberg
This work proposes and test an algorithmic formulation of the notion of authority, based on the relationship between a set of relevant authoritative pages and the set of “hub pages” that join them together in the link structure, and has connections to the eigenvectors of certain matrices associated with the link graph.
Computer NetworksFinding related pages in the World Wide Web
491 Citations1999Jay B. Dean, Monika Henzinger
This paper discusses a different approach to Web searching where the input to the search process is not a set of query terms, but instead is the URL of a page, and the output is aSet of related Web pages.
Communications of the ACMTapping the power of text mining
433 Citations2006Weiguo Fan, Linda Wallace +2 more
Sifting through vast collections of unstructured or semistructured data beyond the reach of data mining tools, text mining tracks information sources, links isolated concepts in distant documents, maps relationships between activities, and helps answer questions.
Visualizing association rules for text mining
145 Citations2003Pak Chung Wong, Paul Whitney +1 more
The results indicate that the design can easily handle hundreds of multiple antecedent association rules in a three-dimensional display with minimum human interaction, low occlusion percentage, and no screen swapping.
News Keyword Extraction for Topic Tracking
116 Citations2008Sungjick Lee, Hanjoon Kim
An unsupervised keyword extraction technique that includes several variants of the conventional TF-IDF model with reasonable heuristics is proposed that can be used for tracking topics over time.
Optimizing Text Summarization Based on Fuzzy Logic
92 Citations2008Farshad Kyoomarsi, Hamid Khosravi +3 more
Comparisons of results show that the proposed new method using fuzzy logic beats most methods which use machine learning as their core.
Automatic Discovery of Similar Words
73 Citations2004Pierre Senellart, Vincent D. Blondel
Three algorithms that extract similar words from a large corpus of documents and consider the specific case of the World Wide Web, and a recent method of automatic synonym extraction in a monolingual dictionary, based on an algorithm that computes similarity measures between vertices in graphs.
Behavior Research MethodsAn introduction to association rule mining: An application in counseling and help-seeking behavior of adolescents
51 Citations2007Dion Hoe‐Lian Goh, Rebecca P. Ang
It is shown that ARM can be used to investigate help-seeking behavior in a sample of secondary school students in Singapore and some guidelines and recommendations for using ARM are presented.
Experiments on Supervised Learning Algorithms for Text Categorization
34 Citations2005Setu Madhavi Namburu, Haiying Tu +2 more
Results show that the performance of PLS is comparable to SVM inText categorization and could be a better candidate for multi-class text categorization.
An approach to sentence-selection-based text summarization
32 Citations2004Fang Chen, Kesong Han +1 more
This paper introduced a newly developed text summarization system that supports both Chinese and English, and describes two new techniques for processing the topic sensitive word feature and the sentence length feature.
Studies in fuzziness and soft computingUnderstanding Text Mining: A Pragmatic Approach
28 Citations2005Sergio Bolasco, Alessio Canzonetti +3 more
The joint analysis of the different case studies has given an adequate picture of TM applications according to the possible types of results that can be obtained, the main specifications of the sectors of applications and the type of functions.
Managing the knowledge contained in electronic documents: a clustering method for text mining
28 Citations2002Salvatore Iiritano, Massimo Ruffolo
This paper describes a prototype of a vertical corporate portal that implements a KDD process for knowledge extraction from unstructured data contained in textual documents.
The research of Web mining
24 Citations2003Lizhen Liu, Junjie Chen +1 more
The research on text mining and usage mining on the Web are introduced and an applicable example of usage mining is given.
Information extraction - a text mining approach
21 Citations2007N. Kanya, S. Geetha
This paper presents a framework for text mining, called DISCOTEX (discovery from text extraction), using a learned information extraction system to transform text into more structured data which is then mined for interesting relationships.
Text to Intelligence: Building and Deploying a Text Mining Solution in the Services Industry for Customer Satisfaction Analysis
20 Citations2008Shantanu Godbole, Shourya Roy
The voice of customer (VoC) and customer satisfaction (C-Sat) analysis settings are described and several unique research challenges brought about by this confluence of text mining and industrial services research are outlined.
Zenodo (CERN European Organization for Nuclear Research)Powerful Tool To Expand Business Intelligence: Text Mining
17 Citations2007Li Gao, Elizabeth Chang +1 more
SPOKEN QA BASED ON A PASSAGE RETRIEVAL ENGINE
13 Citations2006Emilio Sanchis, Davide Buscaldi +3 more
This paper presents a Passage Retrieval-based approach for the development of spoken Question Answering systems in order to study the influence of recognition errors over Question AnSWering systems.
Research on Ontology-Based Text Clustering
11 Citations2008Xiquan Yang, Guo Dina +2 more
Experiments show that the proposed text clustering method based on ontology can improve clustering results performance and to perform better than only single term frequency based method.
Use of NER Information for Improved Topic Tracking
8 Citations2008Xiaowei Wang, Jiang Longbin +2 more
This work proposes a method that using NER information for improved topic tracking using multi-vector model, which extracts proper names, locations and normal terms into distinct sub-vectors of the document representation.
SpeeData: multilingual spoken data entry
7 Citations2002Ulla Ackermann, B. Angelini +6 more
The SpeeData project aims at building a demonstrator that provides a user-friendly interface for spoken data-entry in two languages: Italian and German.
Adapting question answering techniques to the Web
7 Citations2003Jignashu Parikh, M. Narasimha Murty
The paper discusses the issues involved in using Web as a knowledge base for question answering involving simple factual questions and proposes some simple but effective methods to adapt traditional QA methods to deal with efficient these issues and lead to an accurate and efficient question answering system.
TCBPLK: A New Method of Text Categorization
6 Citations2007Jian-Suo Xu
Experimental results confirm that TCBPLK method decreases the number of vector, and enhances the generalization and precision of text categorization.
Research on Application of Improved Text Cluster Algorithm in Intelligent QA System
5 Citations2008Ming Zhao, Jian-Li Wang +1 more
An improved text classify cluster algorithm is proposed that synthetically considers the relationship between keywords and eigenvector representation on base of term frequency statistics, thereby it lessens sensitivity of input sequence and frequency, and effectively raises similarity accuracy of small text and simple sentence.
Research on the Design of the Ontology-Based Automatic Question Answering System
4 Citations2008Bo Wang, Yunqing Li
An ontology-based automatic question answering system model is proposed, at first, build restricted area ontology, then take advantage of the accurate description of concept as well as the definition of the relationship of the concepts, to expand keywords and improve the accuracy and recall rates.
Performing Text Categorization on Manifold
4 Citations2006Guihua Wen, Gan Chen +1 more
This paper presents an approach that performs text categorization on texts manifold with respect to the intrinsic global manifold structure, such as by geodesic distance to measure the distance between two texts.
A Visualization Model for Information Resources Management
4 Citations2008Ning Zhou, WU Jia-xin +2 more
A theoretical method of information visualization system and its corresponding technologies are discussed, which includes constructing strategy of visualization model, environmental configuration, functional module and operation method of prototype system.
