Text mining without document context
Information Processing & ManagementPublished 1 June 2006Open access
Éric SanJuan, Fidelia Ibekwe-Sanjuan
Citations67
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
This out-of-context clustering task led us to adapt multi-word term representation for statistical methods and also to refine an existing cluster evaluation metric, the editing distance, in order to evaluate the methods.
Abstract
International audience
Keywords
Computer Science
Proceedings of the National Academy of SciencesCluster analysis and display of genome-wide expression patterns
16,395 Citations1998Michael B. Eisen, Paul T. Spellman +2 more
A system of cluster analysis for genome-wide expression data from DNA microarray hybridization is described that uses standard statistical algorithms to arrange genes according to similarity in pattern of gene expression, finding in the budding yeast Saccharomyces cerevisiae that clustering gene expression data groups together efficiently genes of known similar function.
Language<b>WordNet: An electronic lexical database</b> . Ed. by Christiane Fellbaum. Cambridge, MA: MIT Press, 1998. Pp. xxii, 423.
11,687 Citations2000Adam Kilgarriff
The lexical database: nouns in WordNet, George A. Miller modifiers in WordNet, Katherine J. Miller a semantic network of English verbs, and applications of WordNet: building semantic concordances are presented.
Choice Reviews OnlineFinding groups in data: an introduction to cluster analysis
10,614 Citations1991
Journal of ClassificationComparing partitions
7,714 Citations1985Lawrence J. Hubert, Phipps Arabie
A measure based on the comparison of object triples having the advantage of a probabilistic interpretation in addition to being corrected for chance is proposed and bounded between ±1.5 and ±2.5.
PsychometrikaAn Examination of Procedures for Determining the Number of Clusters in a Data Set
3,870 Citations1985Glenn W. Milligan, Martha C. Cooper
A Monte Carlo evaluation of 30 procedures for determining the number of clusters was conducted on artificial data sets which contained either 2, 3, 4, or 5 distinct nonoverlapping clusters to provide a variety of clustering solutions.
Word association norms, mutual information, and lexicography
3,674 Citations1990Kenneth Church, Patrick Hanks
Accurate methods for the statistics of surprise and coincidence
2,688 Citations1993Ted Dunning
ComputerChameleon: hierarchical clustering using dynamic modeling
2,102 Citations1999George Karypis, Eui-Hong Han +1 more
Chameleon's key feature is that it accounts for both interconnectivity and closeness in identifying the most similar pair of clusters, which is important for dealing with highly variable clusters.
ScientometricsCo-word analysis as a tool for describing the network of interactions between basic and technological research: The case of polymer chemsitry
1,944 Citations1991Michel Callon, J Courtial +1 more
The co-word analysis techniques developed in this paper should help to build a bridge between research in scientometrics and work underway to better understand the economics of innovation.
BioinformaticsPrincipal component analysis for clustering gene expression data
1,294 Citations2001Ka Yee Yeung, Walter L. Ruzzo
The empirical study showed that clustering with the PCs instead of the original variables does not necessarily improve, and often degrades, cluster quality, and would not recommend PCA before clustering except in special circumstances.
IEEE Transactions on Knowledge and Data EngineeringCLARANS: a method for clustering objects for spatial data mining
1,184 Citations2002Raymond T. Ng, Jiawei Han
A new clustering method is proposed, called CLARANS, whose aim is to identify spatial structures that may be present in the data, and two spatial data mining algorithms that aim to discover relationships between spatial and nonspatial attributes are developed.
Scatter/Gather: a cluster-based approach to browsing large document collections
945 Citations1992Douglass R. Cutting, David R. Karger +2 more
This work presents a document browsing technique that employs document clustering as its primary operation, and presents fast (linear time) clustering algorithms which support this interactive browsing paradigm.
Retrieving collocations from text: Xtract
799 Citations1993Frank Smadja
A set of techniques based on statistical methods for retrieving and identifying collocations from large textual corpora, based on some original filtering methods that allow the production of richer and higher-precision output are described.
Introduction to the bio-entity recognition task at JNLPBA
579 Citations2004Jin-Dong Kim, Tomoko Ohta +3 more
The JNLPBA shared task of bio-entity recognition using an extended version of the GENIA version 3 named entity corpus of MEDLINE abstracts is described and a general discussion of the approaches taken by participating systems is presented.
Journal of the Royal Statistical Society Series B (Statistical Methodology)Estimating the Number of Clusters in a Data Set Via the Gap Statistic
542 Citations2001Robert Tibshirani, Guenther Walther +1 more
A stability based method for discovering structure in clustered data
540 Citations2001Asa Ben‐Hur, André Elisseeff +1 more
The method can be used with any clustering algorithm and provides a means of rationally defining an optimum number of clusters, and can also detect the lack of structure in data.
Multivariate Behavioral ResearchA Study of the Comparability of External Criteria for Hierarchical Cluster Analysis
488 Citations1986Glenn W. Milligan, Martha C. Cooper
The results of the study indicated that the Hubert and Arabie adjusted Rank index was best suited to the task of comparison across hierarchy levels.
Spotting and Discovering Terms through Natural Language Processing
239 Citations1997Christian Jacquemin
Christian Jacquemin shows how the power of natural language processing (NLP) can be used to advance text indexing and information retrieval (IR) and provides a comprehensive account of the method and implementation of this innovative retrieval technique for text processing.
Computational LinguisticsRetrieving collocations from text
211 Citations1993SmadjaFrank
Natural languages are full of collocations, recurrent combinations of words that co-occur more often than expected by chance and that correspond to arbitrary word usages.
Journal of the American Society for Information ScienceMapping of science by combined co-citation and word analysis. II: Dynamical aspects
199 Citations1991Robert R. Braam, Henk F. Moed +1 more
It is shown that, over a period of 10 years, continuity in intellectual base was at a lower level than continuity in topics of current research, which indicates that a series of interesting new contributions are made in course of time, without vast alteration in general topics of research.
Document clustering with committees
174 Citations2002Patrick Pantel, Dekang Lin
A new evaluation methodology that is based on the editing distance between output clusters and manually constructed classes (the answer key) is presented, which is more intuitive and easier to interpret than previous evaluation measures.
Pattern RecognitionBootstrap technique in cluster analysis
147 Citations1987Anil K. Jain, J.V. Moreau
A method to estimate the number of clusters in a data set E, using the bootstrap technique, which involves the generation of several “fake” data sets by sampling patterns with replacement in E (bootstrapping).
Information Processing & ManagementCombining full text and bibliometric information in mapping scientific disciplines
147 Citations2005Patrick Glenisson, Wolfgang Glänzel +2 more
Full text analysis and traditional bibliometric methods are serially combined to improve the efficiency of the two individual methods and confirm the main results of the pilot study that such hybrid methodology can be applied to both research evaluation and information retrieval.
ScientometricsDevelopment of a method for detection and trend analysis of research fronts built by lexical or cocitation analysis
124 Citations1994Michel Zitt, Elise Bassecoulard
The method presented here combines structural analysis and trend detection, by operating on a “thick-slice” of time, starting from co-citation or co-word analysis (applications of either type have already been carried on).
Clustering by committee
88 Citations2003Dekang Lin, Patrick Pantel
This work proposes a general-purpose clustering algorithm called CBC (Clustering By Committee) from which it will organize documents according to topics as well as discover concepts and word senses, and presents an evaluation methodology for measuring the precision and recall of discovered senses.
Journal of ClassificationModel-Based Clustering for Image Segmentation and Large Datasets via Sampling
78 Citations2004Ron Wehrens, L.M.C. Buydens +2 more
These experiments suggest that a stable method with better performance can be obtained with two straightforward modifications to the simple sampling method: several tentative models are identified from the sample instead of just one, and several EM steps are used rather than just one E step to classify the full data set.
Studies in classification, data analysis, and knowledge organizationComparison of Distance Indices Between Partitions
59 Citations2006Lucile Denœud, Alain Guénoche
Five classical distance indices on P n, the set of partitions on n elements, are compared and the distributions of the five index values between P and the elements of P k(P).
Lecture notes in computer scienceContextual Document Clustering
30 Citations2004Vladimir Dobrynin, David Patterson +1 more
A novel algorithm based on distributional clustering where subject related words, which have a narrow context, are identified to form meta-tags for that subject to form thematic clusters of documents is presented.
Computer Speech & LanguageA symbolic approach to automatic multiword term structuring
26 Citations2005Éric SanJuan, James Dowdall +2 more
This paper explores how this three-level term structuring brings to light the knowledge structures from a corpus of genomics and compares the mapping of the domain topics against a hand-built ontology (the GENIA ontology) and ways of integrating the results into a Q-A system are discussed.
Terminology International Journal of Theoretical and Applied Issues in Specialized CommunicationUsing distributional similarity to organise biomedical terminology
24 Citations2005Julie Weeds, James Dowdall +3 more
It is shown that distributional similarity can be used to predict semantic type with a good degree of accuracy and is carried out using the Pro3Gres parser.
E-LIS Repository (University of Naples Federico II)Mining Textual Data through Term Variant Clustering: the TermWatch system
24 Citations2004Fidelia Ibekwe-Sanjuan, Éric SanJuan
An experiment on unsupervised textmining is performed on a corpus of scientific titles and abstracts from 16 prominent IR journals, and preliminary results showed that TermWatch was able to capture low occurring phenomena which the usual clustering methods based on co-occurrence may not highlight.
Terminological variation, a means of identifying research topics from texts
23 Citations1998Fidelia Ibekwe-Sanjuan
After extracting terms from a corpus of titles and abstracts in English, syntactic variation relations are identified amongst them in order to detect research topics and a clustering method has been built to compute such classes of research topics.
Journal of the American Society for Information Science and TechnologyThe clustering power of low frequency words in academic Webs
19 Citations2005Liz Price, Mike Thelwall
Results for the Australian and New Zealand academic Web spaces indicate thatLow frequency words are useful for clustering academic Web sites along subject lines; removing low frequency words results in sites becoming, on average, less dissimilar to sites from other subjects.
Proceedings of the 17th international conference on Computational linguistics -Terminological variation, a means of identifying research topics from texts
13 Citations1998Fidelia Ibekwe-Sanjuan
Model-Based Clustering for Image Segmentation and Large Datasets Via Sampling
9 Citations2003Ron Wehrens, L.M.C. Buydens +2 more
Classification Et Désarticulation De Graphes De Termes
8 Citations2004Anne Berry, Bangaly Kaba +3 more
Underlying graph-theoretical aspects with algorithm CPCL presented in TermWatch and in defining a decomposition algorithm for these classes of hierarchical classification are developed.
