Clustering of scientific fields by integrating text mining and bibliometrics
Experimental DermatologyPublished 23 May 2007
Frizo Janssens
Citations35
SJR quartileQ1
SJR score1.21
SNIP1.07
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
7,12-dimethylbenz(a)anthracene binds to the aryl hydrocarbon receptor and acts as a “ ‘spatially reprogrammable’ ligand” to regulate the activity of the H2O/O2 “ receptors” in the cell.
Abstract
Keywords: 7,12-dimethylbenz(a)anthracene; aryl hydrocarbon receptor; cancer; drug metabolism; pregane X receptor
Keywords
Computer SciencePhysics and Astronomy
Nucleic Acids ResearchGapped BLAST and PSI-BLAST: a new generation of protein database search programs
74,499 Citations1997Stephen F. Altschul
A new criterion for triggering the extension of word hits, combined with a new heuristic for generating gapped alignments, yields a gapped BLAST program that runs at approximately three times the speed of the original.
Nucleic Acids ResearchCLUSTAL W: improving the sensitivity of progressive multiple sequence alignment through sequence weighting, position-specific gap penalties and weight matrix choice
64,940 Citations1994Julie Thompson, Desmond G. Higgins +1 more
The sensitivity of the commonly used progressive multiple sequence alignment method has been greatly improved and modifications are incorporated into a new program, CLUSTAL W, which is freely available.
Molecular Biology and EvolutionThe neighbor-joining method: a new method for reconstructing phylogenetic trees.
60,481 Citations1987Naruya Saitou, M Nei
The neighbor-joining method and Sattath and Tversky's method are shown to be generally better than the other methods for reconstructing phylogenetic trees from evolutionary distance data.
Nucleic Acids ResearchThe Protein Data Bank
39,839 Citations2000Helen M. Berman
The goals of the PDB are described, the systems in place for data deposition and access, how to obtain further information, and near-term plans for the future development of the resource are described.
ScienceEmergence of Scaling in Random Networks
36,366 Citations1999Albert-Ĺaszló Barabási, Réka Albert
A model based on these two ingredients reproduces the observed stationary scale-free distributions, which indicates that the development of large networks is governed by robust self-organizing phenomena that go beyond the particulars of the individual systems.
Journal of Machine Learning ResearchLatent dirichlet allocation
27,049 Citations2003David M. Blei, Andrew Y. Ng +1 more
Journal of Molecular BiologyA simple method for displaying the hydropathic character of a protein
23,075 Citations1982Jack Kyte, Russell F. Doolittle
A computer program that progressively evaluates the hydrophilicity and hydrophobicity of a protein along its amino acid sequence has been devised and its simplicity and its graphic nature make it a very useful tool for the evaluation of protein structures.
BioinformaticsMRBAYES: Bayesian inference of phylogenetic trees
22,173 Citations2001John P. Huelsenbeck, Fredrik Ronquist
The program MRBAYES performs Bayesian inference of phylogeny using a variant of Markov chain Monte Carlo, and an executable is available at http://brahms.rochester.edu/software.html.
RadiologyThe meaning and use of the area under a receiver operating characteristic (ROC) curve.
21,812 Citations1982J A Hanley, Barbara J. McNeil
A representation and interpretation of the area under a receiver operating characteristic (ROC) curve obtained by the "rating" method, or by mathematical predictions based on patient characteristics, is presented and it is shown that in such a setting the area represents the probability that a randomly chosen diseased subject is (correctly) rated or ranked with greater suspicion than a random chosen non-diseased subject.
Journal of Computational and Applied MathematicsSilhouettes: A graphical aid to the interpretation and validation of cluster analysis
20,741 Citations1987Peter J. Rousseeuw
A new graphical display is proposed for partitioning techniques, where each cluster is represented by a so-called silhouette, which is based on the comparison of its tightness and separation, and provides an evaluation of clustering validity.
Reviews of Modern PhysicsStatistical mechanics of complex networks
20,607 Citations2002Réka Albert, Albert-Ĺaszló Barabási
A simple model based on these two principles was able to reproduce the power-law degree distribution of real networks, indicating a heterogeneous topology in which the majority of the nodes have a small degree, but there is a significant fraction of highly connected nodes that play an important role in the connectivity of the network.
BioinformaticsMODELTEST: testing the model of DNA substitution.
20,094 Citations1998David Posada, Keith A. Crandall
The program MODELTEST uses log likelihood scores to establish the model of DNA evolution that best fits the data.
Journal of the American Statistical AssociationHierarchical Grouping to Optimize an Objective Function
19,236 Citations1963Joe H. Ward
SIAM ReviewThe Structure and Function of Complex Networks
18,742 Citations2003Michael Newman
Developments in this field are reviewed, including such concepts as the small-world effect, degree distributions, clustering, network correlations, random graph models, models of network growth and preferential attachment, and dynamical processes taking place on networks.
Proceedings of the National Academy of SciencesCluster analysis and display of genome-wide expression patterns
16,395 Citations1998Michael B. Eisen, Paul T. Spellman +2 more
A system of cluster analysis for genome-wide expression data from DNA microarray hybridization is described that uses standard statistical algorithms to arrange genes according to similarity in pattern of gene expression, finding in the budding yeast Saccharomyces cerevisiae that clustering gene expression data groups together efficiently genes of known similar function.
Computer Networks and ISDN SystemsThe anatomy of a large-scale hypertextual Web search engine
15,828 Citations1998Sergey Brin, Lawrence M. Page
This paper provides an in-depth description of Google, a prototype of a large-scale search engine which makes heavy use of the structure present in hypertext and looks at the problem of how to effectively deal with uncontrolled hypertext collections where anyone can publish anything they want.
Nucleic Acids ResearchA comprehensive set of sequence analysis programs for the VAX
14,423 Citations1984John Devereux, Paul Haeberli +1 more
A group of programs that will interact with each other has been developed for the Digital Equipment Corporation VAX computer using the VMS operating system.
ACM Computing SurveysData clustering
13,184 Citations1999Anil K. Jain, M. Narasimha Murty +1 more
An overview of pattern clustering methods from a statistical pattern recognition perspective is presented, with a goal of providing useful advice and references to fundamental concepts accessible to the broad community of clustering practitioners.
Journal of the American Society for Information ScienceIndexing by latent semantic analysis
12,677 Citations1990Scott Deerwester, Susan Dumais +3 more
The PageRank Citation Ranking : Bringing Order to the Web
12,645 Citations1999Lawrence M. Page, Sergey Brin +2 more
This paper describes PageRank, a mathod for rating Web pages objectively and mechanically, effectively measuring the human interest and attention devoted to them, and shows how to efficiently compute PageRank for large numbers of pages.
Proceedings of the National Academy of SciencesModularity and community structure in networks
12,281 Citations2006M. E. J. Newman
It is shown that the modularity of a network can be expressed in terms of the eigenvectors of a characteristic matrix for the network, which is called modularity matrix, and that this expression leads to a spectral algorithm for community detection that returns results of demonstrably higher quality than competing methods in shorter running times.
ScienceMolecular Classification of Cancer: Class Discovery and Class Prediction by Gene Expression Monitoring
11,622 Citations1999Todd R. Golub, Donna K. Slonim +10 more
A generic approach to cancer classification based on gene expression monitoring by DNA microarrays is described and applied to human acute leukemias as a test case and suggests a general strategy for discovering and predicting cancer classes for other types of cancer, independent of previous biological knowledge.
Modern Information Retrieval
11,544 Citations1999Ricardo Baeza‐Yates, Berthier Ribeiro‐Neto
Proceedings of the National Academy of SciencesAn index to quantify an individual's scientific research output
11,535 Citations2005J. E. Hirsch
The index hbar, defined as the number of papers of an individual that have citation count larger than or equal to the citation count of all coauthors of each paper, is proposed as a useful index to characterize the scientific output of a researcher that takes into account the effect of multiple authorship.
Proceedings of the National Academy of SciencesImproved tools for biological sequence comparison.
11,334 Citations1988William R. Pearson, David J. Lipman
Three computer programs for comparisons of protein and DNA sequences can be used to search sequence data bases, evaluate similarity scores, and identify periodic structures based on local sequence similarity.
Choice Reviews OnlineFinding groups in data: an introduction to cluster analysis
10,614 Citations1991
Journal of Educational StatisticsStatistical Methods for Meta-Analysis
10,479 Citations1988Stanley Wasserman, Larry V. Hedges +1 more
Foundations of statistical natural language processing
9,996 Citations1999Christopher D. Manning, Hinrich Schütze
ScienceQuantitative Monitoring of Gene Expression Patterns with a Complementary DNA Microarray
9,597 Citations1995Mark Schena, Dari Shalon +2 more
A high-capacity system was developed to monitor the expression of many genes in parallel by means of simultaneous, two-color fluorescence hybridization, which enabled detection of rare transcripts in probe mixtures derived from 2 micrograms of total cellular messenger RNA.
Journal of the ACMAuthoritative sources in a hyperlinked environment
9,060 Citations1999Jon Kleinberg
This work proposes and test an algorithmic formulation of the notion of authority, based on the relationship between a set of relevant authoritative pages and the set of “hub pages” that join them together in the link structure, and has connections to the eigenvectors of certain matrices associated with the link graph.
Administrative Science QuarterlyInterorganizational Collaboration and the Locus of Innovation: Networks of Learning in Biotechnology
8,372 Citations1996Walter W. Powell, Kenneth W. Koput +1 more
NatureExploring complex networks
8,339 Citations2001Steven H. Strogatz
This work aims to understand how an enormous network of interacting dynamical systems — be they neurons, power stations or lasers — will behave collectively, given their individual dynamics and coupling architecture.
Journal of Molecular BiologySCOP: A structural classification of proteins database for the investigation of sequences and structures
6,336 Citations1995Alexey G. Murzin, Steven E. Brenner +2 more
This database provides a detailed and comprehensive description of the structural and evolutionary relationships of the proteins of known structure and provides for each entry links to co-ordinates, images of the structure, interactive viewers, sequence data and literature references.
IEEE Transactions on Neural NetworksSurvey of Clustering Algorithms
6,154 Citations2005Rui Xu, D. WunschII
Clustering algorithms for data sets appearing in statistics, computer science, and machine learning are surveyed, and their applications in some benchmark data sets, the traveling salesman problem, and bioinformatics, a new field attracting intensive efforts are illustrated.
Journal of the American Society for Information ScienceCo‐citation in the scientific literature: A new measure of the relationship between two documents
5,183 Citations1973Henry Small
A new form of document coupling called co-citation is defined as the frequency with which two documents are cited together, and clusters of co- cited papers provide a new way to study the specialty structure of science.
Physical Review EFinding community structure in networks using the eigenvectors of matrices
4,801 Citations2006M. E. J. Newman
A modularity matrix plays a role in community detection similar to that played by the graph Laplacian in graph partitioning calculations, and a spectral measure of bipartite structure in networks and a centrality measure that identifies vertices that occupy central positions within the communities to which they belong are proposed.
A Comparative Study on Feature Selection in Text Categorization
4,766 Citations1997Yiming Yang, Jan Pedersen
DF thresholding, the simplest method with the lowest cost in computation, can be reliably used instead of IG or CHI when the computation of these measures are too expensive, and strong correlations between the DF, IG and CHI values of a term are found.
PsychometrikaAnalysis of Individual Differences in Multidimensional Scaling Via an N-way Generalization of “Eckart-Young” Decomposition
4,716 Citations1970J. Douglas Carroll, Jih-Jie Chang
IEEE ExpertData mining and knowledge discovery: making sense out of data
4,643 Citations1996U.M. Feyyad
Find loads of the data mining and knowledge discovery making sense out of data book catalogues in this site as the choice of you visiting this page.
PsychometrikaSome Mathematical Notes on Three-Mode Factor Analysis
4,269 Citations1966Ledyard R Tucker
The model for three-mode factor analysis is discussed in terms of newer applications of mathematical processes including a type of matrix process termed the Kronecker product and the definition of combination variables.
Journal of Molecular BiologyPrediction of complete gene structures in human genomic DNA
4,264 Citations1997Chris Burge, Samuel Karlin
A general probabilistic model of the gene structure of human genomic sequences which incorporates descriptions of the basic transcriptional, translational and splicing signals, as well as length distributions and compositional features of exons, introns and intergenic regions is introduced.
ScienceRapid and Sensitive Protein Similarity Searches
4,090 Citations1985David J. Lipman, William R. Pearson
An algorithm was developed which facilitates the search for similarities between newly determined amino acid sequences and sequences already available in databases and increases sensitivity by giving high scores to those amino acid replacements which occur frequently in evolution.
Journal of Molecular BiologyExpanded sequence dependence of thermodynamic parameters improves prediction of RNA secondary structure
3,813 Citations1999David H. Mathews, Jeffrey Sabina +2 more
An improved dynamic programming algorithm is reported for RNA secondary structure prediction by free energy minimization and experimental constraints, derived from enzymatic and flavin mononucleotide cleavage, improve the accuracy of structure predictions.
PsychometrikaThe Approximation of One Matrix by Another of Lower Rank
3,795 Citations1936Carl Eckart, Gale Young
Nucleic Acids ResearchOptimal computer folding of large RNA sequences using thermodynamics and auxiliary information
3,628 Citations1981Michael Zuker, Patrick Stiegler
A new computer method for folding an RNA molecule that finds a conformation of minimum free energy using published values of stacking and destabilizing energies and is much more efficient, faster, and can fold larger molecules than procedures which have appeared up to now in the biological literature.
Proceedings of the National Academy of SciencesA comprehensive two-hybrid analysis to explore the yeast protein interactome
3,585 Citations2001Takashi Ito, Tomoko Chiba +4 more
The comprehensive analysis using a system to examine two-hybrid interactions in all possible combinations between the budding yeast Saccharomyces cerevisiae is completed and would significantly expand and improve the protein interaction map for the exploration of genome functions that eventually leads to thorough understanding of the cell as a molecular system.
Nucleic Acids ResearchThe SWISS-PROT protein sequence data bank and its supplement TrEMBL in 1999
3,238 Citations1999Amos Bairoch, Rolf Apweiler
Current developments of the database include: cross-references to additional databases; a variety of new documentation files and improvements to TrEMBL, a computer annotated supplement to SWISS-PROT.
Advances In PhysicsEvolution of networks
3,140 Citations2002S. N. Dorogovt︠s︡ev, J. F. F. Mendes
The recent rapid progress in the statistical physics of evolving networks is reviewed, and how growing networks self-organize into scale-free structures is discussed, and the role of the mechanism of preferential linking is investigated.
A Survey of Clustering Data Mining Techniques
2,791 Citations2006Pavel Berkhin
This survey concentrates on clustering algorithms from a data mining perspective as a data modeling technique that provides for concise summaries of the data.
Proceedings of the National Academy of SciencesSearching for intellectual turning points: Progressive knowledge domain visualization
2,723 Citations2004Chaomei Chen
A previously undescribed method progressively visualizing the evolution of a knowledge domain's cocitation network is introduced, demonstrating that a search for intellectual turning points can be narrowed down to visually salient nodes in the visualized network.
Accurate methods for the statistics of surprise and coincidence
2,688 Citations1993Ted Dunning
StructureCATH – a hierarchic classification of protein domain structures
2,657 Citations1997CA Orengo, AD Michie +4 more
Analysis of the structural families generated by CATH reveals the prominent features of protein structure space and a database of well-characterised protein structure families will facilitate the assignment of structure-function/evolution relationships to both known and newly determined protein structures.
Journal of Intelligent Information SystemsOn Clustering Validation Techniques
2,655 Citations2001Maria Halkidi, Yannis Batistakis +1 more
The fundamental concepts of clustering are introduced while it surveys the widely known clustering algorithms in a comparative way and the issues that are under-addressed by the recent algorithms are illustrated.
The EMBO JournalThe relation between the divergence of sequence and structure in proteins.
2,461 Citations1986C. Chothia, Arthur M. Lesk
The root mean square deviation in the positions of the main chain atoms, delta, is related to the fraction of mutated residues, H, by the expression: delta(A) = 0.40 e1.87H.
Physical Review EAnalysis of weighted networks
2,451 Citations2004M. E. J. Newman
It is pointed out that weighted networks can in many cases be analyzed using a simple mapping from a weighted network to an unweighted multigraph, allowing us to apply standard techniques for unweighting graphs to weighted ones as well.
Principles of Data Mining
2,398 Citations2001David J. Hand, Heikki Mannila +1 more
Choice Reviews OnlineSmall worlds: the dynamics of networks between order and randomness
2,393 Citations2000
Foundations of the PARAFAC procedure: Models and conditions for an "explanatory" multi-model factor analysis
2,314 Citations1970Richard A. Harshman
It is shown that an extension of Cattell's principle of rotation to Proportional Profiles (PP) offers a basis for determining explanatory factors for three-way or higher order multi-mode data.
arXiv (Cornell University)Probabilistic Latent Semantic Analysis
2,092 Citations2013Thomas Hofmann
This work proposes a widely applicable generalization of maximum likelihood model fitting by tempered EM, based on a mixture decomposition derived from a latent class model which results in a more principled approach which has a solid foundation in statistics.
Journal of the American Society for Information ScienceA general theory of bibliometric and other cumulative advantage processes
2,063 Citations1976Derek de Solla Price
It is shown that such a stochastic law is governed by the Beta Function, containing only one free parameter, and this is approximated by a skew or hyperbolic distribution of the type that is widespread in bibliometrics and diverse social science phenomena.
Lecture notes in computer scienceOn the Surprising Behavior of Distance Metrics in High Dimensional Space
2,058 Citations2001Charų C. Aggarwal, Alexander Hinneburg +1 more
This paper examines the behavior of the commonly used L k norm and shows that the problem of meaningfulness in high dimensionality is sensitive to the value of k, which means that the Manhattan distance metric is consistently more preferable than the Euclidean distance metric for high dimensional data mining applications.
ScienceOn Finding All Suboptimal Foldings of an RNA Molecule
2,023 Citations1989Michael Zuker
The mathematical problem of determining how well defined a minimum energy folding is can now be solved and all predicted base pairs that can participate in suboptimal structures may be displayed and analyzed graphically.
ScientometricsCo-word analysis as a tool for describing the network of interactions between basic and technological research: The case of polymer chemsitry
1,944 Citations1991Michel Callon, J Courtial +1 more
The co-word analysis techniques developed in this paper should help to build a bridge between research in scientometrics and work underway to better understand the economics of innovation.
Social Science InformationFrom translations to problematic networks: An introduction to co-word analysis
1,777 Citations1983Jean‐Max Noyer, Jean-Pierre Courtial +2 more
Co-clustering documents and words using bipartite spectral graph partitioning
1,696 Citations2001Inderjit S. Dhillon
A new spectral co-clustering algorithm is used that uses the second left and right singular vectors of an appropriately scaled word-document matrix to yield good bipartitionings and it can be shown that the singular vectors solve a real relaxation to the NP-complete graph bipartitionsing problem.
Technology and CultureCitation Indexing-Its Theory and Application in Science, Technology, and Humanities
1,669 Citations1980Jack Goodwin, Eugene Garfield
Citation indexing-its theory and application in science, technology, and humanities, Citation indexing (Citation Indexing) (CIFS), مرکز فناوری اطلاعات (Citations Indexing),
NatureQuantifying social group evolution
1,657 Citations2007Gergely Palla, Albert-Ĺaszló Barabási +1 more
The focus is on networks capturing the collaboration between scientists and the calls between mobile phone users, and it is found that large groups persist for longer if they are capable of dynamically altering their membership, suggesting that an ability to change the group composition results in better adaptability.
Proceedings of the National Academy of SciencesRapid similarity searches of nucleic acid and protein data banks.
1,577 Citations1983W. John Wilbur, David J. Lipman
An algorithm for the global comparison of sequences based on matching k-tuples of sequence elements for a fixed k results in substantial reduction in the time required to search a data bank when compared with prior techniques of similarity analysis, with minimal loss in sensitivity.
Computational Statistics & Data AnalysisAlgorithms and applications for approximate nonnegative matrix factorization
1,534 Citations2006Michael W. Berry, Murray Browne +3 more
The development and use of low-rank approximate nonnegative matrix factorization algorithms for feature extraction and identification in the fields of text mining and spectral data analysis and the interpretability of NMF outputs in specific contexts are provided.
The European Physical Journal BHow popular is your paper? An empirical study of the citation distribution
1,508 Citations1998S. Redner
Proceedings of the National Academy of SciencesImproved free-energy parameters for predictions of RNA duplex stability.
1,493 Citations1986Susan M. Freier, Ryszard Kierzek +5 more
These parameters predict melting temperatures of most oligonucleotide duplexes within 5 degrees C, about as good as can be expected from the nearest-neighbor model.
SIAM ReviewUsing Linear Algebra for Intelligent Information Retrieval
1,484 Citations1995Michael W. Berry, Susan Dumais +1 more
A lexical match between words in users’ requests and those in or assigned to documents in a database helps retrieve textual materials from scientific databases.
ACM SIGKDD Explorations NewsletterSubspace clustering for high dimensional data
1,342 Citations2004Lance Parsons, Ehtesham Haque +1 more
A survey of the various subspace clustering algorithms along with a hierarchy organizing the algorithms by their defining characteristics is presented, comparing the two main approaches using empirical scalability and accuracy tests and discussing some potential applications where sub space clustering could be particularly useful.
Machine LearningConcept Decompositions for Large Sparse Text Data Using Clustering
1,297 Citations2001Inderjit S. Dhillon, Dharmendra S. Modha
The concept vectors produced by the spherical k-means algorithm constitute a powerful sparse and localized “basis” for text data sets and are localized in the word space, are sparse, and tend towards orthonormality.
NatureA new approach to protein fold recognition
1,246 Citations1992D. T. Jones, W. R. Taylort +1 more
A new approach to fold recognition, whereby sequences are fitted directly onto the backbone coordinates of known protein structures, using a given sequence as a guide for the matching of sequences to backbone coordinates.
Annual Review of Information Science and TechnologyScholarly communication and bibliometrics
1,124 Citations2002Christine L. Borgman, Jonathan Furner
The Future of Bibliometrics and its Applications: A Case Study of Agricultural Research Within the European Community Core Journals of the Rapidly Changing Research Front of "Superconductivity" is reviewed.
Genome ResearchA DNA microarray system for analyzing complex DNA samples using two-color fluorescent probe hybridization.
1,120 Citations1996Dari Shalon, Steven J. Smith +1 more
This work describes a general experimental approach, using microscopic arrays of DNA fragments on glass substrates for differential hybridization analysis of fluorescently labeled DNA samples, and demonstrates the utility of DNA microarrays in the analysis of complex DNA samples.
Nucleic Acids ResearchThe PROSITE database, its status in 1999
1,100 Citations1999Kay Hofmann, Philip Bucher +2 more
Europhysics Letters (EPL)Competition and multiscaling in evolving networks
1,056 Citations2001Ginestra Bianconi, Albert-Ĺaszló Barabási
This work finds that competition for links translates into multiscaling, i.e. a fitness-dependent dynamic exponent, allowing fitter nodes to overcome the more connected but less fit ones.
Research PolicyNetwork structure, self-organization, and the growth of international collaboration in science
1,043 Citations2005Caroline S. Wagner, Loet Leydesdorff
Using tools from network analysis, the paper shows that the growth of international co-authorships can be explained based on the organising principle of preferential attachment, although the attachment mechanism deviates from an ideal power-law.
Data Mining and Knowledge DiscoveryOn Bias, Variance, 0/1—Loss, and the Curse-of-Dimensionality
1,029 Citations1997Jerome H. Friedman
This work candramatically mitigate the effect of the bias associated with some simpleestimators like “naive” Bayes, and the bias induced by the curse-of-dimensionality on nearest-neighbor procedures.
Journal of Molecular BiologyHow RNA folds
989 Citations1999Ignacio Tinoco, Carlos Bustamante
A folding algorithm to predict the structure of an RNA from its sequence is suggested, but to solve the RNA folding problem one needs thermodynamic data on tertiary structure interactions, and identification and characterization of metal-ion binding sites.
Springer series in statisticsModern Multidimensional Scaling
817 Citations1997Ingwer Borg, Patrick J. F. Groenen
Semi-supervised Clustering by Seeding
804 Citations2002Sugato Basu, Arindam Banerjee +1 more
Less is More: Active Learning with Support Vector Machines
775 Citations2000Greg Schohn, David Cohn
A simple active learning heuristic is described which greatly enhances the generalization behavior of support vector machines (SVMs) on several practical document classification tasks and frequently does so in less time than the naive approach of training on all available data.
Journal of Molecular BiologyA dynamic programming algorithm for RNA structure prediction including pseudoknots 1 1Edited by I. Tinoco
757 Citations1999Elena Rivas, Sean R. Eddy
This is the first algorithm to be able to fold optimal (minimum energy) pseudoknotted RNAs with the accepted RNA thermodynamic model and a useful graphical representation borrowed from quantum field theory is adopted.
Journal of Molecular BiologyExtracting regulatory sites from the upstream region of yeast genes by computational analysis of oligonucleotide frequencies 1 1Edited by G. von Heijne
742 Citations1998Jacques van Helden, Bruno André +1 more
A simple and fast method allowing the isolation of DNA binding sites for transcription factors from families of coregulated genes, with results illustrated in Saccharomyces cerevisiae.
Lecture notes in computer sciencePajek— Analysis and Visualization of Large Networks
717 Citations2002Vladimir Batagelj, Andrej Mrvar
Computer Networks and ISDN SystemsAutomatic resource compilation by analyzing hyperlink structure and associated text
700 Citations1998Soumen Chakrabarti, Byron Dom +4 more
An evaluation of ARC suggests that the resources found by ARC frequently fare almost as well as, and sometimes better than, lists of resources that are manually compiled or classified into a topic.
ScientometricsNational characteristics in international scientific co-authorship relations
699 Citations2001Wolfgang Glänzel
As expected, international co-authorship, on an average, results in publications with higher citation rates than purely domestic papers, however, the influence of international collaboration on the national citation impact varies considerably between the countries (and within one individual country between fields).
Basic & Clinical Biostatistics
679 Citations2004Beth Dawson, Robert G. Trapp
This book discusses methods of Evidence-Based Medicine and Decision Analysis based on evidence-based medicine and decision analysis, as well as statistical methods for multiple Variables, used in medical research.
Proceedings of the National Academy of SciencesLocating protein-coding regions in human DNA sequences by a multiple sensor-neural network approach.
658 Citations1991Edward C. Uberbacher, Richard Mural
This work describes a reliable computational approach for locating protein-coding portions of genes in anonymous DNA sequence using a set of sensor algorithms and a neural network to localize the coding regions.
Active learning using pre-clustering
653 Citations2004Hieu T. Nguyen, A.W.M. Smeulders
A formal framework that incorporates clustering into active learning with two-class active learning that allows to select the most representative samples as well as to avoid repeatedly labeling samples in the same cluster.
SIAM Journal on Numerical AnalysisGeneralizing the Singular Value Decomposition
648 Citations1976Charles F. Van Loan
No CategoryLucene in action
611 Citations2005Otis Gospodnetić, Erik Hatcher +1 more
Lucene in Action describes what Lucene is and how it works and most importantly how it can be used in a variety of real-world use cases, such at Nutch, an open-source project designed to index the internet very much like Google.
ScientometricsComparison of the Hirsch-index with standard bibliometric indicators and with peer judgment for 147 chemistry research groups
609 Citations2006Anthony F. J. van Raan
Characteristics of the statistical correlation between the Hirsch (h-) index and several standard bibliometric indicators are presented, as well as with the results of peer review judgment, which show that the h-index and the bibliometry ‘crown indicator’ both relate in a quite comparable way with peer judgments.
Information Processing & ManagementDocument clustering using nonnegative matrix factorization
608 Citations2005Farial Shahnaz, Michael W. Berry +2 more
A methodology for automatically identifying and clustering semantic features or topics in a heterogeneous text collection using a low rank nonnegative matrix factorization algorithm to retain natural data nonnegativity, thereby eliminating the need for subtractive basis vector and encoding calculations present in other techniques.
…
