Association, statistical, mathematical and neural approaches for mining breast cancer patterns
Expert Systems with ApplicationsPublished 1 October 1999
Parag C. Pendharkar
Citations110
SJR quartileQ1
SJR score1.85
SNIP2.55
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
The results of the study, based on data obtained from a large medical facility in western Pennsylvania, show that data mining can be a viable tool for breast cancer diagnosis.
Abstract
Using several association and classification approaches to study breast cancer patterns, this study illustrates how these approaches can be used to predict and diagnose the occurrence of breast cancer. The results of the study, based on data obtained from a large medical facility in western Pennsylvania, show that data mining can be a viable tool for breast cancer diagnosis.
Keywords
Computer ScienceBiochemistry, Genetics and Molecular Biology
European Journal of Operational ResearchMeasuring the efficiency of decision making units
28,688 Citations1978A. Charnes, W. W. Cooper +1 more
A nonlinear (nonconvex) programming model provides a new definition of efficiency for use in evaluating activities of not-for-profit entities participating in public programs and methods for objectively determining weights by reference to the observational data for the multiple outputs and multiple inputs that characterize such programs.
Learning Internal Representations by Error Propagation
16,235 Citations1985David E. Rumelhart, Geoffrey E. Hinton +1 more
Machine LearningInduction of Decision Trees
14,815 Citations1986J. R. Quinlan
This paper summarizes an approach to synthesizing decision trees that has been used in a variety of systems, and it describes one such system, ID3, in detail, which is described in detail.
Annals of EugenicsTHE USE OF MULTIPLE MEASUREMENTS IN TAXONOMIC PROBLEMS
14,727 Citations1936Ronald Aylmer Fisher
IEEE Transactions on Knowledge and Data EngineeringData mining: an overview from a database perspective
2,221 Citations1996Ming-Syan Chen⋆, Jiawei Han +1 more
This article provides a survey, from a database researcher's point of view, on the data mining techniques developed recently, a classification of the available data mining Techniques, and a comparative study of such techniques is presented.
Machine LearningThe CN2 Induction Algorithm
2,015 Citations1989Peter Clark, Tim Niblett
A description and empirical evaluation of a new induction system, CN2, designed for the efficient induction of simple, comprehensible production rules in domains where problems of poor description language and/or noise may be present.
Lecture notes in computer scienceEfficient similarity search in sequence databases
1,972 Citations1993Rakesh Agrawal, Christos Faloutsos +1 more
An indexing method for time sequences for processing similarity queries using R * -trees to index the sequences and efficiently answer similarity queries and provides experimental results which show that the method is superior to search based on sequential scanning.
Efficient and Effective Clustering Methods for Spatial Data Mining
1,788 Citations1994Raymond T. Ng, Jiawei Han
The analysis and experiments show that with the assistance of CLAHANS, these two algorithms are very effective and can lead to discoveries that are difficult to find with current spatial data mining algorithms.
Future Generation Computer SystemsMining generalized association rules
1,617 Citations1997Ramakrishnan Srikant, Rakesh Agrawal
A new interest-measure for rules which uses the information in the taxonomy is presented, and given a user-specified “minimum-interest-level”, this measure prunes a large number of redundant rules.
An Efficient Algorithm for Mining Association Rules in Large Databases
1,598 Citations1995Ashoka Savasere, Edward Omiecinski +1 more
This paper presents an efficient algorithm for mining association rules that is fundamentally different from known algorithms and not only reduces the I/O overhead significantly but also has lower CPU overhead for most cases.
Bayesian classification (AutoClass): theory and results
972 Citations1996Peter Cheeseman, John Stutz
It is emphasized that no current unsupervised classi(cid:12)cation system can produce maximally useful results when operated alone, and that it is the interaction between domain experts and the machine searching over the model space, that generates new knowledge.
RePEc: Research Papers in EconomicsNo Free Lunch Theorems for Search
924 Citations1995David H. Wolpert, William G. Macready
It is shown that all algorithms that search for an extremum of a cost function perform exactly the same, when averaged over all possible cost functions, which allows for mathematical benchmarks for assessing a particular search algorithm's performance.
Discovery of Multiple-Level Association Rules from Large Databases
920 Citations1995Jiawei Han, Yongjian Fu
A top-down progressive deepening method is developed for mining multiplelevel association rules from large transaction databases by extension of some existing association rule mining techniques.
National Conference on Artificial IntelligenceThe multi-purpose incremental learning system AQ15 and its testing application to three medical domains
761 Citations1986Ryszard S. Michalski, Igor Mozetič +2 more
The demonstration that by applying the proposed method of cover truncation and analogical matching, called TRUNC, one may drastically decrease the complexity of the knowledge base without affecting its performance accuracy is demonstrated.
Fast Similarity Search in the Presence of Noise, Scaling, and Translation in Time-Series Databases
655 Citations1995Rakesh Agrawal, King-Ip Lin +2 more
A new model of similarity of time sequences is introduced that captures the intuitive notion that two sequences should be considered similar if they have enough non-overlapping time-ordered pairs of subsequences thar are similar.
Knowledge Discovery and Data MiningEfficient algorithms for discovering association rules
630 Citations1994Heikki Mannila, Hannu Toivonen +1 more
An improved algorithm for the problem of mining association rules from large collections of data based on careful combinatorial analysis of the information obtained in previous passes is given, which makes it possible to eliminate unnecessary candidate rules.
New England Journal of MedicineVariability in Radiologists' Interpretations of Mammograms
598 Citations1994Joann G. Elmore, Carolyn K. Wells +3 more
Efforts to improve accuracy and reduce variability in interpretation may increase the effectiveness of mammography in detecting early breast cancers.
An empirical comparison of pattern recognition, neural nets, and machine learning classification methods
470 Citations1989Sholom M. Weiss, Ioannis Kapouleas
For these problems, which have relatively few hypotheses and features, the machine learning procedures for rule induction or tree induction clearly performed best.
IEEE Transactions on Knowledge and Data EngineeringEfficient mining of association rules in distributed databases
345 Citations1996David W. Cheung, Vincent Ng +2 more
An efficient algorithm called DMA (Distributed Mining of Association rules), which generates a small number of candidate sets and requires only O(n) messages for support-count exchange for each candidate set, in distributed databases.
Data mining for path traversal patterns in a web environment
330 Citations2002Ming-Syan Chen⋆, Jong Soo Park +1 more
A new data mining capability which involved mining path traversal patterns in a distributed information providing environment like world-wide-web is explored, where the original sequence of log data is converted into a set of maximal forward references and filter out the effect of some backward references.
Machine LearningSymbolic and Neural Learning Algorithms: An Experimental Comparison
285 Citations1991Jude Shavlik, Raymond J. Mooney +1 more
Experimental results suggest that backpropagation can work significantly better on data sets containing numerical data, and occasionally outperforms the other two systems when given relatively small amounts of training data.
Efficient parallel data mining for association rules
202 Citations1995Jong Soo Park, Ming-Syan Chen⋆ +1 more
An algorithm, called PDM, to conduct parallel data mining for association rules, so designed that the global set of large itemsets can be identified efficiently and the amount of inter-node data exchange required is minimized.
Decision SciencesApplications and Implementation
195 Citations1981N. Freed, Fred Glover
This work proposes an alternative solution to the discriminant problem that requires little more than a minimum familiarity with linear programming and shows promise for eliminating the complexities of conventional statistical approaches without sacrificing the essential power of existing methods.
Lecture notes in computer scienceKnowledge discovery in large spatial databases: Focusing techniques for efficient class identification
160 Citations1995Martin Ester, Hans‐Peter Kriegel +1 more
This paper addresses the task of class identification in spatial databases using clustering techniques using a well-known spatial access method, the R*-tree, and presents several strategies for focusing: selecting representatives from a spatial database, focusing on the relevant clusters and retrieving all objects of a given cluster.
National Conference on Artificial IntelligenceImproving inference through conceptual clustering
127 Citations1987Douglas A. Fisher
COBWEB is presented, a conceptual clustering system that organizes data to maximize inference abilities by capturing attribute inter-correlations at classification tree nodes and generating inferences as a by-product of classification.
An empirical comparison of ID3 and back-propagation
126 Citations1989Douglas Fisher, Kathleen B. McKusick
This work experiments with a system from each paradigm: ID3 and back-propagation, and identifies aspects of each system that may account for distinct performance differences across a variety of domains.
Artificial Intelligence in MedicineFuzzy logic in computer-aided breast cancer diagnosis: analysis of lobulation
112 Citations1997Boris Kovalerchuk, Evangelos Triantaphyllou +2 more
It is shown that fuzzy logic can be an effective tool in dealing with this kind of problem and be used to formalize terms in the ACR Breast Imaging Lexicon.
Induction in noisy domains
112 Citations1987Peter Clark, Tim Niblett
Some of the problems presented by noise are discussed and a top-down induction algorithm for induction in real-world domains is proposed and an experimental comparison of this algorithm with other induction systems is presented using three sets of real- world medical data.
HierarchyScan: a hierarchical similarity search algorithm for databases of long sequences
81 Citations2002Chung‐Sheng Li, Philip S. Yu +1 more
Decision SciencesInductive, Evolutionary, and Neural Computing Techniques for Discrimination: A Comparative Study*
79 Citations1998Siddhartha Bhattacharyya, Parag C. Pendharkar
Simulated data is used to examine how the different learning techniques perform with respect to certain data distribution characteristics and helps relate the findings across a range of discrimination techniques.
Computers & Operations ResearchThe potential use of DEA for credit applicant acceptance systems
72 Citations1996Marvin D. Troutt, Arun Rai +1 more
This paper shows how Data Envelopment Analysis (DEA) may be used to develop an acceptance boundary for use in case based computer systems.
Performance Comparisons Between Backpropagation Networks and Classification Trees on Three Real-World Applications
52 Citations1989Les Atlas, Ronald A. Cole +5 more
There is not enough theoretical basis for the clear-cut superiority of one technique over the other, so a number of empirical tests on three real-world problems in power system load forecasting, power system security prediction, and speaker-independent vowel identification conclude that the multi-layer perceptron performed as well as or better than the trained classification trees.
Decision SciencesRule‐Based Expert Systems and Linear Models: An Empirical Comparison of Learning‐By‐Examples Methods*
28 Citations1992Hyung‐Min Michael Chung, Mark S. Silver
A study that compared linear models derived by logistic regression with rule-based systems produced by two induction algorithms—ID3 and the genetic algorithm performed comparably in modeling the experts at one task, graduate admissions, but differed significantly at a second task, bidder selection.
Knowledge Discovery and Data MiningOptimization and simplification of hierarchical clusterings
27 Citations1995Doug Fisher
This paper evaluates hierarchical redistribution, which appears to be a novel optimization strategy in the clustering literature, and finds that resampling is used to significantly simplify hierarchical clusterings.
