Breast cancer survivability prediction using labeled, unlabeled, and pseudo-labeled patient data
Journal of the American Medical Informatics AssociationPublished 7 March 2013Open access
Juhyeon Kim, Hyunjung Shin
Citations91
SJR quartileQ1
SJR score2.04
SNIP1.95
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
The concept of tagging virtual labels to unlabeled patient data, that is, 'pseudo-labels,' and treating them as if they were labeled is considered, to compensate for the lack of labeled patient data.
Abstract
Our proposed algorithm, 'SSL Co-training', implements this concept based on SSL. SSL Co-training was tested using the surveillance, epidemiology, and end results database for breast cancer and it delivered a mean accuracy of 76% and a mean area under the curve of 0.81.
Keywords
Computer ScienceMathematicsBiochemistry, Genetics and Molecular Biology
Journal of Applied EcologyAssessing the accuracy of species distribution models: prevalence, kappa and the true skill statistic (TSS)
5,571 Citations2006Omri Allouche, Asaf Tsoar +1 more
The MIT Press eBooksSemi-Supervised Learning
4,308 Citations2006Olivier Chapelle, Schölkopf, B. +1 more
This first comprehensive overview of semi-supervised learning presents state-of-the-art algorithms, a taxonomy of the field, selected applications, benchmark experiments, and perspectives on ongoing and future research.
Minds at UW (University of Wisconsin)Semi-Supervised Learning Literature Survey
3,868 Citations2005Xiaojin Zhu
The study clearly indicates that the common practice of stripwise precommercial thinning is unjustified, and the justification of heavy 'chessboard' thinning (with pruning) depends on whether the potential reduction in rotation length and the improvement in wood quality outweigh the discounted costs of pre-commercial thinning and selection and pruning of crop trees.
Breast Cancer Statistics
1,922 Citations2012Jiemin Ma, Ahmedin Jemal
Breast cancer rates vary largely by race/ethnicity and socioeconomic status (SES), and geographic region, and death rates are higher in African American women than in whites, despite their lower incidence rates.
Artificial Intelligence in MedicinePredicting breast cancer survivability: a comparison of three data mining methods
1,216 Citations2004Dursun Delen, Glenn Walker +1 more
The comparative study of multiple prediction models for breast cancer survivability using a large dataset along with a 10-fold cross-validation provided us with an insight into the relative prediction ability of different data mining methods.
PubMedApplications of machine learning in cancer prediction and prognosis.
892 Citations2007Joseph A. Cruz, David S. Wishart
PLoS BiologySemi-Supervised Methods to Predict Patient Survival from Gene Expression Data
750 Citations2004Eric Bair, Robert Tibshirani
Diagnostic procedures are presented that accurately predict the survival of future patients based on the gene expression profile and survival times of previous patients that have been successfully applied to several publicly available datasets.
Lecture notes in computer scienceRegularization and Semi-supervised Learning on Large Graphs
543 Citations2004Mikhail Belkin, Irina Matveeva +1 more
This work considers the problem of labeling a partially labeled graph, which may arise in a number of situations from survey sampling to information retrieval to pattern recognition in manifold settings.
Cluster Kernels for Semi-Supervised Learning
461 Citations2002Olivier Chapelle, Jason Weston +1 more
A framework to incorporate unlabeled data in kernel classifier, based on the idea that two points in the same cluster are more likely to have the same label is proposed by modifying the eigenspectrum of the kernel matrix.
Semi-supervised time series classification
282 Citations2006Wei Li, Eamonn Keogh
This work proposes a semi-supervised technique for building time series classifiers and shows that special considerations must be made to make them both efficient and effective for the time series domain.
Handbook of Measuring System Design
229 Citations2005Peter H. Sydenham, Richard Thorn +1 more
A high-performance semi-supervised learning method for text chunking
223 Citations2005Rie Kubota Ando, Tong Zhang
A novel semi-supervised method that employs a learning paradigm which is to find "what good classifiers are like" by learning from thousands of automatically generated auxiliary classification problems on unlabeled data, which produces performance higher than the previous best results.
BioinformaticsImproved breast cancer prognosis through the combination of clinical and genetic markers
191 Citations2006Yijun Sun, Steve Goodison +3 more
Studies in computational intelligenceArtificial Neural Networks
134 Citations2012Muhammet Ünal, Ayça Ak +2 more
Intelligent systems reference libraryArtificial Neural Networks
124 Citations2011Crina Groşan, Ajith Abraham
Artificial Neural Networks are inspired by the way biological neural system works, such as the brain process information, and involves adjustments to the synaptic connections that exist between neurons.
European Journal of CancerA computer program for period analysis of cancer patient survival
109 Citations2002Hermann Brenner, Olaf Gefeller +1 more
This paper presents a simple and easy-to-use computer program (SAS macro) that enables one to carry out period analysis of both absolute and relative survival rates with the type of data commonly available in population-based cancer registries.
Neural ComputationNeighborhood Property–Based Pattern Selection for Support Vector Machines
94 Citations2007Hyunjung Shin, Sungzoon Cho
The experimental results provide promising evidence that it is possible to successfully employ the proposed algorithm ahead of SVM training, and to select only the patterns that are likely to be located near the decision boundary.
Machine LearningSemi-supervised model-based document clustering: A comparative study
65 Citations2006Shi Zhong
Deterministic annealing can often significantly improve the performance of semi-supervised clustering and the constrained approach is the best when available labels are complete whereas the feedback-based approach excels when available labeled data are incomplete.
Soft-supervised learning for text classification
63 Citations2008Amarnag Subramanya, Jeff Bilmes
A new graph-based semi-supervised learning (SSL) algorithm that generalizes in a straightforward manner to multi-class problems and significantly outperforms the state-of-the-art on two standard tasks.
BioinformaticsGraph sharpening plus graph integration: a synergy that improves protein functional classification
59 Citations2007Hyunjung Shin, Andreas Martin Lisewski +1 more
Victoria University Research Repository (Victoria University)Breast Cancer Survivability via AdaBoost Algorithms
55 Citations2010Jaree Thongkam, Guandong Xu +2 more
Data pre-processing RELIEF attributes selection, and Modest AdaBoost algorithms, are used to extract knowledge from the breast cancer survival databases in Thailand and computational results showed that Modest Ada boost outperforms Real and Gentle AdaBoosts.
Expert Systems with ApplicationsToward breast cancer survivability prediction models through improving training space
50 Citations2009Jaree Thongkam, Guandong Xu +2 more
Results have indicated that the proposed approach leads to improving the performance of breast cancer survivability prediction models by up to 28.34% due to the improved training data space.
PubMedOn Efficient Large Margin Semisupervised Learning: Method and Theory.
47 Citations2009Junhui Wang, Xiaotong Shen +1 more
FigshareGraph-Based Semi-Supervised Learning as a Generative Model
46 Citations2018Jingrui He, Jaime Carbonell +1 more
Experimental results on various datasets show that the proposed method is superior to existing graph-based semi-supervised learning methods, especially when the labeled subset alone proves insufficient to estimate meaningful class priors.
A graph-based semi-supervised learning for question-answering
28 Citations2009Aslı Çelikyılmaz, Marcus Thint +1 more
A graph-based semi-supervised learning approach for the question-answering (QA) task for ranking candidate sentences using match-scores of textual entailment features as similarity weights between data points and shows improvement in generalization performance over state-of-the-art QA models.
Australasian Data Mining ConferencewFDT: weighted fuzzy decision trees for prognosis of breast cancer survivability
27 Citations2008Umer Khan, Hyunjung Shin +2 more
Performance comparisons suggest that predictions of weighted fuzzy decision trees (wFDT) are more accurate and balanced, than independently applied crisp decision tree classifiers; moreover it has a potential to adapt for significant performance enhancement.
Graph-based Semi-supervised Learning Algorithm for Web Page Classification
20 Citations2006Rong Liu, Jianzhong Zhou +1 more
The preliminary experiments on the WebKB dataset show that the algorithm in this paper can effectively exploit unlabeled data in addition to labeled ones to get higher accuracy of Web page classification.
Expert Systems with ApplicationsGraph sharpening
19 Citations2010Hyunjung Shin, N. Jeremy Hill +2 more
This paper presents an approach to ''sharpening'' in which weights are adjusted to meet an optimization criterion wherever they are directed towards labelled points, and shows that it can improve performance on a number of publicly available bench-mark data sets.
Proceedings - International Conference on Pattern Recognition/Proceedings/International Conference on Pattern RecognitionSemi-supervised method for gene expression data classification with Gaussian fields and harmonic functions
15 Citations2008Yunchao Gong, Chuanliang Chen
The semi-supervised learning algorithms which learning with both labeled and unlabeled data to do classification for microarray data are proposed which holds a much higher classification accuracy than the supervised methods and is much more stable when the labeled examples are very few.
International Joint Conference on Artificial IntelligenceSemi-supervised learning of visual classifiers from web images and text
6 Citations2009Nicholas Morsillo, Christopher Pal +1 more
This paper proposes a semi-supervised model for automatically collecting clean example imagery from the web that includes both visual and textual web data in a unified framework and shows that classifiers trained using the method significantly outperform analogous baseline approaches on the Caltech-256 dataset.
Lecture notes in computer scienceSemi-supervised Learning with Ensemble Learning and Graph Sharpening
3 Citations2008Inae Choi, Hyunjung Shin
The proposed ensemble learning and graph sharpening method replaces the hyperparameter selection procedure to an ensemble network of the committee members trained with various values of hyperparameters and improves the performance of algorithms by removing unhelpful information flow by noise.
