Could machine learning fuel a reproducibility crisis in science?
NaturePublished 26 July 2022
Elizabeth Gibney
Citations71
SJR quartileQ1
SJR score18.29
SNIP10.16
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
Abstract
‘Data leakage’ threatens the reliability of machine-learning use across disciplines, researchers warn. ‘Data leakage’ threatens the reliability of machine-learning use across disciplines, researchers warn.
Keywords
Computer ScienceMedicine
Journal of Computational ScienceTwitter mood predicts the stock market
4,995 Citations2011Johan Bollen, Huina Mao +1 more
This work investigates whether measurements of collective mood states derived from large-scale Twitter feeds are correlated to the value of the Dow Jones Industrial Average (DJIA) over time and indicates that the accuracy of DJIA predictions can be significantly improved by the inclusion of specific public mood dimensions but not others.
The Lancet Digital HealthA comparison of deep learning performance against health-care professionals in detecting diseases from medical imaging: a systematic review and meta-analysis
1,703 Citations2019Xiaoxuan Liu, Livia Faes +15 more
A major finding of the review is that few studies presented externally validated results or compared the performance of deep learning models and health-care professionals using the same sample, which limits reliable interpretation of the reported diagnostic accuracy.
Political AnalysisComparing Random Forest with Logistic Regression for Predicting Class-Imbalanced Civil War Onset Data
241 Citations2015David Muchlinski, David S. Siroky +2 more
This article compares the performance of Random Forests with three versions of logistic regression, and finds that the algorithmic approach provides significantly more accurate predictions of civil war onset in out-of-sample data than any of theLogistic regression models.
