Hazy
Communications of the ACMPublished 22 February 2013
Arun Kumar, Feng Niu, Christopher Ré
Citations74
SJR quartileQ1
SJR score1.15
SNIP3.34
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
Abstract
Racing to unleash the full potential of big data with the latest statistical and machine-learning techniques.
Keywords
Computer Science
Cambridge University Press eBooksConvex Optimization
31,266 Citations2004Stephen Boyd, Lieven Vandenberghe
The Elements of Statistical Learning: Data Mining, Inference, and Prediction
19,345 Citations2013Trevor Hastie, Robert Tibshirani +1 more
Machine LearningMarkov logic networks
2,678 Citations2006Matthew Richardson, Pedro Domingos
Experiments with a real-world database and knowledge base in a university domain illustrate the promise of this approach to combining first-order logic and probabilistic graphical models in a single representation.
Data Mining and Knowledge DiscoveryData Cube: A Relational Aggregation Operator Generalizing Group-By, Cross-Tab, and Sub-Totals
2,134 Citations1997Jim Gray, Surajit Chaudhuri +6 more
This paper explains the cube and roll-up operators, shows how they fit in SQL, explains how users can define new aggregatefunctions for cubes, and discusses efficient techniques to compute the cube.
arXiv (Cornell University)HOGWILD!: A Lock-Free Approach to Parallelizing Stochastic Gradient Descent
1,224 Citations2011Feng Niu, Benjamin Recht +2 more
The MIT Press eBooksThe Tradeoffs of Large-Scale Learning
1,185 Citations2011Léon Bottou, Olivier Bousquet
This contribution develops a theoretical framework that takes into account the effect of approximate optimization on learning algorithms and shows distinct tradeoffs for the case of small-scale and large-scale learning problems.
Parallelized Stochastic Gradient Descent
1,074 Citations2010Martin Zinkevich, Markus Weimer +2 more
This paper presents the first parallel stochastic gradient descent algorithm including a detailed analysis and experimental evidence and introduces a novel proof technique — contractive mappings to quantify the speed of convergence of parameter distributions to their asymptotic limits.
Proceedings of the VLDB EndowmentThe MADlib analytics library
369 Citations2012Joseph M. Hellerstein, Christoper Ré +9 more
Large Scale Online Learning
337 Citations2003Léon Bottou, Y. Le Cun
It is argued that suitably designed online learning algorithms asymptotically outperform any batch learning algorithm in situations where training data is abundant and computing resources are comparatively scarce.
Mathematical Programming ComputationParallel stochastic gradient algorithms for large-scale matrix completion
317 Citations2013Benjamin Recht, Christopher Ré
Jellyfish is an algorithm for solving data-processing problems with matrix-valued decision variables regularized to have low rank that is orders of magnitude more efficient than existing codes.
Proceedings of the VLDB EndowmentTuffy
257 Citations2011Feng Niu, Christopher Ré +2 more
This work presents Tuffy, a scalable Markov Logic Networks framework that achieves scalability via three novel contributions: a bottom-up approach to grounding, a novel hybrid architecture that allows to perform AI-style local search efficiently using an RDBMS, and a theoretical insight that shows when one can improve the efficiency of stochastic local search.
Towards a unified architecture for in-RDBMS analytics
178 Citations2012Xixuan Feng, Arun Kumar +2 more
This work proposes a unified architecture for in-database analytics that requires changes to only a few dozen lines of code to integrate a new statistical technique, and demonstrates the feasibility of this architecture by integrating several popular analytics techniques into two commercial and one open-source RDBMS.
Ricardo
166 Citations2010Sudipto Das, Yannis Sismanis +4 more
R Ricardo is part of the eXtreme Analytics Platform (XAP) project at the IBM Almaden Research Center, and rests on a decomposition of data-analysis algorithms into parts executed by the R statistical analysis system and parts handled by the Hadoop data management system.
arXiv (Cornell University)Factoring nonnegative matrices with linear programs
113 Citations2012Victor Bittorf, Benjamin Recht +2 more
A data-driven model for the factorization where the most salient features in the data are used to express the remaining features and this method extends to more general noise models and leads to efficient, scalable algorithms.
Brainwash: A data system for feature engineering
103 Citations2013Michael R. Anderson, Dolan Antenucci +8 more
This work proposes brainwash, a vision for a feature engineering data system that could dramatically ease the ExploreExtract-Evaluate interaction loop that characterizes many trained system projects.
International Journal on Semantic Web and Information SystemsElementary
94 Citations2012Feng Niu, Ce Zhang +2 more
The authors present Elementary, a prototype KBC system that is able to combine diverse resources and different KBC techniques via machine learning and statistical inference to construct knowledge bases and empirically show that this decomposition-based inference approach achieves higher performance than prior inference approaches.
Mathematical ProgrammingTwo “well-known” properties of subgradient optimization
81 Citations2007Kurt M. Anstreicher, Laurence A. Wolsey
Two topics are related in that convergence of the iterates is required to prove correctness of the primal construction scheme, and the construction of primal estimates when subgradient optimization is applied to maximize the Lagrangian dual of a linear program.
