The BellKor Solution to the Netflix Grand Prize
Published 1 January 2009
Yehuda Koren
Citations340
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
TL;DR
Part of the contribution to the “BellKor’s Pragmatic Chaos” final solution, which won the Netflix Grand Prize, is described, which improved the baseline predictors and introduced a new blending algorithm based on gradient boosted decision trees.
Abstract
This article describes part of our contribution to the “Bell-Kor’s Pragmatic Chaos ” final solution, which won the Netflix Grand Prize. The other portion of the contribution was created while working at AT&T with Robert Bell and Chris Volinsky,
Keywords
Computer SciencePhysics and Astronomy
The Annals of StatisticsGreedy function approximation: A gradient boosting machine.
28,973 Citations2001Jerome H. Friedman
A general gradient descent boosting paradigm is developed for additive expansions based on any fitting criterion, and specific algorithms are presented for least-squares, least absolute deviation, and Huber-M loss functions for regression, and multiclass logistic likelihood for classification.
Computational Statistics & Data AnalysisStochastic gradient boosting
6,887 Citations2002Jerome H. Friedman
It is shown that both the approximation accuracy and execution speed of gradient boosting can be substantially improved by incorporating randomization into the procedure.
Factorization meets the neighborhood
3,876 Citations2008Yehuda Koren
The factor and neighborhood models can now be smoothly merged, thereby building a more accurate combined model and a new evaluation metric is suggested, which highlights the differences among methods, based on their performance at a top-K recommendation task.
Restricted Boltzmann machines for collaborative filtering
1,882 Citations2007Ruslan Salakhutdinov, Andriy Mnih +1 more
This paper shows how a class of two-layer undirected graphical models, called Restricted Boltzmann Machines (RBM's), can be used to model tabular data, such as user's ratings of movies, and demonstrates that RBM's can be successfully applied to the Netflix data set.
Collaborative filtering with temporal dynamics
1,170 Citations2009Yehuda Koren
Two leading collaborative filtering recommendation approaches are revamp and a more sensitive approach is required, which can make better distinctions between transient effects and long term patterns.
The Netflix Prize
1,132 Citations2007James R. Bennett, Stan Lanning
Netflix released a dataset containing 100 million anonymous movie ratings and challenged the data mining, machine learning and computer science communities to develop systems that could beat the accuracy of its recommendation system, Cinematch.
Scalable Collaborative Filtering with Jointly Derived Neighborhood Interpolation Weights
503 Citations2007Robert M. Bell, Yehuda Koren
This work enhances the neighborhood-based approach leading to substantial improvement of prediction accuracy, without a meaningful increase in running time, and suggests a novel scheme for low dimensional embedding of the users.
McRank: Learning to Rank Using Multiple Classification and Gradient Boosting
434 Citations2007Ping Li, Qiang Wu +1 more
This work considers the DCG criterion (discounted cumulative gain), a standard quality measure in information retrieval, and proposes using the Expected Relevance to convert class probabilities into ranking scores.
A General Boosting Method and its Application to Learning Ranking Functions for Web Search
185 Citations2007Zhaohui Zheng, Hongyuan Zha +4 more
This work presents a general boosting method extending functional gradient boosting to optimize complex loss functions that are encountered in many machine learning problems, based on optimization of quadratic upper bounds of the loss functions.
The BellKor 2008 Solution to the Netflix Prize
132 Citations2007Robert M. Bell, Yehuda Koren +1 more
The final solution (RMSE=0.8712) consists of blending 107 individual results, and the main approaches behind them are described, which can achieve slightly more accurate results than the newer one, at the expense of a significant increase in running time.
Improved neighborhood-based algorithms for large-scale recommender systems
69 Citations2008Andreas Töscher, Michael Jahrer +1 more
A way to calculate similarities by formulating a regression problem which enables us to extract the similarities from the data in a problem-specific way and leads to increased prediction accuracy.
Machine Learned Sentence Selection Strategies for Query-Biased Summarization
46 Citations2008Donald Metzler, Tapas Kanungo
This work is the first to evaluate SVR andGBDTs for the sentence selection task, and shows that GBDTs provide a robust and powerful framework for the sentences selection task and significantly outperform Svr and ranking SVMs on several data sets.
The BigChaos Solution to the Netflix Prize 2008
28 Citations2008Andreas Töscher
The team “BellKor in Bigchaos” is a combined team of team BellKor and BigChaos, and the solution with a RMSE of 0.8616 is created by a linear blend of the results from both teams.
