Netflix Prize: $1M Winner, Ensemble Method, Privacy Fallout

Netflix awarded $1 million after a blended recommender beat Cinematch by 10%, never deployed the full system, and canceled its sequel over privacy concerns.

Three years for ten percent

Netflix opened the contest on October 2, 2006, offering $1,000,000 to the first team that beat the accuracy of its own recommender, Cinematch, by 10 percent on a held-out set of movie ratings, measured by root mean squared error. Annual $50,000 progress prizes went out in 2007 and 2008 to whichever team was furthest ahead at the time. [1]

On September 21, 2009, Netflix announced that a team called BellKor's Pragmatic Chaos had crossed the 10 percent threshold and would take the full $1,000,000 grand prize. The team was a merger of three groups that had spent the competition's final stretch combining their separately built models: researchers from AT&T Labs, a group called Pragmatic Theory, and BigChaos, a two-person Austrian team. [1]

Atlas interpretation: The final margin over the runner-up team, The Ensemble, was reportedly a fraction of a percentage point, close enough that Netflix took the extra step of comparing submission timestamps to settle the win. That the deciding factor came down to minutes, after three years and over 44,000 submissions from more than 5,000 teams, says something about how flat the top of this kind of leaderboard gets once the easy gains are gone. [1]

The winning approach was blending, not one clever model

No single algorithm reached 10 percent on its own. The winning entries were ensembles: hundreds of individual predictors, from matrix factorization to nearest-neighbor methods to time-aware models that accounted for how a user's rating habits drifted over the years, blended together and weighted to minimize error on the leaderboard's held-out data. [1]

Netflix said it never put the winning ensemble into production. The company had already found that a couple of the individual algorithms folded into the final blend produced most of the accuracy gain on their own, and that the added complexity of running the full ensemble was not worth the small extra lift for a live recommender system. [1]

A second prize, announced the same day

Netflix used the same announcement to open a follow-on contest, this one scored on how well entrants could predict demographic attributes such as age and gender rather than just ratings, with the prize split into a $500,000 payment after six months and $500,000 after eighteen. [1]

The sequel never ran

In 2008, researchers Arvind Narayanan and Vitaly Shmatikov showed that the contest's supposedly anonymized ratings data could be cross-referenced against public reviews on sites such as IMDb to identify individual subscribers, arguing that even a little auxiliary knowledge about a person's viewing habits was enough to single out their record in the dataset. [3]

Four Netflix subscribers filed a federal lawsuit in December 2009 alleging the ratings release violated their privacy and the Video Privacy Protection Act, and the Federal Trade Commission separately raised reidentification concerns with the company. On March 12, 2010, Netflix cancelled the second contest, saying it had reached an understanding with the FTC about handling user data and had settled the litigation. [2]

Atlas interpretation: The sequel died for a reason that had nothing to do with the modeling problem it was built around. The first contest's release schedule had already put the sensitive data in public hands years before the de-anonymization work was published, so cancelling the second round closed a door the field had already learned was easier to open than assumed. [3][2]

Sources

  1. Netflix Awards $1 Million Prize and Starts a New Contest

    The New York Times · Sep 21, 2009

  2. NetFlix Cancels Recommendation Contest After Privacy Lawsuit

    Wired · Mar 12, 2010

  3. How To Break Anonymity of the Netflix Prize Dataset

    arXiv · Oct 18, 2006