IBM Watson on Jeopardy: How DeepQA Worked & What Followed

IBM Watson beat Ken Jennings and Brad Rutter using candidate search and hand-built scorers; the page traces the match and Watson Health's later collapse.

Generating hundreds of guesses before judging any of them

DeepQA parsed each clue, guessed its answer type, and checked whether it needed splitting into subclues. Primary search then queried text engines, passage indexes, and triple stores against a corpus IBM had assembled in advance, with no live web access during play, favoring recall over precision and generating several hundred loosely supported candidate answers per question. [2]

A soft filter narrowed that field to roughly 100 candidates. Survivors were checked against retrieved evidence by more than 50 scoring algorithms measuring passage alignment, source reliability, geospatial and temporal consistency, and grammatical fit. On a clue about a 1974 pardon, one scorer favored Nixon over Ford because only Nixon filled the object position in a matching retrieved sentence. [2]

Fifty scorers and a metalearner, not a neural network

Atlas interpretation: That combination step is the part later retellings compress into deep learning, a term the paper never uses. A metalearner merged each candidate's scores using techniques the paper calls mixture of experts and stacked generalization, trained on questions with known answers, closer to layered logistic regression over hand-built features than to a network learning its own representations. [2]

Scaled onto roughly 2,500 cores with the Apache UIMA-AS framework, the pipeline answered within three to five seconds, the window a contestant has to hear a clue and decide whether to buzz in. An earlier single-processor version had taken two hours per question. [2]

What three days on television actually proved

Watson beat Ken Jennings and Brad Rutter over three days of competition in February 2011, finishing with $77,147 to their $24,000 and $21,600. It also made a memorable, uncorrected error, naming Toronto instead of Chicago on a Final Jeopardy clue about US cities, a reminder that an aggregate score decided the match, not a run of flawless answers. [1]

The healthcare business the Jeopardy win was supposed to fund

Atlas interpretation: That win ran on a corpus IBM curated in advance, with no live web access. IBM's next move, recommending cancer treatment from patient records and medical literature, asked the same architecture to judge evidence nobody had pre-selected. [2]

MD Anderson Cancer Center spent five years and $62 million on a Watson-based Oncology Expert Advisor before letting the contract lapse in 2016 without using it on a patient. An audit blamed procurement and integration problems, including Watson's trouble telling the abbreviation for acute lymphoblastic leukemia apart from the one for a drug allergy. [3]

IBM sold what remained of Watson Health, its oncology, imaging, and clinical-trial data products, to the private equity firm Francisco Partners in January 2022, in a deal reported at just over $1 billion, a fraction of what building the unit had cost. [4]

Atlas interpretation: IBM's later AI strategy shows the lesson taken. Rather than owning end-to-end models again, its 2026 partnership putting OpenAI's models inside its consulting business routes capability through outside vendors. Winning Jeopardy and running a hospital's oncology practice turned out to be different engineering problems. [4]

Sources

  1. Watson and Jeopardy!

    IBM

  2. Building Watson: An Overview of the DeepQA Project

    AI Magazine (AAAI) · Jul 28, 2010

  3. M. D. Anderson Breaks With IBM Watson, Raising Questions About Artificial Intelligence in Oncology

    JNCI: Journal of the National Cancer Institute (Oxford University Press) · May 22, 2017

  4. Francisco Partners scoops up bulk of IBM's Watson Health unit

    TechCrunch · Jan 21, 2022