Gemini 1.0 Launch: Model Sizes, Benchmarks & Demo Video

Google announced Gemini in Ultra, Pro and Nano sizes, staggered their availability and published benchmark claims alongside an edited multimodal demo.

Three sizes, built multimodal from the start

Google announced Gemini 1.0 on December 6, 2023, in three sizes: Ultra for the most complex tasks, Pro for scaling across a wide range of products, and Nano for running on-device. Google described the model as built from the ground up to handle text, code, audio, image and video together, rather than stitching separate models behind one interface. [1]

Availability was staggered. Gemini Pro went into Bard immediately, and Nano began shipping on Pixel 8 Pro. Ultra was held for further safety testing and a broader release the following year, so the model most of the launch coverage described was not the one most people could actually use that week. [1]

A human-expert MMLU score, with an asterisk

Google reported Gemini Ultra scoring 90.0% on MMLU, a 57-subject exam benchmark, and called it the first model to exceed a human-expert threshold on that test. It also cited a 59.4% score on the multimodal MMMU benchmark and claimed state-of-the-art results on 30 of 32 benchmarks used to evaluate large language models, with GPT-4 shown as the comparison point in Google's own charts. [1]

Atlas interpretation: The headline score belonged to Ultra, the size still undergoing safety testing at launch, not to Pro or Nano, which were what shipped that day. A reader who only saw the 90.0% figure and then opened Bard got the mid-tier model. [1]

The demo that was not quite the demo

Alongside the announcement, Google published a video called Hands-on with Gemini: Interacting with multimodal AI, showing what appeared to be smooth, real-time voice and visual exchanges: narrating a sketch as it was drawn, following a ball hidden under cups, and reading hand gestures as a rock-paper-scissors move. Google's own companion post described the clips as illustrations of prompting approaches used to produce the results, not a recording of one continuous live session. [2]

Reporting after the launch found the video was built from still frames and text prompts rather than the fluid spoken back-and-forth it implied, and that some interactions, including the rock-paper-scissors segment, required showing Gemini all three gestures at once with a hint rather than recognizing a single gesture unprompted. A Google spokesperson said the video showed real Gemini outputs edited for pace, and Google DeepMind's Oriol Vinyals said it illustrated what a multimodal experience built with Gemini could look like. [3]

Atlas interpretation: Those two things are both true and do not fully answer the same question. The outputs being real establishes that Gemini could produce them under engineered conditions. It does not establish that a user talking to the product that shipped would see the same latency or the same unprompted recognition the video suggested. [3]

Sources

  1. Introducing Gemini: our largest and most capable AI model

    Google · Dec 6, 2023

  2. How it's Made: Interacting with Gemini through multimodal prompting

    Google Developers Blog · Dec 6, 2023

  3. Google's best Gemini demo was faked

    TechCrunch · Dec 7, 2023