A smaller model, trained by the bigger one
Google announced Gemini 1.5 Flash at I/O on May 14, 2024, describing it as trained through a process it calls distillation, where the most essential knowledge and skills from a larger model, 1.5 Pro, are transferred to a smaller and more efficient one. [1]
Both 1.5 Pro and 1.5 Flash entered public preview that day with a 1 million token context window. Google said 1.5 Pro was separately available with a 2 million token context window through a waitlist; that larger window was a Pro capability, not a Flash one. [1]
Atlas interpretation: The summary above of a two million token context window describes the broader I/O update rather than Flash specifically. Flash's pitch was the opposite of a bigger window: Google positioned it for summarization, chat, captioning and data extraction at high volume, the tasks where a cheaper, faster model beats a larger one that a developer would otherwise over-provision for the job. [1]
Project Astra: a prototype, shown running, not shipped
Google described Project Astra as a prototype agent built to understand and respond to a continuous camera and microphone feed: encoding video frames as it goes, combining what it sees and hears into a timeline of events, caching that information for quick recall, and answering with a wider range of intonation and minimal conversational lag. [1]
Google was explicit that Astra was not a released product on May 14, 2024, stating that some of the demonstrated capabilities would come to the Gemini app and web experience later that year rather than shipping that day. [1]
The price cut three months later is the tell
On August 8, 2024, Google announced that effective August 12, it was cutting Gemini 1.5 Flash's list price for prompts under 128,000 tokens by 78 percent on input, to $0.075 per million tokens, and 71 percent on output, to $0.30 per million tokens, with the reductions cascading to the over-128,000-token tier and to cached tokens as well. [2]
Atlas interpretation: Those percentages work backward to roughly $0.35 per million input tokens and about $1.05 per million output tokens at launch, for prompts under 128,000 tokens. A vendor does not cut a workhorse model's price by three-quarters three months in unless usage justifies it; Google does not appear to have made an equivalent public cut to Astra's availability schedule in the same period, because there was no product to price. That gap between a model getting cheaper by the season and a demo still waiting on a shipping date is the substance behind the summary's line that Flash became the workhorse while Astra stayed a demo. [2]
Sources
- Gemini breaks new ground with a faster model, longer context, AI agents and more
Google · May 14, 2024
- The next chapter of the Gemini era for developers
Google Developers Blog · Aug 8, 2024