Google's next flagship model has reached the last major stage before release. Speaking at The Information's AI Agenda Live Summit on September 23, Koray Kavukcuoglu, who took over operational leadership of Google DeepMind in August, confirmed that Gemini 4 has entered early post-training — the phase where a finished base model is refined with human feedback, reinforcement learning and specialized data to shape its reasoning, safety and reliability.
More striking than the milestone itself was the delivery model that came with it. Kavukcuoglu said Google intends to release an early post-training output "as soon as possible, because we see the results and we are excited," then continue fast-paced iterations from there. That is a genuine break from Google's own habits: every prior Gemini generation arrived as a finished product, with documentation, published benchmarks and a launch event. Gemini 4 will instead be judged on iteration speed, with an early checkpoint refined in public.
The urgency is not hard to read. Google has not shipped a new flagship since the Gemini 3 series in November 2025, followed by the incremental Gemini 3.1 Pro in February 2026. A promised Gemini 3.5 Pro update, announced by Sundar Pichai at I/O in May for a June release, never arrived and was quietly shelved as the company pivoted to its faster, cheaper Flash models. Kavukcuoglu admitted Google "took a little bit of a step back" during that period.
Its rivals kept swinging. OpenAI launched GPT-6 Astra on September 3 and followed with the cheaper GPT-6 Sol and Luna on September 22. Anthropic shipped Claude Fable 5.1 on September 1 and Claude Opus 5.5 on September 22. The market has noticed the gap: Alphabet shares fell nearly 5 percent over two sessions this week as Meta's Muse agent climbed the US app charts during Meta Connect.
Internally, Gemini 4 is already being exercised. Google engineers have used the model to power Antigravity, the company's agentic development environment, as part of the safety testing and guardrail work that post-training involves. An unnamed entry on the benchmark arena LMArena, labeled "gemini-3.8-flash," is widely believed by developers to be an early Gemini 4 Pro checkpoint — a circulating benchmark table ranks it first on agentic coding and knowledge work, though Google has not confirmed the entry's identity and leaked tables are unverified by nature.
The announcement is also the first major test of DeepMind's new command structure. Kavukcuoglu took operational control in August after co-founder Demis Hassabis stepped back to become president and chief scientist of Alphabet. Asked whether Google has fallen behind, he said he has "the utmost trust in the team" and that "in my mind, it's a certainty that we are always gonna be at the frontier."
The open question is safety. An early checkpoint has, by definition, spent less time being hammered with the adversarial prompts that surface jailbreaks and confabulation before the public does — and for Google, Search, Cloud and enterprise deployments all sit on top of Gemini. If an early version ships, millions of users effectively become the last mile of red-teaming. Asked about AGI, Kavukcuoglu reframed the debate: the real issue, he said, is "are we able to build intelligent agents that we can trust." Whether "as soon as possible" means November or January is the part Google is not saying.
Comments (0)
Log in to join the discussion
Log InNo comments yet