The One Test Google’s Best AI Keeps Failing

Illustration of Google Gemini 3.5 Pro delay amid the AI coding benchmark race
KEY POINTS
  • Google pushed Gemini 3.5 Pro’s broad launch back by months after it fell short of internal coding and long-horizon reasoning targets.
  • Unveiled at Google I/O in May and expected in June, the flagship is stuck in a limited Vertex AI enterprise preview with no general-availability date.
  • A late-June retraining meant to fix coding delivered “disappointing” results, per Bloomberg, and Alphabet shares slid more than 4% on the report.
  • While Google stalled, China’s open-weight Kimi K3 hit No. 1 on a frontend coding benchmark with a 76% win rate, topping GPT-5.6 Sol and Claude Fable 5.

More than 4%. That is how far Alphabet’s stock slid in a single stretch after Bloomberg reported that Google had quietly delayed the broad release of Gemini 3.5 Pro, its flagship frontier model, by several months. The cause was not a lawsuit, a safety scare, or a chip shortage. It was one stubborn weakness the company could not engineer away on schedule: writing code.

What Google Actually Delayed

A June Launch That Never Arrived

Gemini 3.5 Pro was previewed at Google I/O in mid-May 2026, with the Pro tier expected to arrive around June. That deadline came and went with no general-availability date, no published benchmarks, and no final pricing. Instead the model has lived inside a limited Vertex AI enterprise preview, open only to a small group of approved customers plus testers on Google’s Antigravity platform and the LMArena leaderboard. For a company that helped define the modern frontier-model race, shipping months late and doing it quietly is exactly the kind of signal rivals notice.

Trend Insight — A slipped launch is not the same as a lost lead. But in a market where OpenAI, Anthropic, Meta and Chinese labs ship on a near-monthly cadence, silence starts to read like weakness.


The Coding Problem Money Could Not Rush

Three Misses, One Pattern

According to Bloomberg, Google is “taking time to try to improve” Gemini 3.5 Pro’s capabilities, “particularly in coding.” Internal testing reportedly surfaced three linked issues: token-efficiency concerns raised by early testers, coding performance below flagship standard, and long-horizon, multi-step reasoning that fell short of the bar Google set on stage at I/O. In late June, engineers updated the data used to train Gemini specifically to lift its coding skills, and the results were described as “disappointing.” Reporting from TechTimes said the rebuilt model missed a third internal deadline. When the delay surfaced, Alphabet shares fell more than 4%, a reminder that investors now price frontier execution as a core business risk.

Trend Insight — Scale, TPUs, DeepMind and oceans of data guarantee capacity, not predictable progress. The last few percent of coding reliability is where raw compute stops being enough.


While Google Stalled, a Chinese Model Surged

Open Weights, Closing Gap

The timing was brutal. In the same week, Beijing-based Moonshot AI released Kimi K3, an open-weight model built for long-horizon workflows, the exact capability Google is struggling to ship. Kimi K3 reached No. 1 on Arena.ai’s Frontend Code Arena with a 76% pairwise win rate, finishing ahead of Anthropic’s Claude Fable 5 and OpenAI’s GPT-5.6 Sol on that benchmark. It scored 88.3 on Terminal Bench 2.1, narrowly trailing GPT-5.6 Sol’s 88.8, and placed ninth overall on the broader Text Arena, a major jump from the prior Kimi generation. Because the weights are open, enterprises can run it on their own infrastructure and tune it for niche tasks, with no single US vendor required.

Trend Insight — The distance between US and Chinese frontier labs is now measured in benchmark decimals, and open weights turn every release into leverage for the buyer.


What the Delay Signals for the Frontier Race

Design for a Multi-Model World

Google still controls one of the most complete AI stacks on earth: custom TPUs, Google Cloud, DeepMind, Android, Search, YouTube, Workspace, and billions of accounts. That breadth is also the trap. A flagship model has to work reliably across products people use every day, which quietly raises the bar for what counts as “done.” For enterprise teams evaluating models right now, the practical lesson is simple: do not architect around an unreleased flagship. Design for a multi-model world, keep an escape hatch to open-weight options, and treat benchmark leadership as temporary. The frontier is no longer a ladder one company climbs alone; it is a field where the lead changes hands by the week.

Trend Insight — The winners of this phase will not be whoever announces first, but whoever ships reliable, agentic coding that survives contact with real production workloads.


Related

Sources

  1. Bloomberg — Google Gemini Launch Delayed as Tech Falls Short of Internal Goals (Jul 16, 2026)
  2. 9to5Google — Gemini 3.5 Pro delays due to coding performance (Jul 16, 2026)
  3. TechStartups — Top Tech News Today, July 17, 2026 (Gemini delay; Moonshot Kimi K3)

AI Biz Insider · AI Trends EN · aibizinsider.com


AI Biz Insider에서 더 알아보기

구독을 신청하면 최신 게시물을 이메일로 받아볼 수 있습니다.

코멘트

댓글 남기기

AI Biz Insider에서 더 알아보기

지금 구독하여 계속 읽고 전체 아카이브에 액세스하세요.

계속 읽기

AI Biz Insider에서 더 알아보기

지금 구독하여 계속 읽고 전체 아카이브에 액세스하세요.

계속 읽기