Google has introduced Gemini, a new family of artificial-intelligence models designed to process text, images, audio, video and code, and is immediately deploying the mid-sized Gemini Pro model in Bard. The company says its largest version, Gemini Ultra, exceeds current state-of-the-art performance on 30 of 32 widely used academic benchmarks, although that model will remain in safety testing and limited evaluation until next year.

The Dec. 6 launch is Google’s most direct technical answer yet to the rapid adoption of OpenAI’s GPT-4 and ChatGPT. Rather than release one model for all uses, Google has built three versions: Ultra for the most complex tasks, Pro for broad cloud-scale use and Nano for efficient on-device computation.

Three models target different computing environments

Gemini is designed from the outset as a multimodal system rather than a text model with separate components attached later, according to Google DeepMind. The company says that approach allows the models to reason across combinations of written language, images, audio and video and to perform programming tasks. Google’s launch collection describes Gemini 1.0 as trained at scale on the company’s Tensor Processing Unit infrastructure.

Gemini Pro is the version most users can experience now. Google has placed a specially tuned English-language version into Bard in more than 170 countries and territories. The company says the upgrade improves reasoning, planning and understanding. A Bard announcement says access to Gemini Ultra will come later through a more advanced Bard experience after additional evaluation.

Gemini Nano is aimed at mobile devices. Google’s Pixel update says the Pixel 8 Pro is using Nano for features including summarization in the Recorder application and Smart Reply in Gboard. Running some generative-AI functions locally can reduce latency, permit offline use and keep some information on the device instead of sending it to a remote data center.

Google makes aggressive benchmark claims

The most striking part of the announcement is Google’s claim that Gemini Ultra sets new results on a broad group of tests. The company says Ultra exceeds existing state-of-the-art results on 30 of 32 benchmarks used for large language models and multimodal systems. It reports a score of 90% on MMLU, a benchmark spanning 57 subjects such as mathematics, physics, history, law, medicine and ethics.

Those results should be treated as company-reported benchmark findings rather than evidence from broad public use. Gemini Ultra is not yet generally available, which means outside developers and researchers cannot independently reproduce the full set of claims in ordinary deployment. Ars Technica noted that Google says Ultra beats GPT-4-associated results across many tests, but also emphasized that benchmark comparisons between proprietary models can be sensitive to prompting, evaluation methods and model access.

The Guardian reported that Google says Ultra outperformed rival systems in 30 of the 32 tests it highlighted. The company also acknowledged that hallucinations — confidently generated false or unsupported answers — remain an unresolved research problem. That caveat is important because benchmark strength does not remove the reliability limitations that have affected every major generative-AI system.

Ultra remains behind a safety gate

Google is not releasing its most powerful model broadly at launch. The company says Gemini Ultra is undergoing external red-team testing, safety evaluation, fine-tuning and reinforcement learning from human feedback. Select customers, developers and experts will receive early access before a wider release planned for early 2024.

The staged approach reflects growing scrutiny of frontier AI systems. President Joe Biden’s October executive order requires developers of certain highly capable models to provide the federal government with information about safety testing under specified conditions. Google says it intends to cooperate with government safety efforts and is also discussing evaluation with the United Kingdom’s new AI Safety Institute.

Gemini Pro will become available to developers and enterprise customers through Google AI Studio and Vertex AI beginning Dec. 13, according to the company. That distribution strategy makes Gemini not just a consumer-chatbot upgrade but a platform technology intended to compete for corporate software, cloud-computing and application-development workloads.

The competitive stakes extend across Google’s product portfolio

Google has unusual advantages in distributing a new foundation model because it controls widely used consumer services, Android, Chrome, Search and a major cloud platform. The company says Gemini will eventually appear across Search, Ads, Chrome and other products. Early experiments in Search have reduced latency in its Search Generative Experience, according to Google’s launch materials.

That breadth also raises the stakes. Microsoft has integrated OpenAI technology into Bing, GitHub and its enterprise software, while OpenAI has built a large direct user base through ChatGPT. Google must show not only that Gemini performs well in controlled tests, but that it can improve products used by billions of people without introducing unacceptable reliability, privacy or safety problems.

The initial release therefore creates a two-stage test. Gemini Pro and Nano can begin accumulating real-world evidence now through Bard and Pixel devices. Gemini Ultra, the model behind the company’s strongest performance claims, must first clear additional internal and external review.

Google’s announcement establishes Gemini as the company’s central AI architecture going forward. Whether it changes the competitive balance will depend less on a single benchmark table than on how the three model sizes perform when developers, enterprises and consumers use them at scale. As of this week, that process has begun for Pro and Nano; the most consequential claims attached to Ultra remain to be tested outside Google.