Google is keeping Gemini’s ability to generate images of people offline after a week of criticism over historically inaccurate and biased outputs, turning a product-quality failure into a broader test of how quickly large technology companies can deploy generative AI without losing control of safety systems layered on top of their models. CEO Sundar Pichai told employees Tuesday that some Gemini responses were “completely unacceptable” and said the company would make structural changes before restoring the feature.
The controversy began after users posted examples in which Gemini produced implausible racial and gender depictions for historical prompts. Google acknowledged the problem and paused people-image generation on February 22. In a detailed company explanation, Google Senior Vice President Prabhakar Raghavan said the system had overcompensated for diversity in some prompts while becoming excessively cautious in others.
A safety adjustment produced a historical-accuracy failure
Generative image systems are typically tuned to reduce stereotyped or exclusionary outputs. Google said Gemini’s safeguards were intended to avoid repeating biases that have appeared in earlier image-generation systems, but the implementation produced the opposite failure in some historical contexts. The model generated people who did not fit the period or prompt and, in other cases, refused benign requests.
The Associated Press reported that Google temporarily suspended the feature after examples circulated widely online. The company said it would improve accuracy before turning people-image generation back on. That decision matters because Gemini is not an experimental research demo; Google has positioned it as the consumer-facing identity for its most important generative-AI products.
Reuters reporting republished by Yahoo Finance noted that the pause came only weeks after Google began offering image generation through Gemini and shortly after Bard was renamed Gemini. The compressed timeline illustrates the competitive pressure facing Google as it races Microsoft and OpenAI across consumer assistants, enterprise AI and developer platforms.
Pichai promises changes to launch and evaluation processes
In an internal memo published by Semafor, Pichai said the problematic text and image responses had offended users and shown bias. He said teams were “working around the clock” and outlined a set of actions that included structural changes, updated product guidelines, improved launch processes, stronger evaluations, expanded red-teaming and technical recommendations.
That list is important because it frames the incident as more than a narrow image-model defect. Pichai is signaling that the failure may involve the way Google translates model capabilities into consumer products: prompt interpretation, policy layers, safety filters, testing coverage and release governance. In other words, fixing the output may require changing the process that allowed the output into production.
The company has not given a firm date for restoring people-image generation. Pichai indicated that Google expects improvements over the coming weeks, but the feature remains disabled as of Saturday. That is a meaningful retreat for a company that has spent February accelerating Gemini’s public rollout.
The pause comes amid an unusually aggressive Gemini expansion
On February 8, Google formally replaced the Bard name with Gemini and introduced Gemini Advanced, powered by Ultra 1.0, as well as a dedicated Android application. In that launch announcement, Pichai described Gemini as an ecosystem spanning consumer products, APIs, cloud services and developer tools rather than a single chatbot.
One week later, Google announced Gemini 1.5, highlighting a context window that can reach 1 million tokens in limited preview and process long documents, video, audio and large codebases. Those technical advances demonstrate that the underlying product strategy is continuing even while one visible consumer feature is paused.
The tension is becoming central to the AI industry: model capability is advancing rapidly, while product safeguards remain difficult to calibrate. A system can be more capable at reasoning, coding or multimodal analysis and still fail because policy tuning or safety interventions distort specific outputs.
Trust becomes a product requirement, not a communications issue
Google’s problem is especially sensitive because the company’s core business is built on information retrieval and user trust. Gemini is being integrated into products that could eventually influence search, productivity, advertising and mobile computing. If users believe the system is politically or culturally biased, or simply unreliable on basic historical questions, the problem can spread beyond one image feature.
The company’s own explanation acknowledges that generative AI will continue to make mistakes. Raghavan wrote that hallucinations and imperfect outputs remain known challenges and cautioned users against relying on Gemini for evolving news or sensitive topics without verification. That admission is realistic, but it also highlights the gap between the probabilistic nature of the technology and the expectations placed on a Google-branded information product.
The immediate engineering objective is to restore people-image generation without recreating stereotypes or producing implausible historical scenes. The broader objective is harder: build a release process that can detect when a well-intended safety intervention creates a different form of distortion. As Google continues to expand Gemini across its product portfolio, this week’s pause is a reminder that the competitive race is not only about larger models and longer context windows. It is also about whether companies can make those systems behave predictably enough to deserve the trust required for mass-market use.