Google Launches Gemini to Catch Up with AI Competitors; Internal Evaluations of Its Programming Performance Vary
The Block
55m ago
Ai Focus
Alphabet's subsidiary Google has begun to launch its flagship AI model Gemini 4 Argon, but there are differing opinions within the company regarding its performance in key areas such as programming. Google claims that the model leads in several benchmark tests, while insiders suggest that its effectiveness in actual employee use is not that outstanding.
Helpful
No.Help

Google, with the code name Alphabet, has begun to launch its much-anticipated flagship artificial intelligence model Gemini 4 Argon. However, there are doubts within the company regarding the actual performance of this model in key areas such as programming.

On Wednesday, Google made this model available to a small group of trusted cybersecurity partners and stated that it would expand its use after conducting more tests, initially opening it to paid users. The company said that Gemini achieved leading results in several benchmark tests, including defeating the Astra model of OpenAI in a test that measures security capabilities.

However, some insiders suggest that these indicators do not reflect the full picture. According to individuals familiar with the work, although Gemini 4 performed exceptionally well in benchmark tests measuring model efficiency, its effectiveness was not as outstanding when actually used by employees, especially for certain programming tasks. Out of concern for discussing internal matters, these individuals have requested to remain anonymous.

Alphabet's stock price rose by about 1.7% in after-hours trading following the company's announcement. However, towards the end of regular trading hours, after media reports indicated that some internal employees were skeptical about the model, Alphabet's stock price gave back some of the gains seen earlier on Wednesday.

Google has recently been committed to developing models that can compete with OpenAI and Anthropic PBC. People familiar with the matter revealed that the company originally planned to release a new version called Gemini 3.5 Pro in June, but ultimately abandoned the project.

Google stated that it is inaccurate to claim that Gemini performs poorly in fields such as programming, and referred to the speech given last week by DeepMind's leader, Koray Kavukcuoglu, who was inspired by the model's performance at that time.

“I have full confidence in the team,” Kavukcuoglu said at a conference organized by the technology news website The Information. “In my opinion, we will always be at the forefront of the industry.”

There are differing opinions within Google. Some employees believe that the Fable and Astra models of Anthropic are progressing faster than Gemini. These individuals argue that even if Gemini performs at its best, it will still fall behind these models in certain aspects. On the other hand, other employees think that the upcoming version has already caught up with leading artificial intelligence laboratories.

A Google employee familiar with model development stated that there is a "general belief" within the company that Gemini 4 is at the forefront of technology. The employee mentioned that the company has conducted rigorous tests on the model and denied any issues with the model in handling complex, real-world coding tasks.

Google urgently needs the success of Gemini. From the artificial intelligence answers in Google's main profit-making engine—the search function, to Maps, Gmail, and Chrome browsers, different versions of this model support nearly all of Google's products. The number of users for these products exceeds one billion each, which is a distribution advantage that some of Google's competitors do not possess.

However, OpenAI and Anthropic are gradually shifting from selling the AI model to developing their own products, including programming agents. If Google is unable to provide cutting-edge models, its competitors may have more time to convince consumers, developers, and businesses that the future of search and software should run on their platforms.

In response to related inquiries, Google stated that although its previous Pro model was released in February, the company's subsequent AI products have seen growth, including the enterprise version Gemini, chatbot applications for consumers, and the AI mode within Google Search. The number of users for the latter two has exceeded one billion each.

Gemini 3

Last November, Google released the highly praised Gemini 3, which was widely seen as a turning point for Google to catch up with OpenAI and Anthropic. In May this year, Google announced Gemini 3.5 at the I/O developer conference and promised to release it the following month, but failed to do so on time. People familiar with the situation said that the company has abandoned 3.5 Pro.

In addition to weighing down Google's ambitions of AI, this decision is also likely to have cost the company a significant amount of time and money. Industry research analyst Mandeep Singh stated that training such a model could cost up to $400 million. Moreover, hiring highly paid artificial intelligence researchers could further increase these costs.

Today, Google is facing many challenges in the development of Gemini. People familiar with the internal evaluation of this model have revealed that its coding capabilities are uneven. One of them stated that Gemini is not adept at front-end design, which determines the appearance and user experience of applications and websites. This could be a serious setback for Google, as it has been striving to compete with its rivals in the highly competitive AI programming tools market.

In addition, another person familiar with the development situation said that this is a very large model. Generally speaking, the operating costs of large models are quite high, which may put pressure on Google's profit margins.

" Benchmaxxing "

Experts say that Google may be influenced by the industry's tendency to place excessive emphasis on benchmark testing – a phenomenon known as “benchmaxxing”, where engineers focus on achieving high scores rather than creating high-quality products. The reason why artificial intelligence laboratories behave this way is that clients typically evaluate models based on their benchmark test scores. Two individuals familiar with this model indicate that Gemini 4 also seems to be affected by this tendency.

The founder of the artificial intelligence startup Surge AI, Edwin Chen, stated that relying on benchmark tests may encourage laboratories to focus on creating models that are adept at writing code in a specific language, rather than developing truly user-friendly or well-designed applications.

He said, "For example, 'Oh, my child SAT did really well' – but SAT's grades cannot be translated into actual performance. This is a problem with great harm."

According to an informed source familiar with the situation, Gemini 4 does indeed have some advantages. For example, the model performs exceptionally well in understanding inputs other than text, such as extracting metadata from videos. In addition, the model also has strengths in terms of security, network safety, and clear, natural communication capabilities.

There is already dissatisfaction within Google. Researchers who were interviewed before attributed this to the bureaucratic approach of attempting to integrate AI technology into almost all of Google's products. They stated that the constantly changing requirements and priorities have made it difficult to focus on a coherent strategy.

At the same time, a group of star researchers including legendary engineer Jeff Dean, Nobel laureate John Jumper, and Noam Shazeer who participated in inventing the key technologies that fueled this AI craze, left Google. In August, Demis Hassabis who had long led Google's artificial intelligence research was appointed as chairman and handed over the day-to-day operations of DeepMind to his long-time deputy, Kavukcuoglu.

Google stated on Wednesday that Gemini 4 is designed specifically for complex, lengthy tasks in fields such as software engineering, finance, law, and cybersecurity. The company claims that Gemini 4 can generate responses of a length surpassing that of its predecessor products, handling up to 1 million tokens at a time, which is approximately 750,000 words.

Google is working hard to catch up, while its competitors are also accelerating their progress. Despite discussions about slowing down the development of some cutting-edge models after a series of incidents where AI intelligents were found to have infiltrated external systems, earlier this month, Meta Platforms Inc released Muse, an artificial intelligence entity claimed to be capable of performing daily tasks such as online shopping and making appointments. This app quickly rose to the top of the download charts.

Tip
$0
Like
0
Save
0
Views 16
WalletJYS reminds readers to view blockchain rationally, stay aware of risks, and beware of virtual token issuance and speculation. All content on this site represents market information or related viewpoints only and does not constitute any form of investment advice. If you find sensitive content, please click“Report”,and we will handle it promptly。
Submit
Comment 0
Hot
Latest
No comments yet. Be the first!
Related
Chainlink provides data support for the new Open USD stablecoin
Open Standard Announces the Launch of Open USD (OUSD), which has been launched on four blockchains: Base, Ethereum, Solana, and Tempo. Chainlink has been designated as the official data oracle, and the five founding partners have committed to providing over $1 billion in initial liquidity.
crypto.news
·2026-10-01 13:26:21
10
DOGE Price: Dogecoin gains an Ethereum-like testnet, available for trading, lending, and stablecoins
DogeOS launches a public test network that allows developers to use the beta version of DOGE to build Ethereum-compatible trading, lending, and stablecoin applications. In the initial phase, this network relies on designated operators and hardware, but in the future, it is hoped that through Dogecoin Core, Dogecoin miners will be able to verify relevant proofs.
CoinDesk
·2026-10-01 13:15:05
11
Servier Globally introduces Veeva OpenData to expand the application of AI
According to Veeva Systems, Servier has already adopted standardized customer reference data based on the Veeva OpenData standard in over 80 countries, and has established a unified data foundation on the basis of Veeva Link Key People to support the expansion of its AI applications globally. Servier states that this approach will help to improve process efficiency and provide high-quality insights for better reaching the right medical professionals in various markets.
PR Newswire
·2026-10-01 13:05:09
12
IBM Launches IBM Bob Self-Managed Deployment to Help Enterprises Advance AI Sovereignty and Governance
IBM announces the launch of self-hosted deployment for IBM Bob, allowing organizations to bring AI software development to local, private cloud, sovereign cloud, and isolated network environments, so as to use AI without moving sensitive code, data, or workflows out of controlled infrastructure. IBM indicates that this solution is designed to help enterprises advance software development and modernization while meeting security, compliance, and data residency requirements.
PR Newswire
·2026-10-01 13:05:08
13
Arthur Hayes predicts Bitcoin to reach $1 million by 2030, pay attention to AI credit risk
Arthur Hayes reiterates that Bitcoin could rise to $1 million by 2030 and believes that the upward trend may accelerate at the end of 2027 or the beginning of 2028, assuming there is pressure on the AI infrastructure funded by debt. Apollo estimates that AI financing could bring in over $2 trillion in new investment-grade debt; US insurance regulators have also made adjustments to the reporting requirements for private credit holdings.
crypto.news
·2026-10-01 12:53:42
15
View More