Researchers from Google and Google DeepMind announced the establishment of DeepMind Institute, aiming to advance discussions surrounding general artificial intelligence (AGI) from principled statements to more concrete aspects of security, evaluation, and policy design. The organization is governed by DeepMind, co-founder Shane Legg, Google executive James Manyika, and Google DeepMind leader Demis Hassabis serving as directors.
First release of four articles
The new organization states that its goal is to allow for open discussions about the different perspectives on AGI among Google, Google DeepMind, and the broader research community, rather than maintaining a single standpoint. The first four articles published cover topics such as how to address the potential economic impacts of AGI, preserving the readable reasoning process of models, principles for promoting human welfare, and an evaluation framework for cutting-edge AI models.
This means that the relevant discussions are shifting from general security concerns to more actionable institutional designs. The focus of the article is not only on risk assessment but also begins to address what disclosure obligations developers should undertake, as well as how external institutions can intervene in the review process.
Focusing on Model Inference Visibility
One of the articles was written by security researchers Rohin Shah and Anca Dragan. The core argument is that it is becoming increasingly difficult for outsiders to observe the model inference process, but this is not an irreversible trend. The two believe that as new architectures make the strongest models more difficult to monitor, developers and regulators need to actively address the trade-off between performance and supervisability.
The directions proposed in the article include restricting the depth of continuous calculations by models without leaving any readable traces of reasoning, or requiring developers to prove that even with a decrease in system transparency, external monitoring can still be carried out to the same extent.
Hassabis Proposes a cutting-edge model evaluation institution
In another article, Hassabis proposed the establishment of a leading AI standard organization led by the United States, dedicated to evaluating the most advanced AI models. According to this concept, developers would initially have up to 30 days before the model is released to voluntarily submit it for review; if this evaluation mechanism proves effective, it may become a prerequisite for deploying leading models in the United States in the future.
In the initial stage, this institution will jointly design evaluation methods with AI company, and then gradually transition to independent testing without disclosing the content in advance, in order to prevent the laboratory from optimizing the model based on known topics. Hassabis also mentioned that if the risk situation further escalates, additional measures can be taken, and it is not ruled out that coordination among cutting-edge AI developers may be slowed down.
Industry discussions have shifted to specific solutions.
At the time of the release of these articles, discussions in the AI industry regarding security are undergoing changes. In the past, principle-based warnings were common; nowadays, more focus is being placed on information disclosure, external audits, and the control of development pace when necessary.
This shift further intensified this week due to the public support from several industry insiders for some of the proposals put forward by Anthropic's Chief Executive Officer, Dario Amodei, regarding "controlling the pace of AI development." The launch of DeepMind Institute also made Google's stance in this discussion more concrete.












