OpenAI New model; Astra Architectural adjustments trigger security controversies
Fortune
52m ago
Ai Focus
OpenAI New model Astra is said to adopt a more efficient architecture, and some security researchers are concerned that this may weaken the monitoring capability of the AI inference process.
Helpful
No.Help

The upcoming cutting-edge model Astra, which is about to be released, has sparked security discussions due to a new internal architecture design. The focus of the controversy is not on the capabilities of the model itself, but rather on how such designs may make it more difficult for researchers to understand how AI arrives at its conclusions when performing tasks.

Adopt a more power-efficient cyclic design

As previously reported by The Information, OpenAI uses " recurrent depth " in some of the internal structures of Astra, which is also known as " looped Transformers ". The characteristic of this method is that it allows a portion of the computation modules to be called repeatedly, rather than having each token go through the entire network completely.

The direct benefit of doing so is to reduce the consumption of computing power. Reports cite relevant research stating that such architectures can achieve results close to those of traditional Transformer with fewer computing resources in certain scenarios, which is particularly important for corporate clients that are already bearing high AI costs.

Security researchers are concerned about the increasing difficulty of monitoring.

The controversy arises from the visibility of the model's reasoning process. Existing reasoning models typically generate a series of intermediate steps, commonly referred to as a “thought chain.” Researchers can use this to examine why the model made a particular decision or whether there was any overstepping of authority.

However, in the cyclic Transformer, some of the intermediate reasoning is not written in natural language but is repeatedly processed within the model in a form that is difficult for humans to understand directly. Researchers sometimes refer to this type of intermediate representation as “neuralese”. This means that what humans can see may be only the final answer, rather than the complete reasoning path.

Some security experts believe that this will weaken one of the few current means available to monitor the behavior of AI proxies. A policy director at the Washington-based think tank AI Policy Network told Fortune that if OpenAI is reducing its reliance on monitoring thought chains, this direction "may be concerning" and could also lead to higher risks.

OpenAI claims to still retain monitoring capabilities

In the face of external doubts, OpenAI's chief scientist, Jakub Pachoki, stated on X that the company has always placed emphasis on monitoring thought chains and is also working hard to retain the relevant capabilities. He mentioned that the scope of use for this type of architecture by Astra is limited, and that the model reasoning remains readable. More details about the architecture will be disclosed in the future.

Pachoki also stated that monitoring thought chains in the future may indeed become more difficult, but the reason may not necessarily lie in this architectural change itself. He said that OpenAI is still investing resources to advance real-time thought chain monitoring, which is also one of the current research priorities.

However, some researchers are concerned not only about Astra itself but also about the industry's demonstration effect. The former OpenAI researcher Daniel Kokotajlo believes that even though OpenAI is currently used in a more restrained manner, other companies may continue to follow this path, ultimately making the model inference process completely opaque to humans.

Cost advantages coexist with industry precedents

The reason why this type of architecture has attracted attention is also that it possesses attractiveness at both the cost and competition levels. In addition to saving computational power, hiding some of the intermediate reasoning processes may also make it more difficult for large models to be distilled, meaning it is harder for smaller models to replicate them by learning the output results.

Against the backdrop of the U.S. government and multiple AI companies continuously paying attention to the risk of cutting-edge models being copied, this design may be seen by more enterprises as a solution that balances efficiency and protection against copying. However, security researchers are concerned that without a unified standard in the industry, even more efficient models in the future may become more difficult to audit and regulate.

Tip
$0
Like
0
Save
0
Views 16
WalletJYS reminds readers to view blockchain rationally, stay aware of risks, and beware of virtual token issuance and speculation. All content on this site represents market information or related viewpoints only and does not constitute any form of investment advice. If you find sensitive content, please click“Report”,and we will handle it promptly。
Submit
Comment 0
Hot
Latest
No comments yet. Be the first!
Related
U.S. stocks open higher and strengthen; seasonal risks in September remain a concern
U.S. stocks rose on Thursday, driven by a decline in U.S. Treasury yields, but oil prices, employment data, and the historically weak performance in September remain focal points for the market.
Coinpaper
·2026-09-04 02:47:50
18
Meta reduces the price of Muse Spark in exchange for user data
Meta launches a low-cost data exchange scheme for Muse Spark, attempting to obtain real usage records through discounts in order to improve the AI intelligent model.
TechCrunch
·2026-09-04 02:34:57
20
web3: Growth rate of Canton destruction rises, pressure on CC token supply eases
Canton Network Weekly destruction and issuance ratio rises to 0.72, CC Supply pressure eases; project adjusts reward mechanism and advances institutional settlement scenarios.
CoinPedia
·2026-09-04 02:34:51
20
web3 : OpenAI releases Astra and makes it available to paid users
OpenAI releases new model Astra, which will integrate paid subscriptions and API, and has sparked discussions due to issues with inferable monitorability.
TechCrunch
·2026-09-04 02:11:12
19
View More