GPT-6 Sol and Luna filling in the mid-to-low price range: What's reduced is the unit price of the model, while workflow costs will be calculated separately.
CoinMeta
09-29 11:16
Ai Focus
When a team has to process tens of thousands of tickets per day with AI, the most expensive ones are not necessarily the most difficult ones; rather, they are the numerous requests that seem simple but occur repeatedly. OpenAI has recently expanded its GPT-6 series to include Sol and Luna precisely for such tasks. The flagship Astra continues to undertake high-difficulty tasks, Sol focuses on complex coding and proxy work, while Luna leans towards high-frequency tasks that are cost-sensitive. OpenAI announced that compared to the previous promotional prices of the GPT-5.6 series, the unit price of the two new models API has been reduced by about half. This is a change in product pricing, but it does not mean that every company...
Helpful
No.Help

When a team has to process tens of thousands of tickets per day with AI, the most expensive ones are not necessarily the most difficult; rather, they are the numerous requests that seem simple but occur repeatedly. OpenAI recently expanded its GPT-6 series to include Sol and Luna precisely for such tasks. The flagship Astra continues to undertake high-difficulty tasks, Sol focuses on complex coding and proxy work, while Luna leans towards high-frequency tasks that are cost-sensitive. OpenAI announced that compared to the previous promotional prices of the GPT-5.6 series, the unit price of the two new models API has been reduced by about half. This is a change in product pricing, but it does not mean that every company's final AI budget will automatically be halved as a result.

The officially listed standard input/output prices are as follows: for every million token, the cost of GPT-6 and Sol is 2/10 US dollars, while for Luna it is 0.10/0.50 US dollars. There is a significant difference between the two prices, but we cannot simply assign all tasks to Luna based on this alone. If a task requires repeated use of tools, tracking of status across multiple files, or if the output must withstand manual review, running the lower-priced models several more times or even redoing the work could result in increased costs and time. For ordinary users, OpenAI has already provided both models as part of the corresponding payment plans for Work and Codex. Luna is also available for some free or Go users to try out on the desktop; however, the official statement also mentions that these models were not yet integrated into the regular Chat interface at that time, and that their rollout would be gradual. Confusing terms like “API available”, “Work available”, and “visible to all chat users” can easily lead to misunderstandings.

Why are there three different tiers for models of the same generation?

For some time now, discussions within the industry about models have often revolved around "who is at the top of the rankings." However, the selection of models by companies is more akin to arranging employees: some are adept at handling complex problems with unclear boundaries, some are suitable for repetitive tasks with clear rules, and others are responsible for finding a balance between speed and quality. OpenAI places Astra in the highest capability category in its official documentation, while Sol strives to make more powerful reasoning capabilities and tool usage accessible at scalable prices. Luna focuses on offering lower costs and faster daily processing. The coexistence of these three categories is due to the fact that real-world work does not come in a single level of difficulty.

OpenAI has published several tests, among which Sol achieved high scores in cross-application business process evaluations such as those conducted by AutomationBench. Luna has also made improvements over its predecessors in various coding and professional tasks. Tests are valuable in that they inform developers about which tasks are worth considering for inclusion in a candidate list. However, the prompts, tools, effort levels, and billing methods used in different tests are not consistent. The official team also notes that comparisons with some competitor models are based on public reports, and there may be differences between the research environments and actual products. Therefore, a single test result should not replace internal acceptance, nor can it be used to conclude that a new model has outperformed in all corporate scenarios.

The other side of price is caching. Long conversations and proxy tasks involve the repeated submission of the same background materials, rules, and historical steps. If the same prefixes can be cached, both the input cost and waiting time have the potential to decrease. OpenAI mentions that improvements have been made to the default cache hit rate for GPT-6, and diagnostic tools have been added for developers. There is a business issue that is easily overlooked: how much savings caching provides depends on whether the company's own workflow has reusable, stable prefixes; if each request is completely different, the savings seen in others' cases should not be directly applied. Those building systems need to consider both the cache hit rate, the total number of calls, and the quality of the final results, rather than just looking at the input unit cost.

When migrating a business, don't mistake being "smarter" for an excuse to avoid acceptance checks.

OpenAI indicates that Sol and Luna have improvements over their respective predecessors in terms of factual accuracy, coding, and computer operations. However, the internal assessment of "fewer factual errors" is based on dialogue samples where users had previously reported errors, and these dialogues do not represent the average error rate of daily requests. For sensitive tasks involving financial, medical, legal, or corporate internal data, it is still necessary to maintain source verification, access control, and manual review. While stronger models can reduce some mechanical labor, they do not automatically eliminate the need for a chain of responsibility.

Developers should be particularly vigilant about version differences that may arise during the upgrade process. The update logs of API for OpenAI also record that on September 25th, an issue with image encoding in Sol and Luna was fixed. Officials recommend that users who rely on image input to re-run evaluations and retry affected processes. Although this incident is separate from the release of the new model, it indicates that the production performance of the same model can also change after fixes are made. Teams responsible for image understanding and interface operations should not use the test scores on the release day as a permanent baseline; they should at least retain representative cases in key processes and re-test after model or platform updates.

A more prudent approach to deployment is to first stratify by task type: tasks that require in-depth judgment should be retained at the higher capability level; for tasks with clear boundaries and that can be automatically accepted, Sol should be considered; only when the quality standards are indeed met should a large number of repetitive tasks be handed over to Luna. Each layer should record the success rate, the rate of manual intervention, delays, and the total cost per task, not just the token expenditure. GPT-6, Sol, and Luna broaden the range of options available, and thus the true competition shifts from a mere performance ranking to whether a company can design its processes clearly enough.

Cover photography: Real photos of Steve Jurvetson taken by Sam Altman, Wikimedia Commons; cropped for editorial purposes; the old photos are not presented as being taken at the location of this release.

Tip
$0
Like
0
Save
0
Views 45
WalletJYS reminds readers to view blockchain rationally, stay aware of risks, and beware of virtual token issuance and speculation. All content on this site represents market information or related viewpoints only and does not constitute any form of investment advice. If you find sensitive content, please click“Report”,and we will handle it promptly。
Submit
Comment 0
Hot
Latest
No comments yet. Be the first!
Related
News reports that AMD Xielong 9006 series servers CPU are in high demand: production for 2027 has been sold out, with an estimated shipment of 6.75 million units.
Tech media Wccftech reports that the production quota for AMD's sixth-generation Ryzen EPYC server processors in 2027 has been sold out, and AMD has already begun accepting orders for production in 2028. Morgan Stanley estimates that shipments in 2027 could reach 6.75 million units, corresponding to revenue of nearly $51 billion.
The Block
·2026-09-30 11:22:47
3
Robbins LLP Reminds Taboola Investors to Pay Attention to Securities Class Actions Alleging Companies Overstate the Value of Their Relationships with Issuers
Robbins LLP reminds investors that a class action lawsuit has been filed against Taboola.com Ltd regarding securities purchased or acquired by investors between May 6 and August 4, 2026. The complaint alleges that Taboola misled investors regarding the value of its relationship with the issuer; after the company announced second-quarter revenue below expectations and lowered its full-year outlook on August 5, 2026, the stock price fell by 27.41% on that day.
PR Newswire
·2026-09-30 11:22:45
4
The report states that before AI got out of control, OpenAI had received warnings from employees, but the management chose to ignore them.
According to The New York Times, long before the artificial intelligence of OpenAI exhibited out-of-control behavior, two employees had already alerted senior management to security issues during the testing phase, but their warnings were not taken seriously. The report also states that OpenAI has recently been exposed to multiple security vulnerabilities and abnormal behaviors. The company has acknowledged that some of its protective measures have failed and has decided to postpone the release of the latest version of the model GPT-6.1 Astra.
The Block
·2026-09-30 11:12:42
10
Former Qualcomm and Apple engineers team up to develop WarpCore; prioritize core technology, claiming to rewrite the rules of the semiconductor industry
Tech media reports that Nuvacore has disclosed the "core-first" development approach for its WarpCore processor, which involves constructing most of the underlying microarchitecture first before determining the instruction sets such as x86, Arm, or RISC-V. The company claims that this approach allows the team to focus on performance, energy efficiency, and chip area initially, but ISA, licensing, and compatibility will still limit the final design.
The Block
·2026-09-30 11:12:41
10
IEHP Celebrates 30 Years of "Joint Care"
Inland Empire Health Plan ( IEHP ) announces the launch of a "30 Years, Hand in Hand for Care" anniversary campaign on World Heart Day to celebrate its 30-year journey in promoting access, safety, and quality of healthcare services in the Inland Empire region of California. They also introduce a community-oriented " IEHP 30 for 30" health challenge.
PR Newswire
·2026-09-30 11:03:18
12
View More