On September 29th, OpenAI launched GPT-6.1 and Sol, just one week after the release of GPT-6, Sol, and Luna. The most prominent selling point of the new versions is that, according to the official statement, they are close to the more advanced GPT-6 and Astra in terms of proxy programming, computer operations, and professional work. Moreover, the standard input and output Token cost is about one-fifth of that of Astra. For developers, this is not just about adding another model to the list. The real question is whether, with the use of 6.1 Sol, the success rate, running time, and the number of tool calls can all be improved simultaneously.
The new model this time can already be called using API with gpt-6.1-sol. The standard price is $2 per million inputs and $10 per million outputs; the cache input cost is $0.10 per million. OpenAI states that the latter is 95% lower than regular inputs and even half lower than the cache input of the previous generation, Sol. The model has also begun to reach the user groups related to Plus, Pro, Business, Enterprise, and Edu within ChatGPT Work and Codex. However, officials also point out that it has not yet entered the regular Chat chat interface. Describing it as "already released" as "available for all ChatGPT chat users" would overestimate its current range of availability.
What's cheap is the unit price, but there are also operating costs for proxy tasks.
When looking at model prices, one cannot consider only the input Token. A workflow may involve repeated searches, web page openings, code executions, waiting for tools to return results, and then sending intermediate results back to the model. Even if each individual call is inexpensive, if it requires more attempts, error reruns, or manual reviews, the final cost can still increase. Therefore, OpenAI includes both "cost per task" and scores in their published materials, which is more informative than just announcing the Token price. When selecting a team, one should first fix a set of tasks, the same set of tools, and acceptance criteria, and then compare the completion rate, latency, and total consumption.
The official DeepSWE 1.1 test involves long-term software engineering tasks in real code repositories. OpenAI states that a score of Sol can be achieved in this test, with a cost of about one-fifth; compared to 6 Sol, the best score is 6.4 percentage points higher. This figure indicates progress in the evaluation, but it does not guarantee that every repository can replicate these results. Real-world engineering also includes test environment stability, dependency version management, permission design, and code review, and benchmark testing will not automatically solve these organizational issues for the team.
Professional document evaluation: GDP.pdf tests whether the model can understand tables, diagrams, and detailed annotations. OpenAI claims that under certain settings, the new model achieves higher scores at less than half of the task cost of Opus. AutomationBench examines multi-step business processes such as sales, marketing, operations, and finance, with a score of 6.1. Sol performs 2.2 percentage points better than Opus at a moderate level of reasoning, and 4.8 percentage points better than its predecessor Sol. These cross-model comparisons are affected by the evaluated versions, tool environments, and cost calculation methods, and can serve as clues for trials, but should not be directly used as procurement conclusions for all corporate processes.
The progress of computer operations is even more worth looking at separately. According to the official statements, under the OSWorld 2.0 offline task and with the highest level of inference, the performance of 6.1 Sol is 7 percentage points higher than that of 6 Sol, and it is only 2.1 percentage points away from Astra. The cost per task is approximately one-seventh of Astra. However, the “offline tasks” mentioned here are different from those in a real system: in a production environment, factors such as login processes, pop-up windows, network interruptions, and permission verifications can cause issues that the evaluation does not take into account. Just because the model can click on buttons correctly does not mean it can bypass user approval, especially for operations with significant consequences such as making payments, deleting data, or publishing content.
A strong benchmark is not a universal pass; it depends on which questions it will still fail on.
OpenAI also announced scientific process testing. Terminal - Bench Science covers tasks such as data analysis, simulation, and theorem proof. The new model scores more than twice that of its predecessors under the highest level of reasoning, with an average cost of about $5.47 per question; Astra still ranks first among the models listed by the official with a score of 68.1%. For research teams, this means that models with lower costs can undertake more exploratory work, but difficult problems require upgraded models and verification by experts. A scientific research answer is not considered complete just by 'running the script'; the source of data, methods, and reproducibility are still indispensable.
Factual tests also have their limits. OpenAI mentions that in a set of high-difficulty conversations where users had pointed out errors in the old model, the proportion of factual error responses from the 6.1 Sol with lower reasoning capability dropped from 11.4% in the previous version to 7.7%. These are deliberately selected error-prone samples, not the overall error rate in normal usage. Describing a reduction of "about 30%" as "the error rate is now only 7.7%" can mislead readers. At the same time, the system manages the 6.1 Sol according to network security " Critical ", biological and chemical capabilities " High ", and continues to use the same level of protection configuration as the Astra level. When the capabilities approach those of the top models, risk control does not automatically become less stringent just because the price has decreased.
There is another point that can easily cause confusion: OpenAI mentions that the Ultrafast speed tier will be available in the coming days, and it should not be confused with the standard API that has already been launched on the same day and collectively referred to as "all launched." Officials claim that this speed tier can offer up to eight times faster performance than the standard speed in Codex, but the price and applicable scenarios need to be verified according to the conditions at the time of launch. The most prudent approach for the technical team at present is to conduct small-scale A/B tests using representative tasks within their unit, recording the input for each task, tool costs, number of retries, completion quality, and the time required for manual review. The new model is worth trying, but being "five times cheaper" is just the starting point; what users are actually paying for is the total cost of getting things done correctly.












