
Optimize AI investment by measuring business results and total costs rather than token unit price
AI-generated summary
OpenAI presented a new framework for measuring the effectiveness of AI deployments, highlighting the performance and economics of its latest model, GPT-5.6.
First, OpenAI explained that the value of AI cannot be measured solely by the price per token. Even with a low-priced model, the total cost may increase if the number of trials, time required, and human confirmation work increase. Even if a model with high performance has a high unit price, it may be possible to reduce the total cost if the objective can be achieved the first time. In evaluation, it was determined that it was necessary to compare the total cost of successfully completing the task and the value generated by the outcome.
The first indicator was useful work results. Actual results were measured, such as the number of customer interactions resolved, number of code fixes, number of contract reviews, time saved for people, and improved decision-making. He explained that value is created not in the number of tokens, but when it is converted into results that can be used in business. In the early stages of implementation, we recommended a method of selecting one business procedure, clearly defining the conditions for completion, and then measuring the results using the system in which the business is actually executed.
The company says that completion conditions vary by division. The results include problem solving in customer service, code modifications that pass testing in the development department, and on-time and accurate contract confirmation in the legal department. The finance department gave examples of work performed before a forecast review, such as collecting data, posting to spreadsheet software, checking differences, checking numerical values, and updating data. He explained that as the business service "ChatGPT Work" handles many processes, people can allocate their time to tasks that require judgment, such as the cause of change and the next response plan.
The cost per successful transaction was used as a second indicator. AI processes a wide range of topics, and while short responses require fewer computational resources, program development, research, and finance-related work may require deep inference, the use of multiple tools, and a large amount of processing. Accordingly, there is a possibility that the value created will also increase.
A method was shown in which costs are calculated based on the total cost, which includes not only the model price, computational resource usage, and probability of reaching the correct result, but also personnel costs, confirmation work, retries, and correction work. The basic calculation was to tally the number of jobs that met the required quality and divide the total cost by that number. He explained that even a model with a low token unit price does not necessarily mean that the total cost will be low; if a high-performance model can produce results in one go, the total cost may be reduced.
The company has prepared three stages of "GPT-5.6" released on July 9, 2026 (local time): Sol, Terra, and Luna. Sol is positioned at the top, Terra is positioned as a balance between performance and cost, and Luna is positioned as high speed and low price. We envisioned choosing Luna for work that requires high-speed processing, Terra for deep analysis, and Sol for advanced inference, depending on the business content (editor's note: On July 30, 2026, OpenAI reduced the price of GPT-5.6 Luna by 80% and the price of Terra by 20%).
GPT-5.6 is said to have been developed with the aim of increasing the results obtained from each token. In the Artificial Analysis Coding Agent Index, GPT-5.6 Sol achieved a new record at maximum inference settings, reducing the number of output tokens by 54% compared to other leading models. GPT-5.6 Sol recorded 72.7% on DeepSWE v1.1, surpassing Anthropic's "Claude Fable 5". The estimated API cost was 36.2% lower. The goal was to increase results per dollar across the entire model group.
Reliability was cited as the third indicator. The company indicated that the use of AI will gradually expand. He explained that the role will begin with document creation assistance, expand to information gathering and use of multiple tools, and eventually reach the stage of being responsible for the entire business, including exception handling. The position is that the wider the scope of use, the greater the value of accuracy and stability. The evaluation items were categorized into three categories: results that were usable immediately after delivery, results that required manual correction and retrials, and results that were handled by a person until the end. He explained that it is possible to grasp practical value that cannot be grasped by model accuracy alone.
In order to ensure reliability, it was also necessary to clarify the scope of use. Before AI can take action, it needs to define what data is available, what systems can be modified, and when human review and approval is required. He explains that it is essential to have an environment that is based on safety, security, privacy, and management functions, and allows users to understand AI operations, data processing methods, and operational rules. ChatGPT Work inherits the various infrastructures of "ChatGPT Enterprise" and allows access to business information and procedures while maintaining management functions.
The fourth indicator was economic efficiency when expanding the scale of use. We showed how to continuously measure the same work, tracking the number of cases that met quality standards, total cost, and cost per successful case. He expressed the belief that the value of AI investment will increase if the number of results grows faster than the total cost and the quality is maintained or improved.
The company positioned computing resources as the foundation for research and development and service provision. He explained that computational resources for research will lead to improved capabilities in the future, and computational resources for inference will play a role in supporting current business results. The company says that improvements in model performance, inference efficiency, dedicated hardware, utilization rates, processing path optimization, and product design will increase the effectiveness of computing resource investments. The research results are reflected in new models, and the improvements lead to improved product quality, expanded usage, and profits, creating a cycle that supports investment in next-generation research and safety measures.
The company supported its generative AI tools "ChatGPT", ChatGPT Work, "Codex", and APIs on a common intelligence platform, and demonstrated a structure in which improvements to the platform would spread to all products and users. Through a four-item evaluation method, we comprehensively measure the results produced by AI, the cost to success, reliability, and the value of expanding its use, with the goal of creating an environment where people can allocate their time to tasks that make use of their judgment and creativity. The company announced its policy of continuing to improve model capabilities, response speed, reliability, and reduce operational costs, aiming to realize AI that is useful to many users and organizations.

加齢による認知・身体機能の衰えでスマートフォンを諦める「スマホリタイア」が確認され、80代の節目を中心に約6人に1人に及ぶと推測される。離れて暮らす高齢の親との連絡手段として、テレビ電話サービスなどの代替策が注目されている。
デジタルコマースは、成人向けAI作品の生成・公開やファンとの接点作りができる新サービス「FANZAスタジオ」の先行体験を24日から開始すると発表した。韓国Onoma AIがパートナーとして協力する。
東京都とGovTech東京は8月20日、防災や暑さ指数など10種類の都民向け地図情報を1つに集約・組み合わせ表示できるWebサイト「Tokyo Map」の正式版を公開した。PCやスマホから無料で利用可能。
LINEヤフーは8月20日、LINEアプリの着せかえの一部でメニューアイコンがデフォルトになる不具合について謝罪し、経緯と対応を発表した。iOSの新デザイン仕様への適応が原因で、年内から9月上旬にかけて順次自動変換を行うほか、対象ユーザーへの返金も受け付ける。
SNS上の情報をAIで収集・可視化するスペクティの「Spectee Pro」について解説。災害時の情報源として自治体や報道機関で活用されており、デマや偽情報を排除する仕組みで危機管理を支援している。

Black Kiteの報告書によると、2023年から2026年にかけて発生したランサムウェア攻撃の72%が北米の中堅企業を標的にしており、特に規模の小さい企業が大きなリスクにさらされていることが明らかになった。