Artificial intelligence models, competition shifts to inference pricing
OpenAI cuts the cost of GPT-5.6 Luna, Anthropic prices Claude Opus 5 at half its flagship model, DeepSeek raises the price of V4 Flash by 93%. In the same landscape, Alibaba releases Qwen 3.8 under an open licence and Google presents a compiler for inference on encrypted data.
The morning edition did not cover the Tech section. The fact that justifies this afternoon piece is a shift in the ground on which model makers compete: no longer benchmark scores, but the price paid per million tokens processed and the speed at which the model responds.
Recent movements point in two opposite directions. OpenAI has substantially cut the price of GPT-5.6 Luna, now the default setting for free ChatGPT users. Anthropic has priced Claude Opus 5 at roughly half the price of its top-tier model, Fable 5. In the opposite direction moved China’s DeepSeek, which on August 14 raised the price of V4 Flash by 93%, from 0.14 to 0.27 dollars per million tokens processed (Tech Startups; AIToolsRecap).
| Provider | Move | Figure |
|---|---|---|
| OpenAI | Price cut for GPT-5.6 Luna | default model for free users |
| Anthropic | Positioning of Claude Opus 5 | roughly half the price of Fable 5 |
| DeepSeek | Price hike for V4 Flash (August 14) | +93%, from 0.14 to 0.27 dollars |
According to data cited by the Financial Times and picked up by the technology review, prices paid by customers for the leading US models have fallen significantly since mid-July. The purchasing criterion emerging from the same sources concerns the ratio between processing cost and work actually produced: a parameter that favours efficient models over simply larger ones.
Why a provider raises prices
DeepSeek’s price hike is the anomaly that makes the rest legible. In a market where the main competitor is cutting prices, raising the list price of a fast variant by 93% means either that demand for that specific configuration exceeds available capacity, or that the previous price was set below cost. The available sources do not allow establishing which of the two hypotheses holds, and this newspaper does not choose between them.
Open licences and encrypted data
The same period saw two releases of a different nature. Alibaba’s Qwen group published Qwen 3.8 27B under an Apache 2.0 licence, a model with integrated vision that claims a native context of 262 thousand text units, extendable up to one million, on a Gated DeltaNet and Gated Attention architecture. The FP8 variant reports 61.7 on SWE-Bench Pro, 73.0 on Terminal-Bench 2.1 and 90.3 on LiveCodeBench v6: these are figures published by the developer itself, therefore claims made by the interested party and not independent verifications.
Google presented HEIR, an open-source toolchain that converts already-trained models to run inference directly on encrypted data, without the server seeing the inputs in the clear. The demonstration involved a recommendation system and a credit card fraud detection case.
On these two releases, the news currently comes from a single source (AI Weekly, which picks up the Qwen group’s publications on Hugging Face and Google’s official blog); no independent confirmation available.
The verifiable point is the price list: the price per million tokens processed is public for all the providers mentioned, and the comparison with mid-July values can be redone in a month on the same price lists.
← Archive · Front page · Past editorials · Report an error · Original article (in Italian)