Parviz Sayed Mohammed believes that revealing the identity of the mysterious Ox Alpha model illustrates an important shift in the AI market: comparisons between models are no longer focused solely on the highest level of intelligence, but on how much intelligence an organization gets for each dollar. On August 26, Z.ai revealed that the model that appeared on OpenRouter was GLM-5.3-Flash, and that it had been deliberately run in front of public traffic before its identity was announced.
The model attracted the attention of enthusiasts and independent developers because it delivered good performance at no cost initially, resulting in several trillion tokens being processed daily, while community estimates of usage during one week ranged from single-digit figures to more than 20 trillion tokens. More importantly, according to the article, the service relied entirely on Chinese chips and infrastructure, not on a U.S. laboratory as some observers guessed during the six days preceding the disclosure.
Price Reorders the Comparison
The announced price for GLM-5.3-Flash is 15 cents for input and 50 cents for output per million tokens. OpenRouter is offering a 50% discount through September 9, bringing the price to 7.5 cents for input and 25 cents for output. The model's weights are also open under an MIT license, while inference is hosted by entities including Z.ai, GMI Cloud, Cloudflare, and other providers in the United States.
According to the Artificial Analysis comparison, in which the model was listed the same day, GLM-5.3-Flash scored 57 on the intelligence index at a cost of about nine cents per task. By contrast, GPT-5.6 Sol (max) achieves a score of approximately 59 at 67 cents per task, while Grok 4.6 scores 61 at 94 cents. The author calculates that just a two-point difference between the first two models corresponds to a price roughly 7.4 times higher, while a difference of approximately four points requires nearly ten times the cost.
These figures are not a general judgment of model quality, nor do they replace each organization's own testing, but they illustrate the article's main argument: when models draw close to one another on performance indicators, inference cost can become a decisive factor in determining usage volume, particularly in programming, agents, and repetitive tasks.
Budget Pressure Comes Before the Decision to Abandon AI
The author cites Uber to illustrate the problem. The Information reported in April that Uber CTO Praveen Neppalli Naga said the expected annual programming budget for 2026 had been consumed within four months, and that a two-hour trial demonstration had personally cost him $1,200. By June, Uber had imposed a cap of $1,500 per person and per tool.
This does not mean that the tools are useless; the article distinguishes between “usefulness” and “value.” The author reports that Andrew Macdonald, Uber's chief operating officer, was unable to link dashboards to a 25% increase in the number of features useful to consumers. The article also cites the results of McKinsey's 2026 survey, according to the source, which found that 80% of respondents said they had become faster, that 37% of companies had seen an impact on earnings before interest and taxes, while 32% of companies or more had avoided purchasing at least one software program because they were able to build the feature internally using coding agents.
The editorial reading here is that organizations face a dual equation: they want to reduce the bill, but they cannot simply stop investing in AI. The discussion therefore shifts from the question “Should we use AI?” to a more practical one: Which model should each team use, for what type of task, and according to what measure of success?
Three Tiers Instead of One Model
Mohammed proposes distributing programming and agentic-work tasks across three tiers, calculating the share of tasks and tokens rather than the share of financial spending. At the top tier are the Fable and Opus models for irreversible decisions, strategy analysis, and detailed execution plans; the author estimates that these tasks account for no more than 5% of the workload.
The middle tier includes Kimi K3, Gemini 3.7 Flash, GPT-5.6 Sol, and Grok 4.6, all of which, according to the source, are around a score of 60 on the intelligence index. He proposes allocating about 50% of task volume to them, particularly for daily programming and routine work. He notes that Kimi is popular for programming, while Grok 4.6 is close behind, although its smaller context window remains a limitation.
The remaining 45%, in the author's proposal, would be allocated to GLM-5.3-Flash as a model for high-volume work. But this percentage is not a fixed formula; the tool used to run the models, the mix of tasks across programming, content, and marketing, and the results of internal evaluations should determine the actual allocation.
What Should Be Measured Before Rebuilding the Budget?
The article concludes that open-weight Chinese models, such as those from Zhipu, Qwen, DeepSeek, and others, should be explicitly included in cost calculations, rather than excluded merely because of their origin. It notes that the share of Chinese models' tokens on OpenRouter exceeded that of U.S. models in early June, and that the top of the list remained largely occupied by Chinese laboratories.
Before the wave of releases expected from Google, xAI, Anthropic, OpenAI, and DeepSeek in September, the author proposes tracking token consumption and linking spending to a clear indicator, such as customer or revenue growth, or at least development speed and productivity. He also calls for preparing an AI budget at the level of each organization or team, then defining a model policy in tiers: a high-end model for strategic decisions, a mid-range model for paid seats and daily programming, and a low-cost model for intensive, repetitive tasks.
The author's viewpoint does not prove that GLM-5.3-Flash will be the best choice for every organization, nor that general comparison metrics reflect performance in every context. But it offers a practical standard for review: not every request should receive the most expensive model available, just as the cheapest model should not be adopted without clear evaluations and value metrics. The open question the article leaves is whether cost-performance curves will continue in this direction after the new releases arrive, and whether AI companies capable of lowering inference costs will actually capture most of the usage volume.