Elon Musk announced that Grok Bot will choose the most suitable artificial intelligence model for each task instead of relying permanently on Grok. Musk wrote that SpaceX will use the “best backend model” to achieve the best result, and mentioned Claude Opus 5.5, MidJourney, Suno, and other interfaces among the potential options.
Grok Bot is a joint product of SpaceXAI, Musk’s artificial intelligence lab that was merged into SpaceX, and Cursor. SpaceX had agreed in June to acquire Cursor for $60 billion in stock, and the deal closed in August. Grok Bot was launched in beta during the same month.
Model routing has become a way to reduce costs
The article describes this approach as “model routing”: sending each request to the model best suited to the type and cost of the work, while reserving the more expensive models for more complex tasks. According to interviews conducted by the writer with leaders of 22 companies at Insight Partners’ ScaleUp:AI event, at least ten of them said they follow this approach.
In one example, a security company uses a rules engine to screen data first, then sends the remainder to small models. One of its executives said that passing a petabyte of data through any model, even a small one, could cost millions of dollars. In another example, a low-cost model is used to understand the user’s request before referring it to a more expensive model to carry out the task.
The main motivation is to control spending on tokens. Some companies initially began with open-ended artificial intelligence budgets, then returned to a more precise question: What did they actually buy in exchange for all those tokens? Other companies are also trying to train employees to choose the appropriate model instead of automatically using the most powerful one.
Cheaper models handle most of the volume
OpenRouter data, a service that provides access to hundreds of models through a single connection, reflects this trend. In the week ending October 7, 2026, four low-cost, fast Flash models were among the ten most-used models by token count. DeepSeek V4.1 Flash topped the list with 33.6 trillion tokens, followed by GLM 5.3 Flash with 10.5 trillion and MiMo-V2.6-Flash with 10.1 trillion.
By contrast, use of Claude Opus 5.5 grew by 74% in one week to reach 3.37 trillion tokens, but it remained ninth by volume. The article explains this by saying that the advanced model is receiving an increasing share of tasks without handling most usage traffic.
The results cited in The New Stack tests indicate that cheaper models may be sufficient in many cases: Claude Sonnet 5.5 passed all 15 programming tests at a cost 42% lower than Opus 5.5, while GPT-6.1 Sol matched GPT-6 Astra’s accuracy at a cost equivalent to 18% of it.
What is changing in practice?
The shift does not mean that the most powerful model has become unnecessary; rather, choosing it has become an operational decision requiring guidance and testing. The article warns of the risks of switching quickly: one engineering team failed after replacing one model with another before an important presentation, then had to roll back the change. The conclusion was that evaluation tests should be used before switching models.
OpenRouter data remains an indicator rather than a complete measure of the market, because it reflects users of a service who already chose to use a model router. Token counts also do not measure the number of users or spending, and enterprise contracts may be more stable. However, the combination of this indicator with the practices of the companies covered in the interviews and the announcement from Grok Bot makes clear that loyalty to a single model is declining in favor of an architecture that chooses the model according to the task and cost.