Artificial intelligence

SpaceXAI Unveils Grok 4.7 for Programming and Knowledge Work at the Previous Model’s Price

SpaceXAI announced Grok 4.7 as its most powerful model for programming and knowledge work, while maintaining the price and speed of Grok 4.6. The model outperforms some models in programming, engineering, and legal tests, but trails competing models in terminal tasks and clinical reasoning.

2026-09-22
4 min read
77 views
certi.news
SpaceXAI Unveils Grok 4.7 for Programming and Knowledge Work at the Previous Model’s Price

SpaceXAI announced the Grok 4.7 artificial intelligence model on September 21 local time, presenting it to businesses and developers as its most powerful model designed for programming and knowledge work. The model is available at the same price and speed as Grok 4.6, while the company says it is twice as fast and half the cost compared with models in the same category.

What changed in Grok 4.7?

The model is based on a larger foundational model, with longer reinforcement-learning training and composite tasks designed to address problems that may take hours to complete. According to SpaceXAI, this improved the model’s ability to review its work and handle long contexts. It was also trained on an original understanding of how the company’s software agent, Grok Bot, operates, with the aim of improving its performance in conversational tasks and general knowledge work.

The improvements include creating documents and presentations, in addition to performing specialized tasks that emulate the work of lawyers, nurses, and financial analysts. In the GDPval test, Grok 4.7 recorded an Elo score of 1695, ahead of Grok 4.6 at 1605 and GPT-6 Astra at 1542, but it remained behind Fable 5.1, which scored 1735.

Advanced results in some tests, but not all

In CursorBench 4.0 for long programming tasks, Grok 4.7 scored 46.3%, compared with 40.4% for Grok 4.6 and 41.7% for GPT-5.6 Sol, while Fable 5.1 led with 51.8%. The model also achieved the highest score among the compared models in EEBench for electrical engineering, at 64.0%, and in the Harvey Legal Agent Benchmark for legal tasks, at 19.6%.

However, the advantage was not universal; in Terminal-Bench 4.0 for long terminal tasks, it scored 38.0%, compared with 57.9% for Fable 5.1. In HealthBench Professional for clinical reasoning, it reached 56.7%, behind GPT-5.6 Sol at 60.5% and Fable 5.1 at 62.1%. These results mean that model selection will remain tied to the type of task, not to overall ranking or price alone.

Price and availability

The price of the Grok API starts at $2 per million input tokens and $6 per million output tokens. By comparison, GPT-5.6 Sol costs $4 for input and $20 for output, while Fable 5.1 costs $10 for input and $50 for output. SpaceXAI offers a faster version with twice the speed of the basic version, while also doubling the price.

Grok 4.7 became available the same day through Cursor and the Grok Build programming tool, in addition to the Grok API and third-party programming tools, model routers, and cloud computing platforms.

Safety and access to testing tools

SpaceXAI said it had completely redesigned its safety mechanisms and improved resistance to attempts to circumvent restrictions and to dangerous-request refusals. The model scored 62.4% in the LatchBio biological safety test, the highest result among the models compared by the company. In HackerBench v0.3, the model passed only 3.3% of dual-use and cyber-risk requests, while the company said it does not block the majority of legitimate security tasks.

SpaceXAI has also begun granting some cybersecurity partners invitation-only access to red-team functions for defensive research purposes. These claims remain tied to the tests selected by the company and its methodologies, while the variation in results across fields requires development and business teams to evaluate the model according to actual use cases.

News source
ITmedia NEWS Japan
Open original source ↗
c
Author

certi.news

In the same category

You may also like

View all news