Skip to content
AI.info

The Pulse

SpaceXAI launches Grok 4.7 for coding and knowledge work

SpaceXAI’s September 21, 2026 announcement presents Grok 4.7 as its most capable model for coding and knowledge work; the source was checked for an attributable speaker and contains no named quotation.

SpaceXAI launches Grok 4.7 for coding and knowledge work

AI.info Team ·

SpaceXAI targets longer software tasks

SpaceXAI launched Grok 4.7 on September 21, 2026, positioning the model as a system for coding, agentic tasks and professional knowledge work rather than a general chatbot update. The company says the model is available in Cursor and Grok Build, through the Grok API, and via third-party coding harnesses, model routers and cloud platforms.

SpaceXAI describes Grok 4.7 as its most capable model for coding and knowledge work. The release says the model works longer on difficult tasks, checks its own work more carefully and includes the company’s best-calibrated safeguards to date. It is served at the same standard price and speed as Grok 4.6.

The company says Grok 4.7 uses a new, larger base model than Grok 4.6 and a longer reinforcement-learning run built around a harder mix of tasks, with greater weight given to problems that can take many hours to complete. SpaceXAI also says the model is better at verifying its own work and managing longer context. It was trained to understand the Grok Bot harness natively, which the company says improves its performance on conversational tasks and general knowledge work.

That focus creates a clear test for the launch: SpaceXAI is presenting Grok 4.7 as a more dependable worker on extended assignments, while its published results show gains over Grok 4.6 but do not place the model first on every comparison. The company’s table puts Fable 5.1 ahead on CursorBench 4.0 and Terminal-Bench 4.0, while GPT-5.6 Sol scores higher on DeepSWE v1.1 and HealthBench Professional.

Grok 4.7 posts gains over Grok 4.6

On CursorBench 4.0, which SpaceXAI describes as a test of longer-running coding tasks, Grok 4.7 scores 46.3%, compared with 40.4% for Grok 4.6 and 41.7% for GPT-5.6 Sol. Fable 5.1 records 51.8% on the same test.

The model shows a larger improvement on Terminal-Bench 4.0, rising to 38.0% from Grok 4.6’s 20.3%. GPT-5.6 Sol scores 37.3%, while Fable 5.1 reaches 57.9%. On DeepSWE v1.1, Grok 4.7 scores 71.0% in the high-effort configuration, compared with 65.2% for its predecessor; GPT-5.6 Sol leads that comparison at 72.7%.

SpaceXAI also reports a score of 1,657 on AA Briefcase v1.1, a benchmark for multi-hour office work. Grok 4.6 scores 1,546, GPT-5.6 Sol scores 1,487 and Fable 5.1 scores 1,678. The results cover work associated with professionals such as lawyers, nurses and financial analysts, including the creation of documents and presentations.

On additional comparisons, Grok 4.7 scores 64.0% on EEBench, compared with 53.0% for Grok 4.6, 39.4% for GPT-5.6 Sol and 56.4% for Fable 5.1. The model also scores 19.6% on the Harvey Legal Agent Benchmark, ahead of Grok 4.6’s 15.8%, GPT-5.6 Sol’s 2.5% and Fable 5.1’s 6.7%.

Pricing and availability

Grok 4.7 starts at $2 per million input tokens and $6 per million output tokens. SpaceXAI says the model is available today in Cursor and Grok Build, as well as through the Grok API, third-party coding harnesses, model routers and cloud platforms.

The company also offers a fast variant with twice the output speed at twice the price. SpaceXAI directs developers to its API console and documentation for access and technical details. The release provides a command for installing the company’s command-line tool and directs users to Grok Build to get started.

SpaceXAI adds a new safety claim

The launch includes a new safeguard stack for Grok 4.7. SpaceXAI says the model is the strongest it has tested on refusals and jailbreak resistance. In dual-use areas such as cybersecurity and biological work, the company says the model combines utility for benign tasks with safer refusal behavior on dangerous requests.

On LatchBio’s biosafety benchmark, SpaceXAI reports a score of 62.4%. The company also says Grok 4.7 allows 3.3% of risky dual-use prompts through on HackerBench v0.3, its benchmark for risky and malicious cyber tasks, while rarely blocking legitimate security work.

SpaceXAI says selected cybersecurity partners have received invite-only access to Grok 4.7’s red-team capabilities for defense research. The published results are self-reported in the launch announcement, which does not identify the participating partners or provide an independent audit of the benchmark results.

For developers, the immediate change is access to a model aimed at longer coding and knowledge-work assignments, with a starting price of $2 per million input tokens and $6 per million output tokens. Grok 4.7 can be accessed through the company’s API and coding products under the model name grok-4.7.

Source

SpaceXAI

Explore

More articles