IT Brief UK - Technology news for CIOs & IT decision-makers
United Kingdom
xAI launches Grok 4.7 for coding & knowledge work

xAI launches Grok 4.7 for coding & knowledge work

Mon, 21st Sep 2026 (Today)
Sean Mitchell
SEAN MITCHELL Publisher

xAI has launched Grok 4.7 for coding and knowledge work.

The model is available through Cursor, the Grok API and other software tools.

According to xAI, Grok 4.7 succeeds Grok 4.6 and uses a larger base model. xAI trained the new version with a longer reinforcement learning run on a harder mix of tasks, placing more weight on problems that can take many hours to complete.

xAI said the model is designed to work longer on difficult tasks, verify its output more carefully and handle longer context windows. Grok 4.7 was also trained to understand the Grok Bot harness natively, which xAI said improves performance in conversational tasks and general knowledge work.

Benchmark results

xAI published benchmark comparisons against Grok 4.6, GPT-5.6 Sol and Fable 5.1 Max across software engineering, legal work, clinical reasoning, electrical engineering and office tasks.

On CursorBench 4.0, which measures longer-running coding tasks, Grok 4.7 scored 46.3%, compared with 40.4% for Grok 4.6 and 41.7% for GPT-5.6 Sol, but below Fable 5.1 Max at 51.8%.

In DeepSWE v1.1, Grok 4.7 posted 71.0% in a high-effort setting, compared with 65.2% for Grok 4.6, 72.7% for GPT-5.6 Sol and 70.0% for Fable 5.1 Max. On EEBench, which covers electrical engineering, the new model reached 64.0%, ahead of Grok 4.6 at 53.0%, GPT-5.6 Sol at 39.4% and Fable 5.1 Max at 56.4%.

xAI also pointed to gains in office and document-based work. On AA Briefcase v1.1, Grok 4.7 scored 1,657, compared with 1,546 for Grok 4.6 and 1,487 for GPT-5.6 Sol, while Fable 5.1 Max scored 1,678. On the Harvey Legal Agent Benchmark, Grok 4.7 scored 19.6%, versus 15.8% for Grok 4.6, 2.5% for GPT-5.6 Sol and 6.7% for Fable 5.1 Max.

On HealthBench Professional, Grok 4.7 scored 56.7%, above Grok 4.6 at 48.5% but below GPT-5.6 Sol at 60.5% and Fable 5.1 Max at 62.1%. In Terminal-Bench 4.0, the model scored 38.0%, ahead of Grok 4.6 at 20.3% and slightly above GPT-5.6 Sol at 37.3%, but below Fable 5.1 Max at 57.9%.

Safety focus

xAI said Grok 4.7 includes a new safeguard stack, and that internal testing showed stronger results on refusals and jailbreak resistance than earlier versions. The company also said the model performed strongly in dual-use areas such as cybersecurity and biological work, where systems must complete benign tasks without responding to harmful prompts.

According to xAI, Grok 4.7 reached 62.4% on LatchBio's biosafety benchmark. It also allowed only 3.3% of risky dual-use prompts through on HackerBench v0.3, while rarely blocking legitimate security work.

xAI has started giving select cybersecurity partners invite-only access to Grok 4.7 red-team functions for defence research. The move suggests the company is positioning the model not only as a coding assistant but also as a tool for controlled security testing.

Pricing and rollout

The model is priced from USD $2 per million input tokens and USD $6 per million output tokens. xAI also offers a faster variant at USD $4 per million input tokens and USD $12 per million output tokens, and has set that version as the default.

Grok 4.7 is being distributed through Cursor on desktop, web, iOS, command-line tools and software development kits, as well as the Grok API and third-party model routers and cloud platforms. The broad rollout places it in both direct developer workflows and external distribution channels, where buyers often compare models on price, task completion and reliability.

The launch comes as AI model providers compete more aggressively on coding and professional task performance while facing closer scrutiny over safety controls. xAI's figures show Grok 4.7 outperformed its predecessor across the disclosed benchmarks, though in several categories it still trailed a rival.

Those mixed results reflect a market in which vendors increasingly frame new releases around trade-offs between cost, speed, benchmark performance and safeguards, rather than claiming clear overall leadership. For xAI, the immediate test will be whether developers adopt Grok 4.7 in enough volume to validate its pricing and its effort to expand beyond headline chatbot use into day-to-day technical and professional work.

Based on the company's figures, Grok 4.7 matched Grok 4.6 on entry pricing while posting higher scores in software engineering, electrical engineering, legal work, office tasks and clinical reasoning.