← News

Official SpaceXAI release graphic with the text Grok 4.7 on a dark grainy background

21 Sep 2026

SpaceXAI

SpaceXAI ships Grok 4.7 for coding and knowledge work

SpaceXAI released Grok 4.7, calling it its most capable model for coding and long knowledge-work tasks. The company says the model uses a larger base than Grok 4.6, longer reinforcement-learning training on multi-hour problems, stronger self-checks, and a new safeguard stack — at the same list price as 4.6.

SOFTWARE desk — another frontier lab shipped a same-day coding-agent model at commodity list pricing, keeping the agent-coding price war live while safety stacks get marketed as product features.

SpaceXAI calls Grok 4.7 its most capable model for coding and knowledge work. It says the model works longer on difficult tasks, checks its own work more carefully, and comes with the company’s best-calibrated safeguards so far. The same page says it is served at the same price and speed as Grok 4.6, and that it is highly competitive in its class. File that most-capable / longer-work / self-check / same-price-and-speed picture as SpaceXAI’s. “Most capable,” “best-calibrated,” and “highly competitive” are the company’s words, not a ranking this desk measured.

The new model uses a new, larger base than Grok 4.6. A base model is the underlying system, before the extra training that teaches it to finish jobs. SpaceXAI says that extra training was a longer reinforcement-learning run on a harder mix of problems, weighted toward tasks that take many hours. Reinforcement learning, or RL, means the model is scored on completed tasks and steered toward better scores. The company also says Grok 4.7 is better at verifying its own work and at managing longer context — keeping a long job straight. That is a skill claim, not a bigger window than Grok 4.6. File the larger-base / longer-RL / multi-hour / self-check picture as SpaceXAI’s. No parameter count, the learned-weight size, was printed. This desk did not count one.

List price starts at $2 per million input tokens and $6 per million output tokens, the same starting list as Grok 4.6. A token is a small chunk of text, a few characters or a short word. Input is what you send. Output is what the model writes back. Per million tokens is the company’s billing unit. The company also serves a fast variant with twice the output speed at twice the price. File the $2 / $6 start and the 2× speed / 2× price fast variant as SpaceXAI’s. The live docs price table shows the same $2 input and $6 output for prompts under 200,000 tokens, for both Grok 4.7 and Grok 4.6. Once a prompt reaches 200,000 tokens, that table bills every token in the request at $4 input and $12 output. Cached input — a repeat of text the service already stored — is listed at $0.50 under that line and $1.00 above it. The same table lists a 500,000-token context window for both models. File that long-context schedule as the docs table’s. “Same price as 4.6” still holds there. This desk did not receive a bill.

SpaceXAI says Grok 4.7 is available today in Cursor and in Grok Build, and also through the Grok API, third-party coding harnesses, model routers, and cloud platforms. Cursor is a coding app. Grok Build is the company’s coding agent. An API is a hosted call over the internet. A harness is the software that runs a coding agent around the model. A router is a service that picks which model answers a request. File that availability list as SpaceXAI’s. This desk did not open Cursor or start a Grok Build session.

On the company’s CursorBench 4.0 chart — a test SpaceXAI says stresses longer-running coding tasks — Grok 4.7 xHigh is listed at 46.3% and Grok 4.6 High at 40.4%. xHigh and High are effort settings on that chart, not a second model. File those two percentages as company-reported. This desk did not rerun CursorBench. The same chart lists other labs’ models on selected benches only. Do not read any column as Grok beating Claude or GPT overall. The rest of that company chart stays in Sources.

SpaceXAI says Grok 4.7 was built with an entirely new safeguard stack — the checks meant to refuse dangerous requests — and that it is the strongest model the company has tested on refusals and on jailbreak resistance. A jailbreak is a prompt written to talk a model past a refusal. On dual-use work, tasks that can help ordinary research or cause harm, the company says the model leads both on useful benign tasks and on refusing dangerous ones. Company-reported figures: it tops LatchBio’s biosafety benchmark at 62.4%. Biosafety here means a test of whether the model helps with dangerous biology. On HackerBench v0.3, which the company calls its benchmark for risky and malicious cyber tasks, it says the model allows only 3.3% of risky dual-use prompts through and rarely blocks legitimate security work. The company also says it has started giving select cybersecurity partners invite-only access to red-team runs of Grok 4.7 for defense research. A red-team run is a deliberate attempt to break the safeguards. File the new stack, the 62.4%, the 3.3%, and the invite-only line as company-reported. This desk did not run LatchBio or HackerBench.

Do not treat the company chart as an independent ranking. Do not claim Grok 4.7 beat Claude or GPT overall. The newsroom compares selected benches only. Do not invent a parameter count, a speed this desk timed, or a third-party leaderboard. None of those was an independent result on the primary page.

Plain English for the rest of the card: SpaceXAI = the name on this x.ai newsroom post. Grok 4.7 = the new model. Grok 4.6 = the previous model the company uses for the price and score comparison. token = a small chunk of text used for billing. input / output = what you send / what the model writes. reinforcement learning / RL = training that scores finished tasks and steers the model toward better scores. context = how much of the job the model can hold at once. Cursor = a coding app. Grok Build = SpaceXAI’s coding agent. API = a hosted call over the internet. harness = software that runs a coding agent. safeguard stack = the checks meant to refuse dangerous requests. jailbreak = a prompt written to bypass a refusal. dual-use = a task that can help ordinary work or cause harm. company-reported = a number from SpaceXAI’s own chart, not a test this desk reran. This filing is the Grok 4.7 release. It is not an independent bake-off.

PRIMARY here: SpaceXAI’s 21 Sep 2026 company newsroom post “Introducing Grok 4.7” — Tier A PRIMARY company source, the original record. The live models index on the docs site is the same-day price card for grok-4.7, not a second newsroom. A dedicated grok-4-7 model-card path was not live at filing. The most-capable-for-coding-and-knowledge-work line, the longer-tasks / self-check / best-calibrated-safeguards wording, the new larger base versus Grok 4.6, the longer reinforcement-learning run on multi-hour tasks, the $2 / $6 starting list and the fast variant at twice the output speed and twice the price, availability in Cursor and Grok Build plus the API, harnesses, routers, and cloud platforms, the CursorBench 4.0 company chart (Grok 4.7 xHigh 46.3% versus Grok 4.6 High 40.4%), the new safeguard stack, the LatchBio 62.4% figure, the HackerBench v0.3 3.3% figure, and the invite-only red-team line are company-reported. The docs long-context schedule — $4 input and $12 output once a prompt reaches 200,000 tokens, the same shape as Grok 4.6, with a 500,000-token window — is docs-attributed. NOT claimed: an independent benchmark rerun, that Grok beat Claude or GPT overall, a parameter count, that this desk called the API or opened Cursor, a stock tip, or investment advice. Distinct from the already-filed amazon-blocks-meta-muse, draftkings-ai-target-losers, xai-grok-build-memory, stepfun-step-5-preview, and zai-zcode-open-source.

RELATED

ONLINE

article thread

guidelines

warming…

warming…

On 21 Sep 2026, SpaceXAI published “Introducing Grok 4.7” on its company newsroom. The page is dated Sep 21, 2026. SpaceXAI is the name on that x.ai post — the same newsroom that earlier shipped Memory in Grok Build under the xAI name. That newsroom page is the filing event. These are the company’s words. This desk did not call the API or rerun a benchmark.

Sources