Large Language ModelsGenerate imagesGenerate videos

Grok 4.7: Release Date, Price, Context Window and Benchmarks

Grok 4.7 arrived on September 21, 2026 with the same $2 and $6 token prices as Grok 4.6, a 500,000 token context window and higher coding scores. See the release facts, the pricing tiers, xAI's reported benchmarks and what independent tests found about cost per task.

Grok 4.7: Release Date, Price, Context Window and Benchmarks
Cristian Da Conceicao
Founder of Picasso IA

Grok 4.7 landed on September 21, 2026, and the first thing people noticed was what did not change: the price. After a launch that slipped past earlier target dates, xAI (now branded SpaceXAI) kept Grok 4.6's rates of $2 per million input tokens and $6 per million output tokens, and it kept the 500,000 token context window. What changed is how hard the model works on each task, and that is where the real cost story sits.

This page puts the release date, every price tier, the context window limits, the benchmark scores xAI reported and the numbers independent testers measured in one place. Where sources disagree, you will see both figures. A vendor score is a claim until somebody else repeats it, so each number below is labeled with who produced it.

💡 Quick read: Same token prices as Grok 4.6, same 500K window, higher scores on all seven of xAI's benchmarks, and about 2.25 times more output tokens per task. The per-task bill is the number to watch, not the per-token rate.

Grok 4.7 Release Date

What Launched on September 21

xAI released Grok 4.7 on September 21, 2026, announced from the SpaceXAI account at about 16:17 UTC. Its predecessor, Grok 4.6, arrived on August 12, 2026, so only 40 days separate the two models, even though reports describe 4.7 as later than earlier target dates. The basic facts fit in one table.

DetailGrok 4.7
Release dateSeptember 21, 2026
Previous modelGrok 4.6, August 12, 2026
API model IDgrok-4.7
InputText and images
OutputText
Reasoning effortlow, medium, high (default), xhigh
Knowledge cutoffMay or June 2026, sources differ

One integrator notes that the dashed spelling grok-4-7 is rejected, so use the dotted ID in API calls. On the technical side, xAI describes a larger base model with extended reinforcement training on tasks that run for hours, plus better habits for checking its own work and managing long context.

Wall calendar with a late September date circled in red marker on a desk

Where You Can Use It

Grok 4.7 reached developers through several doors on day one:

  • xAI API: live at launch under the grok-4.7 ID.
  • Cursor: added it right away, including a faster paid variant.
  • Grok Build: xAI's own coding harness, where several of the reported benchmark runs were made.
  • GitHub Copilot: rolled it out gradually after launch.
  • OpenRouter and Vercel AI Gateway: both listed it for routing.
  • Cloud platforms: Bedrock, Model Garden and Foundry lagged behind the xAI API at launch.

Listed rate limits are 150 requests per second and 50 million tokens per minute, which is generous enough that most teams will hit their budget long before the cap.

Grok 4.7 Price Per Million Tokens

Standard and Long Prompt Rates

Pricing is identical to Grok 4.6. The rate depends on how long your prompt is, and the whole request bills at the higher tier once you cross the line, not just the extra tokens.

Prompt sizeInputCached inputOutput
Under 200K tokens$2.00$0.50$6.00
200K tokens and above$4.00$1.00$12.00

All prices are per million tokens. Cached input costs a quarter of the standard input rate, which makes repeated system prompts and shared project context much cheaper.

There is also a fast variant available only inside Cursor and Grok Build, not through the public API. It is described as twice the output speed at twice the price, so plan on double the standard rates if you choose it.

Overhead view of printed invoices, a calculator and a notebook of numbers on an oak table

Real Bills for Real Prompts

Rates mean little until you apply them. The table below assumes a reply of 5,000 output tokens and no caching.

Prompt sizeInput costOutput costTotal
50K tokens$0.10$0.03$0.13
150K tokens$0.30$0.03$0.33
199K tokens$0.40$0.03$0.43
201K tokens$0.80$0.06$0.86
400K tokens$1.60$0.06$1.66

Look at the middle rows. Adding just 2,000 tokens to a prompt takes the bill from $0.43 to $0.86, a doubling caused by nothing more than crossing the threshold.

💡 Tip: If your prompt sits near 200K tokens, trim it or split the job into two calls. Cached input at the lower tier is usually cheaper than one oversized request.

The 500K Context Window

What 500,000 Tokens Holds

Grok 4.7 accepts 500,000 tokens in a single request, exactly like Grok 4.6, so this is not an upgrade. A common rule of thumb is that one token equals about three quarters of an English word. That puts the window near 375,000 words, or roughly 1,500 pages of plain text. In practice it can hold:

  • A mid-sized code repository with its tests and documentation
  • Several long contracts or research papers at the same time
  • A full week of support tickets for a pattern search
  • A very long agent session with tool output still in memory

Treat these as estimates. Code, tables and non-English text use more tokens per word than ordinary prose, so the real capacity shifts with your content.

When Long Prompts Cost Double

The window is large, but the price tier is what limits you. Anything at or above 200K tokens moves to the $4 input and $12 output rates for the entire request. A full 500K prompt with a 5,000 token reply costs about $2.06 at those rates, against $0.33 for a 150K prompt.

Aerial view of a very long paper scroll unrolled across a library reading table

A bigger window also does not guarantee better recall. xAI says 4.7 is better at managing longer context and checking its own work across multi-hour tasks, but those are launch claims, and you should test retrieval on your own documents before sending half a million tokens in production.

Benchmarks xAI Reported

Coding and Terminal Scores

xAI published a table of seven benchmarks, and Grok 4.7 beat Grok 4.6 on every one. The coding rows show the largest movement.

BenchmarkGrok 4.6Grok 4.7Grok 4.7 setting
CursorBench 4.040.4%46.3%xhigh
DeepSWE v1.165.2%71.0%high
Terminal-Bench 4.020.3%38.0%xhigh, Grok Build harness
SWE-Marathon v1.131.9%46.0%high

Terminal-Bench 4.0 is the headline, a gain of roughly 17 points, and one outlet lists the 4.7 score as 37.6% instead of 38.0%. SWE-Marathon also jumped by more than 14 points, which fits the focus on long, multi-hour work.

One detail deserves attention. The launch table pairs Grok 4.6 at high effort with Grok 4.7 at xhigh on CursorBench, so that comparison is not like for like. Higher effort means more reasoning tokens, a point that comes back in the cost section below.

Knowledge Work Scores

Outside coding, the picture is smaller but still positive.

BenchmarkGrok 4.6Grok 4.7
EEBench53.0%64.0%
HealthBench Professional48.5%56.7%
Harvey legal benchmark15.8%19.6%

The legal score is a reminder of how hard these tests are. Even after a gain of nearly four points, Grok 4.7 lands under 20% on the Harvey benchmark, so legal drafting still needs human review.

Hand pointing a pen at printed bar charts pinned to a cork board

💡 Caveat: Every score in this section comes from xAI's own materials. Third-party sites that republished them state plainly that they did not verify the numbers.

What Independent Tests Found

Index Scores From Third Parties

An independent leaderboard tracker measured Grok 4.7 on launch day and found smaller gains than the vendor table suggests:

  • Intelligence Index: a score of 46, ranked 16th out of 655 tracked models, up 2 points on Grok 4.6.
  • Coding Agent Index: up 9 points on Grok 4.6.
  • Cyber Index (September 28): 56%, tied for first place among 17 models.
  • VulcanBench (October 4): Grok 4.7 ranked in the top three, ahead of the 5.1 Max tier of the Claude Fable 5 family.

The pattern is consistent. The model is clearly better at agentic coding and security work, while general reasoning moved only a little. That fits a model tuned by training on long tasks rather than one built for chat.

The Token Appetite Problem

Here is the catch. The same tracker recorded about 81,000 output tokens per task for Grok 4.7, against 36,000 for Grok 4.6, a ratio of roughly 2.25. Other testers saw the same thing: at xhigh, the new model writes more than twice the output of 4.6 at high.

MetricGrok 4.6Grok 4.7
Output tokens per task36,00081,000
Output cost per task at $6 per million$0.22$0.49
Per-token price$6 per million$6 per million

Output cost only, before input and caching. The rate card did not move, but cost per finished task roughly doubled. A cheap model that talks twice as much is not a cheap model.

Close-up of a yellow taxi fare meter with the numbers rolling during a rainy ride

You can keep that appetite under control with a few habits:

  • Cap output length with a sensible max token limit on every call.
  • Use low or medium effort for the easy steps in an agent loop, and save xhigh for the step that needs it.
  • Cache shared context so the $0.50 input rate applies to your system prompt and project files.
  • Route short tasks to Grok 4.6 or a smaller model, and send only long, tool-heavy jobs to 4.7.
  • Log tokens per finished task, not tokens per request, so retries and verbose replies show up in your numbers.

Two claims remain unconfirmed. xAI gave no parameter count, and a figure of 2.1 trillion floated by Elon Musk has not been backed by launch materials. Musk also mentioned supplemental SpaceX data, which does not appear in the official documents. Until xAI publishes details, treat both as rumor.

Grok 4.7 Against Its Rivals

Grok 4.7 vs Grok 4.6

FeatureGrok 4.6Grok 4.7
ReleaseAugust 12, 2026September 21, 2026
Input price$2 per million$2 per million
Output price$6 per million$6 per million
Context window500K500K
Effort levelslow to xhighlow to xhigh
Output tokens per task36,00081,000
Terminal-Bench 4.020.3%about 38%

The recommendation from the comparisons I reviewed is simple. Upgrade for agentic coding, terminal automation and long repository work. Stay on 4.6 for short answers, classification and chat, where the extra verbosity only adds cost. A rough split looks like this:

  • Switch now: coding agents, multi-file refactors, shell automation, security triage
  • Test first: legal, health and finance tasks, where low absolute scores call for review
  • Stay put: short Q&A, summaries, support replies, anything priced per message

Developer typing at night with two monitors showing code and a warm desk lamp

Fable 5.1 Max and Other Rivals

Against the Claude Fable 5 family, specifically the 5.1 Max tier, xAI's own table shows Grok 4.7 ahead on two of seven benchmarks and behind on the coding rows. The VulcanBench result points the other way, so the honest summary is that the two trade wins depending on the test.

Clean head-to-head numbers against other labs are still thin. If you are choosing between Grok 4.7, GPT 5.6 Sol and Gemini 3.5 Flash, run the same ten real prompts through each and compare both quality and total tokens used.

Four sprinters crouched at the starting blocks of a running track at sunrise

Picking a Reasoning Effort Level

Grok 4.7 offers four settings: low, medium, high and xhigh, with high as the default. Because output tokens drive your bill, effort is a cost dial as much as a quality dial.

  1. Low: quick lookups and formatting jobs
  2. Medium: everyday writing and light code edits
  3. High: the default, a sensible starting point for most agent work
  4. Xhigh: hard multi-step tasks where a wrong answer costs more than extra tokens

Start at high, measure tokens per finished task, and move to xhigh only for the jobs that fail without it.

Hand sliding one of four faders on an analog mixing console

Try the Grok Family on PicassoIA

Grok 4.7 is not listed on PicassoIA at the time of writing, but its predecessor Grok 4 is, and it runs directly from a text field with no API access or setup. It is a quick way to get a feel for xAI's reasoning style before you commit to API spend.

How to Use Grok 4 on PicassoIA

  1. Open the Grok 4 page in the large language models collection.
  2. Type your question into the Prompt field. Multi-step problems and comparisons suit it best.
  3. Keep Temperature at the default of 0.1 for factual work, and raise it a little for brainstorming.
  4. Check Max Tokens. The default is 2,048, so raise it if long answers get cut off.
  5. Leave Top P at 1 and both penalties at 0 unless the answer repeats itself.
  6. Run the prompt, then paste the same text into another model for comparison.

For that last step, Claude Sonnet 5, GPT 5.6 Sol and Gemini 3.5 Flash make good rivals, and the Kimi K2.6 page is worth a look for agent-style coding prompts.

Make Your Own Images Next

Reading about models is half the job. The other half is making something with them, and that is where PicassoIA earns its place. Draft a script, a product description or a blog outline with a language model, then turn the same idea into visuals in a few clicks.

Designer arranging photo prints into a mood board at a long studio table

For photorealistic results, start with Seedream 4.5, GPT Image 2, Flux 2 Pro or Nano Banana Pro. Describe the subject, the light and the lens, such as "85mm, soft window light, film grain", and compare two models on the same prompt. When a still looks right, the text-to-video collection can turn it into a short clip.

Browse every model at picassoia.com/en/all-models, pick one, write your first prompt and see what you can create today.

Share this article