Monday, 31 August 2026 PDT | 03:54 PM
The 1 News Alt Logo Text Smart News for Global Indians

Grok Gets Smarter, Gemini Faster: AI Model War Rages Over Performance, Speed and Price

AI News August 18, 2026 06:00 AM
Grok Gets Smarter, Gemini Faster: AI Model War Rages Over Performance, Speed and Price

Recently, major artificial intelligence (AI) companies have released a succession of new models, further intensifying competition in the AI model market. DeepSeek unveiled V4 Pro 0813, SpaceX (xAI) introduced Grok 4.6, and Google launched Gemini 3.7 Flash.

According to Artificial Analysis (AA), an independent AI model evaluation site, each model has adopted different strategies regarding intelligence, speed, and cost. DeepSeek faced controversy as its performance stagnated while prices rose by up to 12 times, Grok 4.6 boasted significantly enhanced intelligence, and Gemini struck a balance between speed and capability. Additionally, OpenAI previewed an "Ultrafast" mode reaching an extraordinary speed of up to 750 tokens per second.

Grok 4.6 Re-armed with Enhanced Performance in Just Over a Month

SpaceX (xAI) updated its flagship model, Grok, to version 4.6 in just over a month. Despite the rapid release, its performance has been substantially upgraded. The "Intelligence Index" for Grok 4.6 evaluated by AA stands at an impressive 61 points.

Previously, the only models exceeding 60 points on the AA Intelligence Index were Anthropic's Claude Opus 5 and Fable 5, OpenAI's GPT-5.6 Sol, and Moonshot AI's Kimi K3, but Grok 4.6 has now joined their ranks. Tied with GPT-5.6 Sol, it effectively ranks as co-third just beneath the Claude family. Compared to its predecessor Grok 4.5, it rose by 5 points, and compared to Grok 4.3, it jumped by 23 points.

Despite these gains, its pricing remained unchanged at $2 for input and $6 for output per 1 million tokens. By comparison, GPT-5.6 Sol—which shares the same score—costs $5 for input and $30 for output.

According to SpaceX, Grok 4.6 underwent extended pre-training, utilizing AI-generated datasets and high-quality engineering data to enhance reasoning capabilities. During the post-training stage, Supervised Fine-Tuning (SFT) was conducted using the previous version Grok 4.5 in a prompt optimization workflow, followed by Reinforcement Learning (RL).

SpaceX stated that the objective of the SFT stage was to refine responses to user-friendly formats and improve the handling of scientific and programming tasks. Furthermore, behavioral tendencies were improved so that the model automatically checks its work for errors when performing long-horizon projects.

Aligning with SpaceX's claims, Grok 4.6 demonstrated outstanding performance in agentic tasks during AA benchmarks. On the GDPval-AA v2 benchmark, it recorded 62%, placing third behind Claude Opus 5 at 67% (max) and 66% (xhigh), and tying with Claude Fable 5. Grok 4.5 previously scored 51%. GDPval-AA v2 measures practical competence across 44 professions—including finance, law, medicine, software development, and mechanical engineering—evaluating the ability to generate actual workplace deliverables rather than simple answers.

On the τ3\tau^3 τ3-Banking benchmark, which evaluates practical processing capabilities in financial tasks, Grok 4.6 tied for first place with Alibaba's Qwen Max 3.8 at 51%. This test measures whether an AI can accurately navigate vast unstructured knowledge bases and invoke multi-step API tools to solve complex financial tasks.

Additionally, on the AA-Briefcase benchmark, it recorded an Elo of 1577, ranking fourth just below Claude Opus 5 and outperforming Claude Fable 5. The top three spots in this benchmark were all taken by Claude Opus 5 across different reasoning modes. AA-Briefcase tests whether a model can synthesize hundreds to thousands of fragmented information pieces—such as messengers, emails, meeting minutes, and various business documents—like a real corporate environment to complete complex, long-term projects such as reports or financial models.

In Terminal-Bench v2.1, Grok 4.6 scored 88%, tying for third place with three other models, and tied for first place on the GPQA Diamond benchmark (94.9%) alongside Google's newly announced Gemini 3.7 Flash (high).

Terminal-Bench v2.1 measures autonomous problem-solving capability in Linux terminal (CLI) environments for development, DevOps, and data processing. GPQA Diamond evaluates graduate-level difficult scientific questions through multiple-choice tests, where even human experts achieve only around 65% to 70%.

AA particularly praised Grok 4.6 for its low resource consumption per task. According to AA, Grok 4.6 resolved tasks using an average of about 53 turns and approximately 500 million input tokens, whereas Claude Opus 5 (max) required about 103 turns and roughly 2 billion input tokens for the same type of task.

A turn represents a single cycle of "think →\rightarrow → use tools [execute code, read files, etc.] →\rightarrow → result →\rightarrow → determine next action." It can be thought of as the number of times a model checks results after attempting an action. As turns increase, input tokens—the volume of data that must be re-read every turn, including past conversation logs, code, and file contents—grow exponentially.

This means that to reach the same answer, Claude Opus 5 attempted twice as many tries as Grok 4.6 and had to re-read four times as much text in the process.

Because token usage translates directly to API fees, minimizing trial and error can result in lower total costs even if token unit prices are equal or slightly higher. This represents a hidden cost-saving effect in real-world usage. On the AA-Briefcase benchmark, Grok 4.6's cost per task stood at 84 cents, matching Moonshot AI's Kimi K3 while achieving slightly higher intelligence.

AA explained, "Because long-horizon agentic tasks accumulate context rapidly, the fact that Grok 4.6 reaches comparable results in half the turns and with a quarter of the input tokens gives it a fundamental unit-economics advantage that goes beyond nominal token pricing." While its Intelligence Index alone is on par with GPT-5.6 Sol, Grok 4.6 clearly demonstrates distinct strengths in agentic practical work and cost efficiency.

DeepSeek V4 Pro 0813 Lands Just Below Muse Spark 1.2

While Grok 4.6 advanced into the frontier model group, DeepSeek V4 Pro 0813 is being evaluated as having failed to achieve significant performance gains beyond expectations.

The AA Intelligence Index for DeepSeek V4 Pro 0813 is 53 points, marking a noticeable rise from the 45 points of the V4 Pro preview version released in April. However, this is only a 1-point difference from its lower-tier sibling, V4 Flash 0731 (max), which scored 52 points. A score of 53 places it 7th, directly below Meta's Muse Spark 1.2 (xhigh) at 57 points and tied with China's Zhipu AI (Z.ai) GLM-5.2 (max).

Because V4 Flash 0731 showed dramatic performance improvements compared to its preview version, high expectations surrounded the official V4 Pro release as well.

Since DeepSeek previously stated that V4 Flash 0731 shared the exact same architecture and parameter size as its preview version and achieved performance upgrades solely through post-training, the official V4 Pro launch drew identical expectations. The South China Morning Post (SCMP) evaluated it by stating, "Although there were improvements compared to predecessors, the results are somewhat sluggish compared to competing models."

While official analysis posts from AA are not yet published, DeepSeek V4 Pro 0813's benchmark results show that its relatively strong suit is GPQA Diamond, where it scored 93%. However, a score of 93% ranks 8th and is shared by 10 models, including Claude Opus 5 (max) and Kimi K3 (low).

In GDPval-AA v2, it also ranked 8th with a score of 55%, whereas Claude Opus 5 (max) led at 67% and Grok 4.6 (high) took 2nd at 62%.

Its weaker area lies in coding-related metrics; it scored 49% on the SciCode benchmark, placing it near the bottom of comparison graphs just ahead of Nvidia's Nemotron 3 Ultra. SciCode evaluates the ability to generate code pipelines and simulations for actual scientific research, requiring models to implement 338 detailed sub-problems step-by-step into logical code to solve 80 main tasks.

Furthermore, its score on the CritPt (Critical Point) benchmark—which evaluates physics reasoning performance—stood at 18%, presenting a wide gap from the top-ranked GPT-5.6 Sol (max) at 32%. Nevertheless, this evaluation saw Muse Spark 1.2 (xhigh) also scoring 18% and Grok 4.6 (high) scoring 17%, with top-tier models like Claude, GPT, and Kimi K3 dominating the upper tiers. CritPt is an unpublicized benchmark testing advanced physical reasoning, including hypothesis formulation and physical-law-based reasoning at research paper levels across quantum, astrophysics, and high-energy physics.

In Terminal-Bench v2.1, it placed in the mid-tier with 79%, trailing 10 percentage points behind Claude Opus 5's 89%. However, DeepSeek V4 Pro 0813's cost per task was 6 cents, making it the 2nd cheapest among all 104 models based on the Intelligence Index, following GPT-5.6 Luna at 5 cents.

Its output speed reached 83 tokens per second, ranking 18th out of 104 models. It can be characterized as a Pareto Frontier model delivering the highest intelligence relative to its cost tier.

However, DeepSeek announced an increase in its API pricing. The hike is larger than anticipated, introducing a pricing structure that distinguishes between peak and off-peak hours. Peak hours are set from 9:00 AM to noon and from 2:00 PM to 6:00 PM Beijing time. The company explained that this was intended to distribute server loads as current computing capacity cannot satisfy demand.

The price increases vary significantly across four categories: input tokens for cache misses (newly inputted queries), cache hits (discounted re-use of previously processed content), and output generation during peak versus off-peak hours. For example, processing cache-hit inputs with V4 Pro during peak hours increases costs by nearly 12 times.

This has sparked criticism that DeepSeek's identity as a low-cost, high-value AI may be shaken. Even at the hiked peak-hour prices, however, DeepSeek V4 Pro's peak-hour output pricing of $3.96 per 1 million tokens remains significantly cheaper compared to Claude Fable 5's output unit price of $50. Nevertheless, based on peak-hour rates, DeepSeek's price advantage vanishes when compared to GPT-5.6 Luna, which recently slashed prices by 80%.

Google's Gemini 3.7 Flash Released in Three Weeks, Achieving 340 Tokens-Per-Second Output

Google has released its third Flash model in just three weeks. However, delays in launching Google's top-tier model, Gemini 3.5 Pro, overshadowed the announcement of Gemini 3.7 Flash, a point highlighted in foreign media headlines. Axios reported that the rollout came amid responses to new model offensives from competitors such as Grok 4.6 and DeepSeek V4 Pro 0813.

Gemini 3.7 Flash has been strengthened with a focus on coding, debugging, and agentic workflows. Its AA Intelligence Index stands at 56 points, representing a 4-point increase over version 3.6 Flash and falling slightly behind GPT-5.6 Terra and Muse Spark 1.2 at 57 points. AA analyzed that most of this score increase stems from enhanced performance in agentic tasks such as τ3\tau^3 τ3-Banking and Terminal-Bench v2.1.

Evaluations in τ3\tau^3 τ3-Banking and Terminal-Bench v2.1 rose by 3 and over 8 points respectively compared to the predecessor, while GDPval-AA v2 showed the largest margin of increase. Furthermore, DeepSWE v1.1 results (evaluating long-form software engineering task performance on actual open-source repositories) jumped significantly from 49.0% to 65.3%. However, AA pointed out that this version consumes roughly 40% more output tokens than its predecessor, averaging about 37,000 tokens per task.

Gemini 3.7 Flash claimed the top spot in several specific benchmark categories, including AA-AnalystAgent (complex Q&A based on spreadsheets and documents), AutomationBench-AA (simulated SaaS environment agent tasks), WebDev Arena (blind A/B test web development benchmark operated by Arena.ai), and FrontierCode (evaluating code quality for merge viability).

Its pricing is set at $0.75 for input and $3.75 for output per 1 million tokens until the end of 2026—half the price of 3.6 Flash—before increasing to $1.50 and $7.50 starting in 2027.

AA evaluated Gemini 3.7 Flash as a Pareto Frontier model harmonizing speed and intelligence. A Pareto Frontier model denotes an implementation achieving an optimal balance (Pareto optimality) in trade-offs where improving one metric degrades another. In particular, its output speed achieved remarkable growth.

Gemini 3.7 Flash's output speed reaches 340 tokens per second, nearly three times faster than GPT-5.6 Terra and GLM 5.2. According to AA, the average time per task under the "high" reasoning mode is 1.7 minutes, which is 40% faster than GPT-5.6 Terra (max). It earned praise as the smartest model at equivalent speeds and the fastest model at equivalent intelligence.

OpenAI's "Ultrafast" Mode Up to 14x Faster: "Speed Without Intelligence Loss"

While SpaceX's Grok, Meta's Muse Spark, and Google's Gemini engage in fierce contention to enter the frontier model group, OpenAI is also rallying its forces. OpenAI announced via its official blog on Aug. 13 (local time) that it is adding an "Ultrafast" mode to the GPT-5.6 Sol model, initially provided as a preview version to a limited group of customers.

According to OpenAI, Ultrafast mode is an API service tier that operates up to 14 times faster than the standard processing speed of GPT-5.6 Sol, claiming an ability to output up to 750 tokens per second. This was implemented through a partnership with AI chip startup Cerebras.

OpenAI emphasized "speed without intelligence loss," explaining that while achieving near-real-time inference speeds traditionally required reducing capacity (sacrificing model intelligence) or utilizing specialized models, Ultrafast satisfies both speed and model intelligence. While AA previously designated Google's Gemini 3.7 Flash as a Pareto Frontier model, OpenAI's development can be understood as achieving this within top-tier frontier models.

OpenAI introduced that GPT-5.6 Sol with Ultrafast mode can be applied across instant incident response, financial research and security, customer support, and commerce.

In system incident response, it helps analyze logs, recent code changes, and engineer reports in real-time to identify root causes and resolve issues. In finance, it can analyze fluctuating market signals in real-time to reflect them in transaction screening and anomaly detection. In biology and science, real-time experimentation becomes feasible, allowing researchers to run multiple repeated tests during daytime working hours instead of leaving jobs running overnight to check the next morning.

Meanwhile, Anthropic remains quiet. While some claims surfaced on X suggesting that Fable 5.1 has already been developed and is timing its release to match OpenAI's upcoming flagship model announcements, no official announcements or media reports have been found.

However, given that Anthropic's new model release cycle has narrowed significantly to intervals of 4 to 8 weeks this year—following Claude Sonnet 5 in late June, the supply resumption of Fable 5 and Mythos 5 in early July, and Claude Opus 5 in late July—a new model release is likely to occur in late August or early September at the latest.