Artificial Analysis 评测 Gemini 4 Argon,智能指数追平 GPT-6 Astra 且成本仅 60%

📌 One-Sentence Summary
Artificial Analysis benchmarks Google DeepMind's Gemini 4 Argon at 53 on its Intelligence Index, matching GPT-6 Astra (max.) while costing 60 percent less per task under launch discounts.
📝 Summary
Artificial Analysis evaluated Google DeepMind's new Gemini 4 Argon, the first non-Flash proprietary Gemini model in more than seven months. At high reasoning it scores 53 on the Artificial Analysis Intelligence Index, tying GPT-6 Astra (max, 53) and edging GPT-6.1 Sol (max, 52) by one point, a 23-point jump over Gemini 3.1 Pro Preview (30). With a 50 percent launch discount and a 95 percent cached input discount, each Intelligence Index task costs $1.99, about 60 percent of GPT-6 Astra (max, $3.26) but 2.7x GPT-6.1 Sol (max); standard pricing would raise this to $3.98. Gains come from lower hallucination (15 percent on AA-Omniscience, the lowest among the over 45 models) and stronger agentic performance (77.5 percent on AutomationBench-AA, 57 percent on Term Bench 4). The model offers a 1 million-token context window, text, image, video, and audio input, and a new Long Decode Continuation feature. The model is rolling out to select users and is not publicly available yet.
💡 Main Points
Gemini 4 Argon matches the top model on the Intelligence Index, restoring Google to the top tier of labs.
It scores 53 at high reasoning, tying GPT-6 Astra (max, 53) and beating GPT-6.1 Sol (max, 52) by one point, a 23-point jump over Gemini 3.1 Pro Preview (30) and 12 points over Gemini 3.8 Flash at high reason.
Launch discounts make it cost-competitive, but the advantage is temporary.
At the current 50 percent launch discount and a 95 percent cached input discount, each Intelligence Index task costs $1.99, 60 percent of GPT-6 Astra (max, $3.26) but 2.7x GPT-6.1 Sol (max). Standard pricing would raise this to $3.98, about 1.2x GPT-6 Astra. The efficiency comes from lower token prices, not fewer tokens: it outputs 62k tokens per task versus 27k for GPT-6 Astra.
Agentic performance, historically a Gemini weakness, shows clear improvement.
It ranks first on AutomationBench-AA at 77.5 percent, ahead of Claude Sonnet 5.5 (max, 71.3 percent), and reaches 57 percent on Term Bench 4, a 53-point gain over Gemini 3.1 Pro Preview, trailing only Claude Sonnet 5.5 (64 percent), Claude Opus 5.5 (60 percent), and GPT-6 Astra (59 percent).
It has the lowest hallucination rate among leading models, though raw accuracy is slightly lower.
On AA-Omniscience its hallucination rate is 15 percent, versus 51 percent for GPT-6 Astra (max.) and 54 percent for GPT-6.1 Sol (max.), meaning it more often admits uncertainty. Accuracy is 50 percent, down 5 points from Gemini 3.1 Pro Preview and 13 points below GPT-6 Astra (63 percent), leaving its overall AA-Omniscience score of 42 in line with GPT-6 Astra (43) and GPT-6.1 Sol (42).
The model ships with a large context window and a new long-output feature.
It supports a one-million-token context window and text, image, video, and audio input with a text output. The Long Decode Continuation, a new Gemini API feature, allows long-sentence responses to be paused and resumed across calls, enabling up to 1 million output tokens without request timeouts.
💬 Key Quotes
Google's new Gemini 4 Argon scores a 53 on Artificial Analysis's Intelligence Index, matching GPT-6 Astra (max.), while costing 60 percent less per task at launch.
With a 50 percent launch discount and a 95 percent cached input discount, each Intelligence Index task costs $1.99, about 60 percent of GPT-6 Astra (max., $3.26), but 2.7x GPT-6.1 Sol (max.) Standard pricing would raise this to $3.98. Gains come from lower hallucination (15 percent on AA-Omniscience, the lowest of over 45 models) and stronger agentic performance (77.5 percent on AutomationBench-AA, 57 percent on Term Bench 4).
At the current launch discounts, each skill and domain task costs $3.98, about 40 percent of GPT-6 Astra (max., $9.79) but 5x GPT-6.1 Sol (max.). If those discounts were never applied, each task would cost $4.98, about double the cost of the top model in the Smashing Business Edition (about $2.50).
In the $0.20 Plan, service costs are flat whether you use just one skill or the full suite in a single request. For requests outside the $0.20 Plan, which covers 25 percent of queries, and in the Free Plan, costs are higher for each additional skill or domain.
📊 Article Meta
AI Screening: 85
Source: AIHOT — 精选
Author: noreply@aihot.news (X:Artificial Analysis (@ArtificialAnlys))
Category: 人工智能
Language: 英文
Read Time: 6 min
Word Count: 1362
Tags:
AI 与智能应用 , 模型评测与基准 , 模型发布 , 大语言模型 (LLM) , AI Agent
这个比较实用,已转发给同事。