Google has released Gemini 3.7 Flash, improving 4 points over Gemini 3.6 Flash and reaching the Intelligence vs. Time per Task Pareto frontier Read our full analysis in the article below
Artificial Analysis
Technology, Information and Internet
San Francisco, California 33,744 followers
Independent analysis of AI: Understand the AI landscape and analyze AI technologies http://artificialanalysis.com/
About us
Leading independent analysis of AI. Backed by Nat Friedman, Daniel Gross and Andrew Ng.
- Website
-
https://artificialanalysis.ai
External link for Artificial Analysis
- Industry
- Technology, Information and Internet
- Company size
- 11-50 employees
- Headquarters
- San Francisco, California
- Type
- Privately Held
Locations
-
Primary
Get directions
101 Montgomery St
500
San Francisco, California 94104, US
Employees at Artificial Analysis
Updates
-
We’re launching Optima. Now anyone can create a custom benchmark for their use case, leveraging Artificial Analysis’ leading research and platform Building and running benchmarks is difficult. We have distilled Artificial Analysis’ research and experience developing benchmarks into Optima, a new platform for benchmarking models on your own workloads and comparing performance, speed and cost efficiency. Optima allows you to find the best model for your task, or an equally performant alternative to your current setup at 10x lower cost or time per task. We’ve integrated Artificial Analysis' research and experience in benchmarks across the Optima workflow: ➤ Build benchmarks based on your own data and use cases: There are three ways to build a benchmark with Optima. Upload an existing evaluation dataset from your own files or Hugging Face, or import agent traces from platforms including Arize AI, Braintrust and Langfuse. Install the Optima skill to build a benchmark using context from your coding environment and previous sessions. Or simply describe your use case and provide example inputs and outputs, and Optima will build the benchmark for you ➤ Run across the latest models: Run the same benchmark across leading models in a single click, and keep your leaderboard up to date as soon as new models are released ➤ Bring Artificial Analysis grading to your own benchmark: Evaluate responses against objective rubric criteria or using the same pairwise judging approach used for Artificial Analysis benchmarks including GDPval-AA and AA-Briefcase. For pairwise judging, select your preferred responses from a sample and Optima uses those preferences to rank models across your test set ➤ Compare performance, cost and time efficiency: Optima measures more than model performance. Cost per Task and Time per Task are tracked alongside benchmark scores, with category-level results and support for custom metrics, allowing you to compare the tradeoffs between models for your specific use case Ahead of launch, here are examples questions our beta testers answered with Optima: ➤ Which model can save me 10x the cost without a meaningful decrease in quality for my finance & accounting agent? ➤ Which model best matches the writing style of lawyers for my legal agent? ➤ Which model can best identify different elements in my custom image dataset? Optima is available today. Build your own benchmark at https://lnkd.in/e_5UQhCb
-
Korean AI lab Upstage has released Solar Pro 4, scoring 42 on the Artificial Analysis Intelligence Index, a significant increase from Solar Pro 3’s 14 Solar Pro 4 is Upstage's new proprietary flagship reasoning model, replacing Solar Pro 3 from April 2026. At 42 on the Intelligence Index it sits alongside Inkling (xhigh, 42) and just behind MiMo-V2.5-Pro (43), and shows a 27-point increase over Solar Pro 3. Pricing increases to $0.30/$1.20/$0.06 per 1M input/output/cache hit tokens from Solar Pro 3's $0.15/$0.60/$0.02 via Upstage’s first-party API. Key results: ➤ Solar Pro 4’s largest improvements on Solar Pro 3 are on agentic and long context work. Terminal-Bench v2.1 improves from 12% to 57%, AA-LCR from 31% to 71%, and τ³-Banking from 9% to 23%. GDPval-AA v2 shows strong progress on real-world agentic tasks, where Solar Pro 3 scored an Elo of 498, well below the human baseline of 1000, Solar Pro 4 scores 1276. ➤ AA-Omniscience improvement from -53 to -1 was driven by abstaining on more questions. Solar Pro 4 attempts only 41% of questions against 92% for Solar Pro 3, and its hallucination rate is 24%, higher than Command A+ (14%) and MiniMax-M3 (18%), and a vast improvement from Solar Pro 3’s 88%. AA-Omniscience Accuracy remains unchanged at 19%. ➤ Solar Pro 4 is more token efficient than Solar Pro 3, though still verbose for its intelligence level. It uses 43k output tokens per Intelligence Index task, around 17% fewer than Solar Pro 3's 52k. ➤ The intelligence gain comes with a hit to latency. Solar Pro 4 takes 9.5 minutes to complete an average Intelligence Index task, against 6.9 minutes for Solar Pro 3, despite using fewer output tokens per task. ➤ Pricing is $0.30/$1.20 per 1M input/output tokens. This is in line with MiniMax's first-party pricing for MiniMax-M3, which scores 3 points higher at 45, and is more expensive than DeepSeek V4 Flash 0731 (Reasoning, Max Effort) at $0.14/$0.28 and 52 on the Intelligence Index. Cache hits are priced at $0.06 per 1M, an 80% discount on input token price. Additional model details: ➤ Context window: 512K tokens ➤ Max output tokens: 128K ➤ Modalities: Text input and output only ➤ Pricing: $0.30 / $1.20 / $0.06 per 1M input/output/cache hit tokens ➤ Inference providers at time of launch: Upstage first-party API, OpenRouter, TimelyRouter See the full results for Solar Pro 4 at https://lnkd.in/ebgsKHrf
-
-
SpaceXAI's Grok 4.6 scores 61 on the Artificial Analysis Intelligence Index, joining the frontier in line with GPT-5.6 Sol, with standout agentic performance at lower cost Grok 4.6 gains 5 points over Grok 4.5 on the Intelligence Index just over one month after its release, or +23 points compared to Grok 4.3. This brings SpaceXAI back to the intelligence frontier alongside OpenAI, behind only Anthropic. Key takeaways: ➤ Grok 4.6 joins the frontier of the Artificial Analysis Intelligence Index: It scores 61, in line with GPT-5.6 Sol (max), behind Claude Opus 5 (max, 63) and Claude Fable 5 (max with fallback, 62), and just ahead of Kimi K3 ➤ Strong agentic performance: Grok 4.6 achieves a GDPval-AA v2 Elo of 1753, behind only Claude Opus 5 and with overlapping confidence intervals with Claude Fable 5 and Qwen3.8 Max. It scores 50.7% on 𝜏³-Banking, among the top two scores alongside Qwen3.8 Max (51.3%), and 88.4% on Terminal-Bench v2.1, in line with the leading models ➤ Frontier-level intelligence at lower cost: Headline pricing is unchanged from Grok 4.5 at $2/$6 per 1M input/output tokens, 60%+ below Claude Opus 5 ($5/$25) and GPT-5.6 Sol ($5/$30). It cost $0.84 per task, the same as Kimi K3 with slightly higher intelligence, placing it on the Intelligence vs. Cost per Task Pareto frontier ➤ Grok 4.6 sits at Fable 5-tier on AA-Briefcase, our private benchmark of long-horizon agentic knowledge work tasks, with an Elo of 1577 - behind the Claude Opus 5 family. It is notably turn-efficient, completing tasks in ~53 turns and ~0.5B input tokens on average vs. ~103 turns and ~2.0B input tokens for Claude Opus 5 (max) Other model details: ➤ Context window of 500k tokens (unchanged from Grok 4.5) ➤ Pricing of $2/$6 per 1M tokens of input/output; cache hits discounted to $0.5 per 1M tokens, an increase over Grok 4.5’s $0.3 per 1M tokens for cache hits Further analysis on Artificial Analysis: https://lnkd.in/gKPHU5Vv
-
-
Artificial Analysis reposted this
Announcing AA-AnalystAgent, our new agentic benchmark for quantitative analysis on real-world spreadsheets & documents. Claude Opus 5 leads at 54%, followed by GPT-5.5 at 50% and Claude Fable 5 at 49% In real analyst roles, professional judgment and expertise are as important raw quantitative capabilities. AA-AnalystAgent tests this and requires models to interpret sources, decide which exceptions and caveats apply, and settle on a methodology to successfully complete a task. Because handing analyst tasks to agents requires not just correct answers but consistent ones, AA-AnalystAgent runs each task five times and reports pass^5 as its headline metric (we call this ‘pass-all-5’). pass^5 means models must get a task correct every time it tries across five independent attempts to pass. AA-AnalystAgent overview: 🔢 80 questions across 14 business and scientific domains, including healthcare expenditure reports, trade and commodity statistics, hydrology and weather data, government appropriations, energy cost models, financial models, environmental reporting, and project schedules 🧠 Five workflow buckets from across real analyst work: source lookup and diagnosis, filter and total, ratios/trends/sensitivities, P&L modeling, and cash/balance sheet/valuation modeling Methodology details: 🤖 Agentic harness: each task is solved by an agent running in our open-source Stirrup reference harness, with tools for code execution, web fetch, image viewing, and answer submission 📊 Scoring: each task is run five times per model, with final answers compared to reference solutions by an equality checker. Three metrics: pass^5, pass@1 (average pass rate), pass@5 🔒 Privately-held question set to limit contamination risk. Two example questions from California Medicaid expenditure reports are publicly shown on the methodology page with full prompt and source material Key findings: 🥇 Anthropic's Claude Opus 5 (max) leads at 54%, followed by OpenAI's GPT-5.5 (xhigh) at 50% and Claude Fable 5 (max, Opus 4.8 Fallback) at 49%. Anthropic holds three of the top five places, and the top three are separated by three net tasks out of 80. 🎯 Reliability separates the top of the leaderboard more than raw capability: GPT-5.5 (xhigh) has the highest pass@1, but Opus 5 leads on pass^5 because it repeats what it gets right. 🔎 Committing early to a wrong interpretation is the most widespread way models fail, appearing in 57% of the failures we classified. 💰 The price of a given score varies enormously: Claude Sonnet 4.6 and Xiaomi Technology's MiMo-V2.5-Pro both score 20%, at $1.34 and $0.05 per task. 🥈 Kimi (Moonshot AI)'s Kimi K3 (max) is the top open weights model at 39%, 15 points behind the closed frontier. AA-AnalystAgent launches as a standalone leaderboard and is not part of the Artificial Analysis Intelligence Index. See below for further detail ⬇️
-
-
-
-
-
+1
-
-
Ant Group has released Ling 3.0 Tiny, which scores 25 on the Artificial Analysis Intelligence Index. With just 7.9B total and 1.3B active parameters, it sits on the Pareto frontier for Intelligence vs. Active Parameters Ant Group has released Ling 3.0 Tiny, an open weights reasoning model with a 262K-token context window. It scores 25 on the Artificial Analysis Intelligence Index, comparable to gpt-oss-120b (high, 24) with 15x fewer total parameters and 4x fewer active parameters. However, this parameter efficiency comes with high token usage: Ling 3.0 Tiny used 213M output tokens to run the Intelligence Index, nearly as many as the larger Ling 3.0 Flash (240M). It follows the recent release of Ling 3.0 Flash, which scores 38 with 124B total and 5.1B active parameters. The weights are available on Hugging Face under an MIT license. Key results: ➤ Ling 3.0 Tiny extends the open weights Pareto frontier for Intelligence vs. Active Parameters. It scores 25 on the Artificial Analysis Intelligence Index, and at 7.9B total and 1.3B active parameters, the model is small enough to run locally in many settings ➤ Ling 3.0 Tiny makes a 59-point improvement in AA-Omniscience over Ling-mini-2.0, primarily driven by improvements in hallucination rate. Ling 3.0 Tiny scores -19 with 9% accuracy and a 30% hallucination rate, similar to Qwen3.6 27B (-20), but its accuracy is relatively low among similarly sized open weights models. Compared to the previous generation, where Ling-mini-2.0 had a 96% hallucination rate, Ling 3.0 Tiny improves on hallucination while maintaining the same accuracy. Instead of guessing when it doesn’t know, it attempted just 37% of questions in the AA-Omniscience evaluation. ➤ Ling 3.0 Tiny is relatively capable in tool use but performs modestly in agentic work tasks. Ling 3.0 Tiny scores 21% on 𝜏³-Banking, ahead of Qwen3.6 27B (17%). On GDPval-AA v2, however, it reaches an Elo rating of 772, behind Nemotron 3.5 Lightning at 824 despite sitting slightly ahead on the overall Intelligence Index. Additional model details: ➤ Size: 7.9B total parameters, 1.3B active (mixture of experts) ➤ Context window: 262K tokens ➤ License: MIT ➤ Providers: InclusionAI first-party API and Novita
-
-
Announcing AA-AnalystAgent, our new agentic benchmark for quantitative analysis on real-world spreadsheets & documents. Claude Opus 5 leads at 54%, followed by GPT-5.5 at 50% and Claude Fable 5 at 49% In real analyst roles, professional judgment and expertise are as important raw quantitative capabilities. AA-AnalystAgent tests this and requires models to interpret sources, decide which exceptions and caveats apply, and settle on a methodology to successfully complete a task. Because handing analyst tasks to agents requires not just correct answers but consistent ones, AA-AnalystAgent runs each task five times and reports pass^5 as its headline metric (we call this ‘pass-all-5’). pass^5 means models must get a task correct every time it tries across five independent attempts to pass. AA-AnalystAgent overview: 🔢 80 questions across 14 business and scientific domains, including healthcare expenditure reports, trade and commodity statistics, hydrology and weather data, government appropriations, energy cost models, financial models, environmental reporting, and project schedules 🧠 Five workflow buckets from across real analyst work: source lookup and diagnosis, filter and total, ratios/trends/sensitivities, P&L modeling, and cash/balance sheet/valuation modeling Methodology details: 🤖 Agentic harness: each task is solved by an agent running in our open-source Stirrup reference harness, with tools for code execution, web fetch, image viewing, and answer submission 📊 Scoring: each task is run five times per model, with final answers compared to reference solutions by an equality checker. Three metrics: pass^5, pass@1 (average pass rate), pass@5 🔒 Privately-held question set to limit contamination risk. Two example questions from California Medicaid expenditure reports are publicly shown on the methodology page with full prompt and source material Key findings: 🥇 Anthropic's Claude Opus 5 (max) leads at 54%, followed by OpenAI's GPT-5.5 (xhigh) at 50% and Claude Fable 5 (max, Opus 4.8 Fallback) at 49%. Anthropic holds three of the top five places, and the top three are separated by three net tasks out of 80. 🎯 Reliability separates the top of the leaderboard more than raw capability: GPT-5.5 (xhigh) has the highest pass@1, but Opus 5 leads on pass^5 because it repeats what it gets right. 🔎 Committing early to a wrong interpretation is the most widespread way models fail, appearing in 57% of the failures we classified. 💰 The price of a given score varies enormously: Claude Sonnet 4.6 and Xiaomi Technology's MiMo-V2.5-Pro both score 20%, at $1.34 and $0.05 per task. 🥈 Kimi (Moonshot AI)'s Kimi K3 (max) is the top open weights model at 39%, 15 points behind the closed frontier. AA-AnalystAgent launches as a standalone leaderboard and is not part of the Artificial Analysis Intelligence Index. See below for further detail ⬇️
-
-
-
-
-
+1
-
-
For the same model, output speed can vary by over 15x depending on the inference provider. We're unpacking why those differences exist and what they mean in practice on Wednesday, August 12 in San Francisco Hear directly from Artificial Analysis and leading inference providers on how they think about speed, cost, and model support. Register to attend 👉 https://lnkd.in/gACmB5Wa
-
We're building a range of new AI benchmarking products at Artificial Analysis to help developers evaluate AI models. We’re looking for a small group of users to test them and share feedback before launch. Let us know if you would like early access! Apply for early access here: https://lnkd.in/g6cXK_VU