Best AI Models in September 2026: Frontier Models Compared
TeamdayΒ· 24 min readΒ· 2026-02-20Β· Updated 2026-09-05
AI ModelsGPT-6Claude Fable 5.1GeminiDeepSeek V4Grok 4.6GLM-5.3QwenKimi K3Mistral2026Frontier AI

Best AI Models in September 2026: Frontier Models Compared

The best AI model in September 2026 depends on the work. GPT-6 Astra and Claude Fable 5.1 are the premium choices for difficult, long-running work. Claude Sonnet 5 is the practical daily driver. Grok 4.6 is the price-performance coding pick. Gemini 3.8 Flash is the fast multimodal option. GLM-5.3 Flash and DeepSeek V4 Flash set the floor on price.

Last verified: September 5, 2026. Prices are provider list prices per one million tokens, before volume discounts, long-context surcharges, gateway markups, and taxes. Benchmark results mentioned below are vendor-reported unless stated otherwise.

Best AI Models: Quick Picks

JobBest starting pointCheapest credible alternative
Hardest long-horizon agent workGPT-6 Astra or Claude Fable 5.1 ($10 / $50)Claude Opus 5 ($5 / $25) or GPT-5.6 Sol ($4 / $20 promo)
Daily coding and knowledge workClaude Sonnet 5 ($2 / $10)GPT-5.6 Terra ($2 / $12)
Coding with aggressive price-performanceGrok 4.6 ($2 / $6, 500K context)DeepSeek V4 Pro ($1.32 / $3.96 peak, half off-peak)
Fast multimodal workGemini 3.8 Flash ($0.75 / $3.75 promo)Qwen3.8-Flash ($0.15 / $0.47, image and video input)
High-volume routing, extraction, draftsGPT-5.6 Luna ($0.20 / $1.20)GLM-5.3 Flash ($0.075 / $0.25 promo on OpenRouter)
Cheapest first-party frontier APIDeepSeek V4 Flash ($0.44 / $1.32 peak; $0.22 / $0.66 off-peak)β€”
Cheapest model in this guideGLM-5.3 Flash on OpenRouter ($0.075 / $0.25 promo; $0.15 / $0.50 list)β€”
Long-context reasoning at scaleKimi K3 ($3 / $15, 1M context)GLM-5.3 ($1.40 / $4.40, 1M context)
Open-weight coding specialistKimi K2.7 Code ($0.95 / $4, 262K context)GLM-5.3 Flash (MIT, native vision, 1M context)
European open modelMistral Small 4 (Apache 2.0, multimodal)β€”
Self-hosted ecosystem breadthLlama 4β€”

September 2026 Frontier Model Comparison

ProviderCurrent modelStatusContextDirect API price: input / outputBest fit
OpenAIGPT-6 AstraGA, Sep 31.05M$10 / $50Hardest reasoning, coding, computer use, research
OpenAIGPT-5.6 Sol / Terra / LunaGA, Jul 91.05M$4 / $20 (promo); $2 / $12; $0.20 / $1.20Value tiers below Astra
AnthropicClaude Fable 5.1GA, Sep 11M$10 / $50Difficult autonomous work with a premium ceiling
AnthropicClaude Opus 5GA, Jul 311M$5 / $25Hard coding and long-running agents at half Fable's input price
AnthropicClaude Sonnet 5GA, Jun 301M$2 / $10 (now permanent)Daily coding, agents, computer use
GoogleGemini 3.8 FlashGA, Sep 21M$0.75 / $3.75 (promo through Dec 31)Fast multimodal and long-horizon software engineering
GoogleGemini 3.1 ProPreviewTiered at 200K$2 / $12 up to 200K; $4 / $18 abovePro-tier reasoning while Google's Pro successor is unshipped
xAIGrok 4.6GA, Aug 12500K$2 / $6 (2x above 200K)Coding, engineering, tool use
DeepSeekV4 Pro (0813) / V4 Flash (0731)GA, Aug 13 / Jul 311M$1.32 / $3.96; $0.44 / $1.32 at peak β€” half off-peakLow-cost reasoning and coding
Z.aiGLM-5.3 / GLM-5.3 FlashReleased, Aug 18 / Aug 261M$1.40 / $4.40; $0.15 / $0.50 list ($0.075 / $0.25 promo on OpenRouter)Long-horizon engineering; budget multimodal
AlibabaQwen3.8-Max / Qwen3.8-FlashReleased, Aug 3 / Aug 261M$2 / $6; $0.15 / $0.47 (Token Plan)Multimodal agents and coding
Moonshot AIKimi K3 / Kimi K2.7 CodeReleased, Jul 16 / Jun1M / 262K$3 / $15; $0.95 / $4Frontier reasoning; long coding-agent trajectories
MiniMaxM3Open weight, Jun 1Up to 1M$0.30 / $1.20 up to 512K (July check)Efficient multimodal agents
MistralMedium 3.5 / Small 4Released, Apr / Mar256K$1.50 / $7.50; $0.15 / $0.60 (July check)European deployment and open models
MetaLlama 4 Scout / MaverickOpen weight, Apr 2025Up to 10MHost-dependentSelf-hosting and ecosystem flexibility

Context size is not the same as usable memory. A model may accept one million tokens yet lose accuracy across a noisy history. Long-context price tiers, caching, compaction, and the agent harness often matter more than the headline window.

How We Chose the Models

This is not a single-score leaderboard. We compare:

  • official release status and exact API model IDs;
  • capability for coding, research, tool use, computer use, and knowledge work;
  • input and output limits;
  • direct list price and cache economics;
  • open-weight availability and license;
  • whether the model can be used in a durable agent harness;
  • deprecations and regional access restrictions.

Provider benchmarks are useful evidence, but they are not interchangeable. A score can change with the harness, tool set, reasoning budget, retry policy, test subset, and grader. We use them to understand a model's intended strengths, not to manufacture a false universal ranking.

1. OpenAI GPT-6 Astra, with GPT-5.6 as the Value Tier

OpenAI rolled out GPT-6 Astra (gpt-6-astra) from September 3. It is the first GPT-6 model and OpenAI's most capable model for the hardest reasoning, coding, computer-use, and research work.

Model IDPositionPrice: input / cached / cache write / outputContextMax output
gpt-6-astraFlagship$10 / $1 / $12.50 / $501.05M128K
gpt-5.6-solValue flagship (promo through Nov 21)$4 / $0.40 / $5 / $201.05M128K
gpt-5.6-terraBalanced$2 / $0.20 / $2.50 / $121.05M128K
gpt-5.6-lunaEfficient$0.20 / $0.02 / $0.25 / $1.201.05M128K

Astra takes text and image input, has an April 30, 2026 knowledge cutoff, and supports reasoning effort from low through medium, high, and xhigh to max. The GPT-5.6 "none" effort tier is gone. Batch and Flex processing cost 50% of list; Fast mode costs 2x. Requests above 272K input tokens bill at $20 / $2 / $25 / $75, so dumping a whole workspace into every call is still poor architecture.

Two facts matter for budgeting. First, Astra costs 2.5x GPT-5.6 Sol, so Sol, Terra, and Luna remain the sensible defaults for most work. Second, OpenAI bills prompt-cache writes at 1.25x the input price on every tier. Any estimate that assumes "no cache-write fee" is wrong.

OpenAI retired GPT-5.4 from Codex on August 31, 2026.

Sources: GPT-6 Astra launch, OpenAI API pricing.

2. Anthropic Claude: Fable 5.1, Opus 5, Sonnet 5, and Haiku 4.5

Anthropic now has four relevant choices.

Claude Fable 5.1 (claude-fable-5-1) shipped on September 1 as the successor to Fable 5 in the same tier at the same base price: $10 input and $50 output, with a 1M-token context and up to 128K output tokens. The practical improvement is cache economics β€” cache reads cost $0.25 per million tokens, four times cheaper than Fable 5's $1. Fable 5 stays callable as a legacy model.

Claude Opus 5 is Anthropic's Opus-tier model at $5 input and $25 output. It is built for complex agentic coding and long-running agents β€” the hard jobs where you want Opus-class capability at half Fable 5.1's input price.

Claude Sonnet 5 is the best default for most teams at $2 input and $10 output. Anthropic made the launch price permanent; the planned September 1 increase to $3 / $15 did not happen. Its newer tokenizer produces more tokens than Sonnet 4.6 for the same input, so the lower per-token price does not translate one-to-one into a lower invoice.

Claude Haiku 4.5 remains the fastest and cheapest Claude at $1 input and $5 output with a 200K context. Anthropic has set October 15, 2026 as the earliest retirement date, so plan a migration path for anything pinned to it.

Sources: Claude models overview, Claude Fable 5.1 on Teamday, Claude Opus 5 on Teamday.

3. Google Gemini 3.8 Flash and Gemini 3.1 Pro

Google's Flash line moved fast this summer: Gemini 3.6 Flash on July 21, Gemini 3.7 Flash on August 13, and Gemini 3.8 Flash (gemini-3.8-flash) generally available on September 2.

Gemini 3.8 Flash is Google's most intelligent Flash model, built for long-horizon software engineering and autonomous agents, with a 1M-token context. It costs $0.75 input, $0.075 cached, and $3.75 output β€” the same 50% promotional price as 3.7 Flash. The promotion runs through December 31, 2026; the price doubles on January 1, 2027. Google keeps 3.7 Flash available for compute-efficient workflows. A gated Gemini 3.8 Flash Cyber variant exists for vulnerability work.

Gemini 3.5 Flash still lists at $1.50 input, $0.15 cached, and $9 output β€” double the newer Flash models β€” and has left Google Antigravity's model selector. There is no reason to start new work on it.

Gemini 3.1 Pro remains the Pro-tier option while Google's flagship Pro successor is still unshipped. Its price is tiered: $2 input and $12 output up to 200K tokens, then $4 and $18 above that.

For production, stable model IDs matter. Use Google's deprecation table before pinning any preview alias.

Sources: Gemini API pricing, Gemini model lifecycle, Gemini models.

4. xAI Grok 4.6

xAI released Grok 4.6 on August 12 as its recommended model. It has a 500K-token context window and costs $2 input, $0.50 cached input, and $6 output. Prices double above 200K tokens of context, so keep prompts under that line where you can.

Grok 4.6 replaces Grok 4.5 ($2 / $0.30 cached / $6) as the recommended pick at the same headline price. Grok 4.3 is the budget tier at $1.25 / $0.20 / $2.50 with a 1M context, and Grok Build 0.1 ($1 / $0.20 / $2) is the coding-agent model behind the Grok Build CLI.

Source: xAI models and pricing.

5. DeepSeek V4 Pro (0813) and V4 Flash (0731)

DeepSeek V4 is out of preview. V4 Flash build 0731 shipped on July 31 and V4 Pro build 0813 went GA on August 13. Both have a 1M-token context, and DeepSeek now bills by time of day.

ModelPeak price: input / cache hit / outputOff-peak price
deepseek-v4-pro (0813)$1.32 / $0.044 / $3.96$0.66 / $0.022 / $1.98
deepseek-v4-flash (0731)$0.44 / $0.014 / $1.32$0.22 / $0.007 / $0.66

Peak hours are 01:00–04:00 and 06:00–10:00 UTC on weekdays; everything else is off-peak at half price. OpenRouter lists deepseek/deepseek-v4-pro-0813 at $0.66 / $0.022 / $1.98 around the clock. If your agents run overnight in Europe or during the US working day, the off-peak rate is the one you actually pay.

DeepSeek also released an experimental V4 Flash Vision in late August. Treat it as experimental. DeepSeek V5 has not been announced.

DeepSeek's cache-hit price is exceptionally low, but cached pricing only applies when requests actually reuse the provider's cacheable prefix. The useful metric is cost per completed, accepted task β€” not the cheapest line in a pricing table.

Sources: DeepSeek pricing, DeepSeek changelog, DeepSeek V4 Pro on OpenRouter.

6. Z.ai GLM-5.3 and GLM-5.3 Flash

Z.ai released GLM-5.3 on August 18 as the successor to GLM-5.2, with a 1M-token context. Z.ai lists it at $1.40 input, $0.26 cached, and $4.40 output β€” the same price as GLM-5.2, which is now the previous generation.

GLM-5.3 Flash followed on August 26 and is the more interesting release for cost-sensitive teams. It is a 320B-parameter model with 18B active parameters, MIT-licensed, natively multimodal (text, image, and video input), with a 1M-token context. On OpenRouter (z-ai/glm-5.3-flash) it costs $0.075 input, $0.015 cached, and $0.25 output on a 50% launch promotion; list price is $0.15 / $0.03 / $0.50. Even at list price it is the cheapest model in this guide.

Sources: Z.ai documentation, GLM-5.3 on OpenRouter, GLM-5.3 Flash on OpenRouter.

7. Qwen3.8-Max, Qwen3.8-Flash, and Qwen 3.7 Plus

Alibaba's Qwen line now has three tiers that matter, and the billing plan you hold decides which you can use.

Qwen3.8-Max (August 3) is the premium tier: 1M-token context, $2 input and $6 output. Qwen3.8-Flash (August 26) is the efficient tier: 125B parameters, open weights, text, image, and video input, 1M context, 131K max output, at $0.15 input and $0.47 output. Both are Token Plan (API key) models only β€” neither is included in the $50 Coding Plan.

Qwen 3.7 Plus remains the Coding Plan flagship at $0.40 input and $1.60 output with a 256K context. If you pay for the Coding Plan, this is still your model.

Alibaba also published Qwen3.8-Flash-Next on August 28, an experimental preview of the architecture behind Qwen 4. Qwen 4 itself has no announced date. Use a dated snapshot for production when provider aliases can move.

Sources: Alibaba Model Studio model list, Alibaba model pricing, Qwen3.8-Flash on Teamday.

8. Kimi K3, Kimi K2.7 Code, MiniMax M3, and Mistral

Kimi K3 (July 16) is Moonshot AI's frontier reasoning model: 2.8 trillion total parameters and a 1M-token context. It costs $3 input, $0.30 cached, and $15 output through Vercel AI Gateway or OpenRouter, or draws on a Kimi for Coding plan. It is the pick when you want a non-US-lab frontier model for long-context reasoning.

Kimi K2.7 Code is the coding-focused open-weight mixture-of-experts model below K3: 262K context, multimodal input, $0.95 input and $4 output. It is designed for long agent trajectories rather than short code completion. See the official Kimi K2.7 Code model card.

MiniMax M3 is open weight, natively multimodal, and supports up to one-million-token context, with a guaranteed minimum of 512K depending on the route. Its direct price up to 512K was $0.30 input and $1.20 output when last checked in July. See MiniMax M3.

Mistral Medium 3.5 is the current general frontier tier at $1.50 input and $7.50 output. Mistral Small 4 is Apache 2.0, multimodal, 256K context, and only $0.15 input and $0.60 output. Mistral Large 3 remains a major open-weight model but is no longer the newest Mistral default. Mistral prices were last checked in July. See Mistral Medium 3.5 and Mistral Small 4.

9. Meta Llama 4

Meta has not published a newer general-purpose Llama release than Llama 4 Scout and Maverick, released April 5, 2025.

That makes Llama 4 old by frontier-model standards, but not irrelevant. Scout's 10M headline context and ability to fit on one H100 with Int4 quantization remain useful. Maverick offers a larger 128-expert architecture. More importantly, the Llama ecosystem has broad hosting, fine-tuning, and deployment support.

Do not describe Llama 4 Behemoth as released. Meta previewed it as a teacher model that was still training.

Source: Meta's Llama 4 announcement.

What You Can Use in Teamday Today

Teamday does not claim that every model in this comparison is a one-click option. As of September 5, the first-class routes in the model picker are:

Teamday harnessCurrent choices
Claude CodeClaude Fable 5.1 (recommended), Opus 5, Sonnet 5, Haiku 4.5
CodexGPT-6 Astra (recommended); GPT-5.6 Sol, Terra, Luna
Google AntigravityGemini 3.8 Flash (recommended; Low, Medium, High); Gemini 3.7 Flash; Gemini 3.1 Pro (Low, High)
Qwen CodeQwen 3.7 Plus (Coding Plan); Qwen3.8-Max and Qwen3.8-Flash (Model Studio API key); Qwen 3.6 Plus
Grok BuildGrok 4.6 (recommended), Grok 4.5, Grok Build 0.1
PiKimi K3 (Kimi for Coding, Vercel AI Gateway, or OpenRouter); GLM-5.3 (OpenRouter or Together); GLM-5.3 Flash (OpenRouter); DeepSeek V4 Pro (direct, OpenRouter, or Together); DeepSeek V4 Flash (direct)
OpenCodeKimi K3 (recommended), Kimi K2.7 Code

Saved agents keep their model. When a model retires from the picker, Teamday maps it to the same tier β€” Fable 5 to Fable 5.1, GLM-5.2 to GLM-5.3, Gemini 3.5 Flash to Gemini 3.8 Flash β€” so nothing moves to a pricier tier on its own.

That means you can select a current frontier model for an AI employee, give the underlying agent workspace context and tools, and let it produce durable work rather than a disposable chat answer. Browse AI employees, compare agent harnesses, or see finished work.

The exact model matters, but the harness decides whether the model can inspect files, use tools, recover from failure, run for more than one turn, and leave an auditable result.

How to Choose a Model for Real Work

Use this five-part test:

  1. Define the accepted output. A merged code change, reviewed research memo, updated forecast, or campaign package is testable. β€œBe smart” is not.
  2. Run the same job on two models. Keep tools, context, and acceptance criteria constant.
  3. Measure total task cost. Include output tokens, cache writes, retries, failed tool calls, and human review.
  4. Test long-horizon reliability. Many models are excellent for five minutes and fragile after fifty tool calls.
  5. Pin the model and record the date. Preview aliases and gateway routes change.

For most companies, the winning architecture is tiered:

  • a fast, inexpensive model for routing, extraction, and drafts β€” GPT-5.6 Luna, Gemini 3.8 Flash, or GLM-5.3 Flash;
  • a strong daily model for most agent work β€” Claude Sonnet 5, Grok 4.6, or GPT-5.6 Terra;
  • a premium model for difficult or high-consequence tasks β€” GPT-6 Astra, Claude Fable 5.1, or Claude Opus 5;
  • an open or alternative provider path, such as the OpenRouter free models, for cost control and resilience.

What Changed Since July 2026

The July version of this article was current for about six weeks. The September refresh records the real changes:

  • OpenAI: GPT-6 Astra (September 3) replaced GPT-5.6 Sol as the top model at 2.5x the price. OpenAI cut GPT-5.6 prices over the summer (Sol $4 / $20 through November 21; Terra $2 / $12; Luna $0.20 / $1.20) and retired GPT-5.4 from Codex on August 31. Cache writes cost 1.25x input on every tier β€” the earlier "no cache-write fee" claim was wrong.
  • Anthropic: Claude Fable 5.1 (September 1) succeeded Fable 5 at the same price with cache reads four times cheaper. Sonnet 5's $2 / $10 launch price became permanent. Haiku 4.5 has an October 15 retirement floor.
  • Google: three Flash releases in six weeks β€” 3.6 (July 21), 3.7 (August 13), 3.8 (September 2) β€” all at $0.75 / $3.75. Gemini 3.5 Flash left the Antigravity selector. The Pro successor still has not shipped.
  • xAI: Grok 4.6 (August 12) replaced Grok 4.5 as the recommended model at the same $2 / $6.
  • DeepSeek: V4 went GA as Flash 0731 and Pro 0813, with weekday peak / off-peak billing and off-peak at half price.
  • Z.ai: GLM-5.3 (August 18) replaced GLM-5.2 at the same price; GLM-5.3 Flash (August 26) became the cheapest model in this guide.
  • Alibaba: Qwen3.8-Max (August 3) and Qwen3.8-Flash (August 26) arrived above and below Qwen 3.7 Plus, but only for Token Plan customers.
  • Moonshot: Kimi K3 (July 16) arrived above Kimi K2.7 Code as a 1M-context frontier reasoning model.

The durable lesson is unchanged: never turn a rumor into a row in a comparison table, and never let a price table go two months without a check.

Frequently Asked Questions

What is the best AI model in September 2026?

There is no universal winner. GPT-6 Astra and Claude Fable 5.1 are the top choices for the hardest long-horizon work; Claude Sonnet 5 is the practical daily driver; Grok 4.6 is the price-performance coding pick; Gemini 3.8 Flash is the fast multimodal option; and GLM-5.3 Flash and DeepSeek V4 Flash set the price floor for high-volume work.

What is the newest OpenAI model in September 2026?

GPT-6 Astra, rolled out from September 3, 2026. It has a 1,050,000-token context window, 128,000-token maximum output, text and image input, and costs $10 per million input tokens and $50 per million output tokens. The GPT-5.6 family (Sol, Terra, Luna) stays available as the cheaper tier.

Which AI model is cheapest in September 2026?

GLM-5.3 Flash on OpenRouter is the cheapest model in this comparison at $0.075 per million input tokens and $0.25 per million output tokens on a 50% launch promotion (list price $0.15 / $0.50). The cheapest first-party frontier API is DeepSeek V4 Flash at $0.44 / $1.32 during weekday peak hours and half that off-peak. Among the big three labs, GPT-5.6 Luna is cheapest at $0.20 / $1.20.

Which current AI models are open weight?

Important current open-weight options include DeepSeek V4, GLM-5.3 Flash (MIT), Qwen3.8-Flash, Kimi K2.7 Code, MiniMax M3, Mistral Large 3 and Small 4, and Meta Llama 4. Licenses differ, so open weight does not always mean unrestricted open source.

Can I use these AI models in Teamday?

Teamday's model picker directly exposes GPT-6 Astra and GPT-5.6, Claude Fable 5.1, Opus 5, Sonnet 5 and Haiku 4.5, Gemini 3.8 Flash, 3.7 Flash and 3.1 Pro, Grok 4.6, Qwen 3.8 Max, 3.8 Flash and 3.7 Plus, Kimi K3 and K2.7 Code, GLM-5.3 and GLM-5.3 Flash, and DeepSeek V4 Pro and Flash. MiniMax, Mistral, and Llama are not first-class picks.

Next scheduled verification: October 2026, or sooner after a major provider release or pricing change.