Evergreen LLM guide
The leading large language models, compared
A practical overview of ChatGPT, Claude, Gemini, DeepSeek, Kimi and Z.ai's GLM family—what they are, what their APIs cost and why parameter counts rarely tell the whole story.
The short version
There is no single “best LLM.” The right choice depends on the job, the quality and latency you need, the tools and context the model must use, and the cost of its complete workflow. The table below uses each provider's representative flagship API model so that prices are easier to compare.
Flagship LLM price comparison
USD per one million tokens, standard API processing. Consumer chat subscriptions are not included.
LAST CHECKED 17 AUG 2026
| LLM | API model | Input / 1M | Output / 1M | Parameters | Pricing note |
|---|---|---|---|---|---|
| ChatGPTOpenAI | GPT-5.6 Sol | $5.00 | $30.00 | Not disclosed | Cached input: $0.50 |
| ClaudeAnthropic | Claude Opus 5 | $5.00 | $25.00 | Not disclosed | Cache hits: $0.50 |
| GeminiGoogle | Gemini 3.1 Pro Preview | $2.00 | $12.00 | Not disclosed | Prompts up to 200K tokens |
| DeepSeekDeepSeek | DeepSeek V4 Pro | $1.32 | $3.96 | Not disclosed | Peak rate; off-peak is 50% less |
| KimiMoonshot AI | Kimi K3 | $3.00 | $15.00 | 2.8T total; active not disclosed | Cache-hit input: $0.30 |
| Z.ai (GLM)Z.ai | GLM-5.2 | $1.40 | $4.40 | Not disclosed | Cached input: $0.26 |
Prices can change without notice. Gemini's row is for prompts up to 200K tokens; longer prompts cost more. DeepSeek's row shows its peak, cache-miss rate. Batch, flex, priority, regional, long-context, tool-use and tax charges may differ. Follow each linked model name to verify current official pricing.
How to read the numbers
Input tokens
Everything sent to the model: your prompt, system instructions, conversation history and retrieved context.
Output tokens
Everything the model generates. Output commonly costs more, and some providers also bill hidden reasoning tokens as output.
Parameters
The learned weights inside a model. A larger count does not guarantee a better result, and an MoE model may activate only part of its total for each token.
Why most parameter counts say “not disclosed”
OpenAI, Anthropic and Google do not publish parameter counts for their current proprietary flagships. DeepSeek and Z.ai have published architecture details for some open releases, but that does not justify assigning an old model's count to a newer API model. This guide reports a number only when the provider ties it directly to the exact model in the table.
Kimi K3 is the exception here: Moonshot AI describes it as a 2.8-trillion-parameter model. Even then, total parameters are not directly comparable with an always-dense model because modern mixture-of-experts architectures route each token through a subset of the network.
A quick overview of the six families
ChatGPT / GPT
OpenAI's consumer product is ChatGPT; developers select GPT models through the API. The family is a broad default for reasoning, coding, multimodal work and agent workflows.
Claude
Anthropic's Claude family is widely used for writing, coding, analysis and long-running agents. Opus targets the hardest work, while Sonnet and Haiku trade capability for lower cost and latency.
Gemini
Google's Gemini models connect naturally with a multimodal and Google-centered ecosystem. Pro targets demanding tasks; Flash variants are designed for speed and volume.
DeepSeek
DeepSeek competes aggressively on API price and publishes open model releases alongside its hosted service. Time-of-day and cache pricing make workload shape especially important.
Kimi
Moonshot AI's Kimi focuses on long context, reasoning, knowledge work and coding. Kimi K3 offers a one-million-token context window, but its output rate means verbose workflows should be budgeted carefully.
Z.ai / GLM
Z.ai provides the GLM model family, including paid flagship and lower-cost variants. It is worth evaluating when agentic coding, long context and open-model availability matter.
Choose with a real workload, not one benchmark
Build a small evaluation set from your own prompts. Measure correctness, instruction following, tool-call success, latency and total tokens consumed. Then calculate the cost per successful task—not merely the price of one million input tokens. That comparison survives model launches better than any leaderboard snapshot.
Update policy
This is a living guide. We update the representative models, standard token prices and disclosed specifications as vendors release material changes. Preview models may be replaced before their names settle; every table row therefore links to its provider's pricing documentation.
