Evergreen LLM guide

The leading large language models, compared

A practical overview of ChatGPT, Claude, Gemini, DeepSeek, Kimi and Z.ai's GLM family—what they are, what their APIs cost and why parameter counts rarely tell the whole story.

9 min read

The short version

There is no single “best LLM.” The right choice depends on the job, the quality and latency you need, the tools and context the model must use, and the cost of its complete workflow. The table below uses each provider's representative flagship API model so that prices are easier to compare.

Flagship LLM price comparison

USD per one million tokens, standard API processing. Consumer chat subscriptions are not included.

LAST CHECKED 17 AUG 2026

LLMAPI modelInput / 1MOutput / 1MParametersPricing note
ChatGPTOpenAIGPT-5.6 Sol$5.00$30.00Not disclosedCached input: $0.50
ClaudeAnthropicClaude Opus 5$5.00$25.00Not disclosedCache hits: $0.50
GeminiGoogleGemini 3.1 Pro Preview$2.00$12.00Not disclosedPrompts up to 200K tokens
DeepSeekDeepSeekDeepSeek V4 Pro$1.32$3.96Not disclosedPeak rate; off-peak is 50% less
KimiMoonshot AIKimi K3$3.00$15.002.8T total; active not disclosedCache-hit input: $0.30
Z.ai (GLM)Z.aiGLM-5.2$1.40$4.40Not disclosedCached input: $0.26

Prices can change without notice. Gemini's row is for prompts up to 200K tokens; longer prompts cost more. DeepSeek's row shows its peak, cache-miss rate. Batch, flex, priority, regional, long-context, tool-use and tax charges may differ. Follow each linked model name to verify current official pricing.

How to read the numbers

Input tokens

Everything sent to the model: your prompt, system instructions, conversation history and retrieved context.

Output tokens

Everything the model generates. Output commonly costs more, and some providers also bill hidden reasoning tokens as output.

Parameters

The learned weights inside a model. A larger count does not guarantee a better result, and an MoE model may activate only part of its total for each token.

Why most parameter counts say “not disclosed”

OpenAI, Anthropic and Google do not publish parameter counts for their current proprietary flagships. DeepSeek and Z.ai have published architecture details for some open releases, but that does not justify assigning an old model's count to a newer API model. This guide reports a number only when the provider ties it directly to the exact model in the table.

Kimi K3 is the exception here: Moonshot AI describes it as a 2.8-trillion-parameter model. Even then, total parameters are not directly comparable with an always-dense model because modern mixture-of-experts architectures route each token through a subset of the network.

A quick overview of the six families

ChatGPT / GPT

OpenAI's consumer product is ChatGPT; developers select GPT models through the API. The family is a broad default for reasoning, coding, multimodal work and agent workflows.

Claude

Anthropic's Claude family is widely used for writing, coding, analysis and long-running agents. Opus targets the hardest work, while Sonnet and Haiku trade capability for lower cost and latency.

Gemini

Google's Gemini models connect naturally with a multimodal and Google-centered ecosystem. Pro targets demanding tasks; Flash variants are designed for speed and volume.

DeepSeek

DeepSeek competes aggressively on API price and publishes open model releases alongside its hosted service. Time-of-day and cache pricing make workload shape especially important.

Kimi

Moonshot AI's Kimi focuses on long context, reasoning, knowledge work and coding. Kimi K3 offers a one-million-token context window, but its output rate means verbose workflows should be budgeted carefully.

Z.ai / GLM

Z.ai provides the GLM model family, including paid flagship and lower-cost variants. It is worth evaluating when agentic coding, long context and open-model availability matter.

Choose with a real workload, not one benchmark

Build a small evaluation set from your own prompts. Measure correctness, instruction following, tool-call success, latency and total tokens consumed. Then calculate the cost per successful task—not merely the price of one million input tokens. That comparison survives model launches better than any leaderboard snapshot.

Update policy

This is a living guide. We update the representative models, standard token prices and disclosed specifications as vendors release material changes. Preview models may be replaced before their names settle; every table row therefore links to its provider's pricing documentation.