What is a token
A token is a chunk of text, typically 3–4 characters. On average 1000 tokens ≈ 750 English words. Cyrillic and Cyrillic-script languages consume 40–60% more tokens per word.
Providers bill separately for input tokens (what you send) and output tokens (what the model writes). Output is usually 3–5× more expensive.
2026 price table (USD)
| Model | Input / 1K | Output / 1K | Sweet spot |
|---|---|---|---|
| GPT‑5 | $0.00125 | $0.010 | general |
| GPT‑5 mini | $0.00025 | $0.002 | bulk work |
| Claude Sonnet 4.5 | $0.003 | $0.015 | code, engineering |
| Claude Haiku 4 | $0.0008 | $0.004 | fast + smart |
| Gemini 2.5 Pro | $0.00125 | $0.010 | long context |
| Gemini 2.5 Flash | $0.0003 | $0.0025 | volume |
| Llama 3.3 70B | $0.00035 | $0.0004 | self-hosted alt |
| DeepSeek V3 | $0.00027 | $0.0011 | cheapest solid |
Real-world example
You feed the model a 2000-word article (~2700 tokens) and ask for a 300-word summary (~400 output):
- GPT‑5: $0.0034 + $0.004 ≈ $0.007
- Claude Sonnet 4.5: $0.008 + $0.006 ≈ $0.014
- Gemini 2.5 Flash: $0.0008 + $0.001 ≈ $0.002
One summary costs less than a US cent even on the flagship. A thousand summaries — $2 to $14.
How to save
- Short system prompts. Don't say "you are a world-class assistant…" — say "be concise".
- Simple tasks → Flash / Haiku / DeepSeek. Complex → Sonnet / GPT‑5.
- Don't upload the whole document — clip the relevant section.
- Prompt caching (Claude and Gemini) gives up to 90% off repeated system prompts.
Don't want to manage paid accounts?
AskHub bills in credits — 1 credit = $0.0001. Packs from $5.
Start free →