Live AI API Pricing & LLM Cost Calculators
CostPerPrompt is a live AI API pricing tracker covering 232+ models from vendors such as OpenAI, Anthropic, Google, Meta, DeepSeek, Alibaba, Cohere, Mistral, xAI, and more. It shows per-token input/output prices, cached input prices, and context window sizes, with data refreshed automatically (last update 2026-08-02).
The tool solves the problem of fragmented and frequently changing LLM pricing by pulling per-token rates from public listings into a single sortable table, then converting those rates into estimated monthly bills using built-in calculators that support caching and batch scenarios.
It's built for developers, product teams, and budget owners who need to compare model costs across vendors and plan API spending accurately.
Browse per-token input/output prices for 232+ models, including cached input pricing and context windows, all updated automatically.
Filter models by provider – OpenAI, Anthropic, Google, Meta, DeepSeek, Alibaba, Cohere, Mistral, xAI, and more – to quickly narrow down options.
Enter requests per day and token usage per request to get an estimated monthly bill based on real per-token prices.
The full calculator accounts for cached input pricing and batch scenarios, giving a more realistic view of production costs.
Sort models by input or output price per million tokens to find the cheapest option for your workload.
See each model's context length (e.g., 1M, 262K, 128K) alongside pricing, helping you pick a model with the right capacity.
Prices are refreshed automatically. The page shows the last update date as 2026-08-02.
CostPerPrompt lists 232+ models from vendors including OpenAI, Anthropic, Google, Meta, DeepSeek, Alibaba, Cohere, Mistral, Moonshot, NVIDIA, xAI, and Z.ai.
Yes. The full calculator includes caching and batch options, allowing you to estimate a monthly bill that reflects real-world usage scenarios.
Yes. The site provides an 'All vendors' dropdown with providers such as Alibaba, Amazon, Anthropic, Cohere, DeepSeek, Google, Meta, MiniMax, Mistral, Moonshot, NVIDIA, OpenAI, Z.ai, and xAI.
Each model listing shows input price per 1M tokens, output price per 1M tokens, cached input price per 1M tokens (where available), and context window size.