Docs
LLM API cost estimation: input, output and cached tokens
Estimate a text API workload with separate input, output and cache rates. Run an offline Python example, then reconcile the estimate with your account’s usage.
Start with the request your product actually sends
For simple text token billing, multiply each token category by its rate and add the results. If rates are per million tokens, divide each count by 1,000,000 before multiplying. An input rate alone does not describe the cost of a complete response.
Include system instructions, conversation history, retrieved passages and tool definitions in the input. Use the model’s reported usage where available; character counts are only a rough planning aid. An agent may make several model calls for one user task, so budget all calls that contribute to it.
In EasyAI, start from the model’s public pricing page, then confirm your account’s rate and billing unit in the console. Public reference rates can differ from account rates. This guide covers text tokens; use the model’s billing rules for image, audio, video or other metered inputs.
Separate cached input before applying the rates
The calculation below assumes total input tokens already include cache-read tokens. Subtract cached input once, charge the remaining input at the regular rate, and charge cache reads at the cache-read rate. Check your provider’s field definitions before applying this convention.
Cache writes may have a separate price. Cache eligibility, expiry and request requirements also differ by model. Do not apply a cache-read discount to every repeated prompt. Add cache writes, tools, searches or other charges separately in their documented units; this sample excludes them.
OpenRouter’s usage-accounting guide is an example of explicit token and cache fields. Those fields describe OpenRouter. They do not establish that EasyAI returns the same cost fields or supports the same caching behavior for every model.
Run an estimate without making an API call
Save the example as token_budget.py and run python token_budget.py. It uses the Python standard library and sends no requests. All rates are hypothetical USD per million tokens, not live EasyAI or provider prices. Replace them with your account’s rates in a consistent currency.
For 10,000 requests with 2,000 total input tokens and 500 output tokens each, the no-cache estimate is $35.00. With 1,000 eligible cached input tokens per request at the example cache-read rate, it is $27.00. These are arithmetic examples; real request lengths vary.
from decimal import Decimal
def estimate(input_tokens, cached_tokens, output_tokens, requests,
input_rate, cache_rate, output_rate):
counts = (input_tokens, cached_tokens, output_tokens, requests)
if any(type(n) is not int or n < 0 for n in counts):
raise ValueError("Counts must be non-negative integers")
if cached_tokens > input_tokens:
raise ValueError("Cached input cannot exceed total input")
rates = tuple(Decimal(str(r)) for r in (input_rate, cache_rate, output_rate))
if any(not r.is_finite() or r < 0 for r in rates):
raise ValueError("Rates must be finite and non-negative")
regular, cached, output = rates
per_request = (
(input_tokens - cached_tokens) * regular
+ cached_tokens * cached
+ output_tokens * output
) / Decimal(1_000_000)
return per_request * requests
# Hypothetical USD rates per million tokens, not a live price list.
rates = dict(input_rate="1.00", cache_rate="0.20", output_rate="3.00")
for cached in (0, 1000):
total = estimate(2000, cached, 500, 10_000, **rates)
print(f"{cached} cached input tokens/request: USD {total:.2f}")Turn the example into a workload forecast
Sample short and long conversations, retrieval requests and agent tasks. Record the model ID, date, input, output, cache reads and any other billable categories. Keep separate rows when models, rates or billing settings change.
Estimate each workload separately and add the totals. A chatbot with short answers has a different cost mix from an agent that repeatedly sends a long history. Track completed tasks and total model calls so retries and follow-up calls remain in the forecast.
Prepare an expected scenario and a heavier mix of long prompts or outputs. A configured output limit is a ceiling, not a prediction of average output. Save the assumptions beside the calculation so another teammate can reproduce it.
Reconcile with a small paid test
Before increasing volume, test a small workload within a budget you choose. Record response usage, request identifiers when provided and the corresponding account usage records. Keep API keys, private prompts and customer information out of shared worksheets.
If the estimate differs from the charge, compare model ID, account rate, currency, billing unit and request settings. Then check cache writes, extra calls, reasoning or other billed token categories, and non-token charges where applicable. A failed or timed-out request is not necessarily free; check the final state and recorded charge.
Use the account charge for reconciliation and retain the calculation as a forecast. For unresolved differences, share the non-sensitive request identifiers, time and model with support. For recurring volume, bring the measured workload and purchasing period to a quote discussion.
Sources
Reviewed .
Ready to build?
Create a key, check the enabled models and follow the quickstart for your first live request.
Get an API key