DevDock
All toolsEveryday

LLM Token Calculator

Exact OpenAI token counts (tiktoken) plus $ estimates from published rates. Nothing uploaded.

Tokens exact (tiktoken) · $ estimates as of 2026-08-11

Mode

Model

Output tokens

OpenAI only. Claude/Gemini omitted — different tokenizers. Verify $ on OpenAI pricing before budgeting. Nothing uploaded.

Sponsored

Ad space

How to count LLM tokens

  1. Paste plain text, or switch to Chat messages and paste a JSON array.
  2. Pick the OpenAI model (encoding is shown in the selector).
  3. Optional: set expected output tokens to estimate completion cost.
  4. Read input tokens, context %, and $/request — compare models in the table.

Exact counts, honest costs

In 2026–2028 the hard part is not “how many words” — it is which tokenizer and which price row. This tool runs OpenAI's BPE encodings locally so GPT-4o / GPT-4.1 / GPT-5 / o-series / embedding prompts get exact token counts. Dollar figures multiply those counts by published per-million rates (last aligned 2026-08-11) — always recheck OpenAI before budgeting. We do not pretend Claude or Gemini are tiktoken.

Need a checksum instead of tokens? Try the Hash Generator. Encoding payloads for APIs? Use Base64 Encode / Decode.

FAQ

Are token counts exact?

Yes for the OpenAI models listed. Counting uses official tiktoken encodings (o200k_base / cl100k_base) via gpt-tokenizer in your browser.

Are the dollar amounts exact?

No — they are estimates. Exact tokens × published $/1M rates (table last aligned 2026-08-11). OpenAI changes prices; confirm on the official pricing page before budgeting.

Why isn’t Claude or Gemini included?

They use different tokenizers. Using OpenAI tiktoken for them is often wrong by 10–20%+. We only show counts we can make exact client-side.

What does Chat messages mode do?

It counts a JSON array of { role, content } with ChatML overhead (encodeChat) for the selected model — closer to chat/completions prompt billing than raw string length.

Why can output cost dominate?

You enter assumed completion tokens. Reasoning models (o3, o4-mini, …) also bill hidden reasoning tokens as output, so real bills can exceed the visible reply.

Is my prompt uploaded?

No. Tokenization runs entirely in your browser.