Text Tools

How Many Words Is 1,000 Tokens? Estimating AI Token Counts From Text

A token is roughly four characters or three-quarters of a word in English. Here is the quick reference table, how to estimate from your own text in seconds, and the five cases where the ratio breaks badly.

September 12, 2026 7 min read Toolio Editorial
How Many Words Is 1,000 Tokens? Estimating AI Token Counts From Text
Summarize with:
Share:

Every AI API bills in tokens, every context window is measured in tokens, and every rate limit is expressed in tokens — but nobody writes in tokens. If you need to know whether your document fits, or what a prompt will cost, you need a way to get from words to tokens without running a tokeniser.

Direct Answer: For ordinary English text, 1,000 tokens is roughly 750 words or about 4,000 characters — OpenAI's own published rule of thumb is approximately 4 characters or 0.75 words per token. Inverted: multiply your word count by about 1.33 to estimate tokens, or divide your character count by 4. The estimate holds well for prose and breaks down for code, JSON, non-English languages, technical terms and heavily punctuated text, where token counts run considerably higher than the ratio predicts. For anything where cost or a context limit actually matters, estimate first to plan, then read the usage field on a real API response for the exact number.

The two rules of thumb

Rule Direction Example
1 token ≈ 0.75 words words × 1.33 = tokens 1,000 words ≈ 1,333 tokens
1 token ≈ 4 characters characters ÷ 4 = tokens 8,000 characters ≈ 2,000 tokens

The character rule is the more robust of the two. Word counting quietly assumes words are of typical English length, so it degrades on text full of long technical terms or short symbols. Character counting degrades more gracefully because tokenisers ultimately work on character sequences.

Quick reference table

Tokens Approximate words Approximate characters Roughly equivalent to
100 75 400 A short paragraph
500 375 2,000 A long email
1,000 750 4,000 A short blog post
4,000 3,000 16,000 A detailed article
16,000 12,000 64,000 A long report
128,000 96,000 512,000 A full-length novel
922,000 ~691,500 ~3,688,000 Roughly seven novels

That last row is the maximum input you can send to a current frontier model such as GPT-6 Astra, whose 922,000-token input cap sits inside a 1,050,000-token context window. The 128,000 row is that model's maximum single output — meaning it can, in principle, generate a novel's worth of text in one response.

Estimating from your own text in about ten seconds

You do not need a tokeniser for planning purposes:

  1. Paste your text into our Word Counter and read the word count. Multiply by 1.33.
  2. Or paste it into the Character Counter and divide the character count by 4.
  3. If the two estimates disagree by more than about 20%, trust the character-based one — the disagreement itself is a signal that your text is not ordinary prose.

That third step is the useful part. When the two methods diverge, it is almost always because the text is code, structured data, or a language the ratio was not derived from.

Where the ratio breaks

Code. Source code tokenises far less efficiently than prose. Indentation, brackets, operators, camelCase and snake_case identifiers all fragment into multiple tokens. Assume code runs meaningfully above the prose ratio, and estimate it by characters rather than words.

JSON and structured data. Every quote, brace, colon and comma is doing tokeniser work. A compact JSON payload can consume noticeably more tokens than prose of the same character length, and repeated key names are paid for on every single object in an array.

Non-English text. The ratio was derived from English. Languages that do not use Latin script — Hindi, Chinese, Japanese, Arabic, Russian and others — can consume substantially more tokens per character, sometimes several times more, because the tokeniser's vocabulary allocates fewer whole-word entries to them. If you are working in a non-English language, do not plan a budget on the English ratio.

Rare and technical vocabulary. Common words are usually a single token. Unusual ones split into pieces — a specialist medical, legal or chemical term can be four or five tokens on its own. Text dense with domain jargon runs high.

Whitespace and formatting. Long runs of spaces, tabs, newlines and markdown decoration all cost tokens. A heavily formatted document is more expensive than its word count suggests.

Broadly:

Content type Behaviour versus the English prose ratio
Plain English prose Matches well
Business or marketing writing Matches well
Technical documentation Slightly higher
Source code Notably higher
JSON / XML / YAML Notably higher
Non-Latin-script languages Substantially higher

What this means for cost

Token estimates turn straight into money. On GPT-6 Astra at $10 per million input tokens and $50 per million output tokens:

  • A 3,000-word input document ≈ 4,000 tokens ≈ $0.04 in input cost
  • A 1,500-word generated response ≈ 2,000 tokens ≈ $0.10 in output cost

Which surfaces the fact that catches people out: output tokens cost five times input tokens. The response is usually shorter than the prompt and still the larger share of the bill. If you want the full arithmetic including caching and batching, GPT-6 Astra API Cost Explained works through it.

When to stop estimating and count exactly

Estimation is for planning. Use exact counts when:

  • You are near a context limit. A 15% estimation error on a 900,000-token prompt is 135,000 tokens — comfortably enough to get your request rejected.
  • You are budgeting real production volume. At scale, a 20% error is a 20% budget error.
  • You are working in a non-English language or with code, where the ratio is least reliable.
  • You are debugging a cost anomaly. Estimates cannot tell you that tool definitions and conversation history are eating your input budget.

The exact number is always available: every API response includes a usage object with the real input and output token counts. That is what you are billed on, so that is what to trust. Anything else — including this article's table — is an approximation.

Frequently Asked Questions (FAQs)

How many words is 1,000 tokens?

About 750 words of ordinary English, or roughly 4,000 characters. Expect fewer words per token for code, structured data and non-English text.

How many tokens is 500 words?

About 665 tokens (500 × 1.33). Round up when you are working near a limit.

Is a token the same as a word?

No. A token is a fragment the model's tokeniser produces — often a whole common word, but frequently a word-piece. "Unbelievable" may be several tokens; "the" is one.

Do spaces and punctuation count as tokens?

They contribute to token counts, yes. Leading spaces are typically bundled with the following word, and punctuation is often its own token. This is part of why heavily formatted text costs more than its word count suggests.

Why does my Hindi or Chinese text use so many more tokens?

Tokeniser vocabularies are weighted toward the languages that dominated their training data, so non-Latin scripts tend to split into far more, far smaller pieces. Estimate these by character count and expect to be well above the English ratio.

Do input and output tokens count the same toward the context window?

They occupy the same window but are billed at different rates. On GPT-6 Astra the window is 1,050,000 tokens total, with input capped at 922,000 and output at 128,000 — and output is five times the price.

The short version

1,000 tokens ≈ 750 English words ≈ 4,000 characters. Multiply words by 1.33 or divide characters by 4, and prefer the character method whenever your text is not plain prose. Expect the estimate to run low for code, structured data and non-English languages.

For a fast estimate, paste your text into our AI Token Counter & API Cost Calculator, which classifies it by writing system rather than applying one flat ratio — useful precisely because the English rule of thumb breaks on the content types above. The Word Counter and Character Counter give you the raw counts if you would rather do the multiplication yourself. For anything that touches a budget or a hard limit, read usage on a real response instead.

The limits quoted above belong to GPT-6 Astra; GPT-6 Astra Explained sets out where those numbers come from.

Sources: OpenAI Help Center guidance on tokens and counting them (the ~4 characters / ~0.75 words per token estimate); OpenAI API model reference for gpt-6-astra (context window, input and output caps, per-token pricing). Ratios for code, structured data and non-Latin scripts are stated directionally rather than as fixed figures because they vary by tokeniser and by content.

Free Calculator

Put this guide into action

Stop guessing — use our Word Counter to run real numbers, compare scenarios, and get instant results you can trust.

Use Free Word Counter
Toolio Editorial

Toolio Editorial Senior Technical Editors & UX Content Engineers

Digital Utilities, Web Engineering & Tool Guides

The Toolio Editorial Board is dedicated to delivering clear, transparent, and accurate technical guides across digital utilities, developer tools, unit conversion standards, date-time algorithms, and decision science. The board maintains rigorous editorial standards, factual accuracy, and step-by-step clarity for every guide published.

Try Calculator Word Counter
Use Word Counter

Continue Reading