The exact split
Every subword piece in order, with `##` continuation markers, so you can see where a word broke and why.
Tokenize textBERT WordPiece · browser-side · version-pinned
Token counts decide whether a prompt fits, what it costs, and whether
your model can read it at all. Paste text and get the real split — the
ids, the character offsets, the [UNK]
warnings — from the exact tokenizer version your model loads.
Runs entirely in this tab. Nothing is uploaded, stored or logged.
No text tokenized yet
Paste text and run the tokenizer to see the exact split, ids and character offsets.
Nothing to tokenize
That input has no non-whitespace characters. WordPiece drops whitespace, so the encoding is only the two special tokens and no content token is charged for it.
Three steps, no account, no upload. The point is that the output is something you can paste into a bug report and someone else can reproduce.
Any language, any length up to 20,000 characters. It stays in your browser tab.
Choose the exact WordPiece model your pipeline loads. Two BERT tokenizers are supported, both version-pinned.
Tokens, ids, offsets and `[UNK]` warnings, with `[CLS]` and `[SEP]` counted so a context window means what you think.
Not a rounded estimate. The same structure your own pipeline produces, so you can compare it directly.
Every subword piece in order, with `##` continuation markers, so you can see where a word broke and why.
Tokenize textThe integer ids a model receives, plus the character span each token came from — including the special-token mask.
Read the output formatResults name the exact repo, commit and checksum of the tokenizer that produced them. No model-agnostic guesses.
See source data
The same sentence produces different tokens, different counts and
different [UNK] behaviour
depending on which BERT you load. That is the most common reason a
count you measured somewhere else does not match yours.
| Input | bert-base-uncased | multilingual-cased |
|---|---|---|
| tokenization | 2 tokens | 3 tokens |
| 分词测试 | 1 known + 3 × [UNK] | 4 chars, all known |
| BERT base uncased | 4 tokens | 6 tokens |
Counts shown are content tokens, excluding
[CLS] and
[SEP]. Verify any of them yourself in
the tool.
Short, specific pieces that explain what the tool shows and why it shows it that way.
Greedy longest-match, `##` continuations, and why `tokenization` costs two tokens in the uncased model and three in the cased one.
Read guideWhat a character span means, how `[CLS]` and `[SEP]` are represented, and where offsets stop being useful.
Read guideWhy “words” and “tokens” diverge sharply, and how to budget a prompt before you send it.
Read guideWhy text becomes `[UNK]`, which languages trigger it most, and how to pick a tokenizer that fits your input.
Read guideNo. Tokenization runs entirely in JavaScript in your browser tab. The site has no upload endpoint, no analytics script and no third-party code on the page, so there is no code path that could transmit what you paste.
Only if your model uses one of the two tokenizers listed here, at the listed commit. BERT-family models shipped with these vocabularies will agree. Models using BPE, SentencePiece or tiktoken count differently, and this site will not guess for them.
The uncased English BERT vocabulary covers very few Chinese characters, so most become [UNK]. The multilingual tokenizer covers them. The tool flags [UNK] explicitly so you notice instead of silently paying for a broken token.
Usually a different tokenizer, a different revision, or a count that excludes [CLS] and [SEP]. This site reports both the content count and the total, and shows exactly which revision it loaded.
The fastest way to settle a tokenization argument is to paste the exact string into the tool and look at the offsets.