Skip to content
wordpiece.org

BERT WordPiece · browser-side · version-pinned

See exactly what your tokenizer does to your text.

Token counts decide whether a prompt fits, what it costs, and whether your model can read it at all. Paste text and get the real split — the ids, the character offsets, the [UNK] warnings — from the exact tokenizer version your model loads.

Text never leaves the browser 150,069 vocabulary entries verified 40 reference vectors pinned in tests

Input

Tokenizer
0 / 20,000

Runs entirely in this tab. Nothing is uploaded, stored or logged.

No text tokenized yet

Paste text and run the tokenizer to see the exact split, ids and character offsets.

How it works

Three steps, no account, no upload. The point is that the output is something you can paste into a bug report and someone else can reproduce.

  1. 01

    Paste your text

    Any language, any length up to 20,000 characters. It stays in your browser tab.

  2. 02

    Pick the tokenizer

    Choose the exact WordPiece model your pipeline loads. Two BERT tokenizers are supported, both version-pinned.

  3. 03

    Read the real output

    Tokens, ids, offsets and `[UNK]` warnings, with `[CLS]` and `[SEP]` counted so a context window means what you think.

What you get

Not a rounded estimate. The same structure your own pipeline produces, so you can compare it directly.

The exact split

Every subword piece in order, with `##` continuation markers, so you can see where a word broke and why.

Tokenize text

Ids and offsets

The integer ids a model receives, plus the character span each token came from — including the special-token mask.

Read the output format

Pinned versions

Results name the exact repo, commit and checksum of the tokenizer that produced them. No model-agnostic guesses.

See source data

Two tokenizers, and why the difference matters

The same sentence produces different tokens, different counts and different [UNK] behaviour depending on which BERT you load. That is the most common reason a count you measured somewhere else does not match yours.

Full comparison
Input bert-base-uncased multilingual-cased
tokenization 2 tokens 3 tokens
分词测试 1 known + 3 × [UNK] 4 chars, all known
BERT base uncased 4 tokens 6 tokens

Counts shown are content tokens, excluding [CLS] and [SEP]. Verify any of them yourself in the tool.

Guides

Short, specific pieces that explain what the tool shows and why it shows it that way.

All guides

Questions people actually hit

Is my text uploaded anywhere?

No. Tokenization runs entirely in JavaScript in your browser tab. The site has no upload endpoint, no analytics script and no third-party code on the page, so there is no code path that could transmit what you paste.

Are these the token counts my model uses?

Only if your model uses one of the two tokenizers listed here, at the listed commit. BERT-family models shipped with these vocabularies will agree. Models using BPE, SentencePiece or tiktoken count differently, and this site will not guess for them.

Why does my Chinese text come out as [UNK]?

The uncased English BERT vocabulary covers very few Chinese characters, so most become [UNK]. The multilingual tokenizer covers them. The tool flags [UNK] explicitly so you notice instead of silently paying for a broken token.

Why is my count different from another tokenizer website?

Usually a different tokenizer, a different revision, or a count that excludes [CLS] and [SEP]. This site reports both the content count and the total, and shows exactly which revision it loaded.

Start with the text you are actually stuck on

The fastest way to settle a tokenization argument is to paste the exact string into the tool and look at the offsets.