Tokenization maps text to pieces and IDs from a model vocabulary. A word is not necessarily one token; spaces, symbols, and language affect segmentation.
The demo breaks sample text into pieces and counts them. Exact boundaries depend on the model tokenizer and may differ from this illustration.
When to use
Check tokens when estimating input size, cost, or chunk boundaries.