Tokenization

토큰화

Split input text into the token pieces a model processes.

···
html
<div class="scene"><div class="heading">TEXT → PIECES</div><div class="source" id="source">AI builds tools.</div><div class="parts" id="parts"></div><div class="count mono" id="count">0 tokens</div></div>
css
.scene{width:min(94vw,800px);height:min(88vh,326px);padding:clamp(10px,2.7vmin,19px);border:1px solid var(--line);border-radius:13px;background:var(--surface);font:500 clamp(13px,2.4vmin,17px)/1.3 var(--font-sans,sans-serif);position:relative;overflow:hidden}.scene .mono{font-family:ui-monospace,SFMono-Regular,monospace}.scene .muted{color:var(--muted)}.scene .accent{color:var(--accent)}.scene .heading{font-weight:750;color:var(--accent);margin-bottom:clamp(5px,1.6vmin,12px)}.source{padding:10px;border-bottom:1px solid var(--line);font-size:1.15em}.parts{height:46%;display:flex;align-items:center;justify-content:center;gap:5px;flex-wrap:wrap}.part{border:1px solid var(--accent);color:var(--accent);background:color-mix(in srgb,var(--accent) 10%,var(--surface));padding:7px;border-radius:5px;animation:rise .35s both}.count{text-align:right;color:var(--muted)}@keyframes rise{from{transform:translateY(12px);opacity:0}to{transform:translateY(0);opacity:1}}
js
const samples=[['AI builds tools.',['AI',' builds',' tools','.']],['hello-world',['hello','-','world']]];let i=0;function draw(){const [source,parts]=samples[i];document.getElementById('source').textContent=source;document.getElementById('parts').innerHTML=parts.map((p,n)=>'<span class=part style=animation-delay:'+n*90+'ms>'+p+'</span>').join('');document.getElementById('count').textContent=parts.length+' example pieces';i=1-i}draw();setInterval(draw,2400)

Tokenization maps text to pieces and IDs from a model vocabulary. A word is not necessarily one token; spaces, symbols, and language affect segmentation.

The demo breaks sample text into pieces and counts them. Exact boundaries depend on the model tokenizer and may differ from this illustration.

When to use

Check tokens when estimating input size, cost, or chunk boundaries.

Open as page ↗