Tokens

The model has never seen a word. It sees tokens: short chunks of letters. Watch your sentence get cut into them below.

A model never sees your words. A tokenizer cuts your text into tokens first: short chunks of letters that come up often. The model reads those chunks. Nothing else.

Why pieces

A list of every word misses names, code, and typos. A short list of pieces covers everything. Common words stay whole. Rare ones split into pieces the list has.

So “strawberry” is roughly straw + berry. Ask a model to count the r’s and it cannot see letters. It sees pieces, and it must reason across them. The model is not careless. It never saw the word you think it saw.

Like a new hire’s notes in shorthand. Common phrases come through whole. Rare ones arrive in parts.

You pay by the piece

Billing and capacity are both counted in tokens. The rough rule: 100 tokens ≈ 75 words. A “128,000-token context” is the size of the model’s desk. Fill the desk and the earliest pieces fall off the edge.

Like a phone plan that charges per megabyte. The bill follows the size of what you send.
A lookalike, not the real tokenizer. Real ones are trained on text, so their pieces and counts differ. This shows the shape of the idea.

Working on this?

If AI is stuck somewhere in your business, tell me where. I read every one of these.