A model never sees your words. A tokenizer cuts your text into tokens first: short chunks of letters that come up often. The model reads those chunks. Nothing else.
Why pieces
A list of every word misses names, code, and typos. A short list of pieces covers everything. Common words stay whole. Rare ones split into pieces the list has.
So “strawberry” is roughly straw + berry. Ask a model to count the r’s and it cannot see letters. It sees pieces, and it must reason across them. The model is not careless. It never saw the word you think it saw.
You pay by the piece
Billing and capacity are both counted in tokens. The rough rule: 100 tokens ≈ 75 words. A “128,000-token context” is the size of the model’s desk. Fill the desk and the earliest pieces fall off the edge.
Working on this?
If AI is stuck somewhere in your business, tell me where. I read every one of these.