| ISSUE 05/30 · COMPUTER SCIENCE |
~3 MIN |
FOUNDATION / LANGUAGE MODELS · WEEK 01 — HIDDEN RULES
One token, then another
START HERE
A computer does not handle written language as one smooth stream of meaning. It first divides the text into manageable pieces. Each piece is called a token. A token can be a whole word, part of a word, punctuation, or even a space-like marker. A language model is an AI program trained to continue text. It writes by choosing one next token, adding it, and repeating the choice.
|
/ WHY NOW
Watch a chatbot answer and the words stream out from left to right. That is not just a visual flourish. Modern language models usually generate by repeatedly predicting the next token—a word, word fragment, punctuation mark or other text unit. Because each new piece normally depends on the pieces already written, a long answer requires many prediction steps in sequence.
|
/ THE IDEA
The model reads the tokens so far and assigns possible next tokens different probabilities. It selects one, appends it, then asks the same kind of question again with the longer history. It is closer to extremely informed autocomplete than to retrieving a finished paragraph from a filing cabinet. The surprising capability comes from the patterns learned and the scale of the computation, not from changing the basic next-step loop.
THE FORMAL IDEA
P(next token | tokens so far)
| P = a set of percentages assigned across possible next tokens, not one guaranteed answer | | the vertical bar means given | | tokens so far = the prompt plus everything already generated |
|
RUN THE TINY EXAMPLE
Change the context, change the odds
After “The capital of France is”: Paris 92%, Lyon 2%, other 6% After “The capital of Italy is”: Rome 94%, Milan 2%, other 4% Choose one token, append it, then calculate a new set of percentages
|
These percentages are illustrative, and exact token pieces vary between systems. The changed place name changes the next-token distribution, and every selected token changes the following calculation.
/ SO WHAT?
This predicts two things you can notice. Longer answers require more sequential generation steps, and randomness settings can produce different valid continuations from the same prompt. It also explains why checking the final result matters: locally plausible next steps can accumulate into a globally wrong answer.
ONE CAVEAT |
| Next-token prediction describes the output mechanism, not a complete theory of what a trained model represents internally. It does not mean the system merely copies the most common phrase. |
KEEP THIS
A language model builds an answer by turning one next-token probability problem into a chain.
|
NEXT: How a ripple in space calibrates its detector
|