Technical primer
Building an LLM
How text becomes tokens, training turns prediction error into learned weights, embeddings organize those patterns, and inference generates one token at a time.
01
Overview
A large language model is a pattern-prediction machine. It receives tokens, uses learned weights to estimate a probability distribution over possible next tokens, selects one, appends it to the context, and repeats.
Building that machine requires four different operations that are easy to blur together:
- Input: a tokenizer converts text into token IDs.
- Training: prediction error changes the model's weights.
- Model and embeddings: the learned weights represent statistical regularities in language.
- Inference: fixed weights turn a prompt into one next-token prediction after another.
Core thesis
Training changes the weights. Inference uses the weights.
What the model learns is a structure of conditional patterns: given this context, which token is likely to come next?
This is not a complete recipe for producing a frontier model. It leaves out data acquisition, distributed infrastructure, architecture search, post-training, evaluation, and deployment. Its purpose is narrower: to make the computational loop legible from input to generated text.
02
1. Input: Text Becomes Tokens
The model does not receive words directly. A tokenizer divides text into units from a fixed vocabulary and assigns each unit an integer ID. A token may be a word, part of a word, punctuation, whitespace, or a byte-level fragment.
One influential family uses byte-pair encoding. The tokenizer learns frequently recurring symbol pairs and gives them reusable vocabulary entries. The same tokenizer must encode the prompt and decode generated token IDs back into text.
text → tokenizer → token IDs → tokenizer decoder → text
Tokenization is reversible, but it is not semantically neutral. Different vocabularies divide the same sentence differently, changing sequence length and the units whose relationships the model must learn.
For a sequence of tokens x₁, x₂, …, xₙ, a decoder language model estimates:
P(xₜ | x₁, x₂, …, xₜ₋₁)
Every position asks the same question: given the tokens so far, what token came next in the training text?
03
2. Training: Prediction Error Changes Weights
During pretraining, almost every token becomes the answer to a prediction made from its preceding context. Given The cat sat on the …, the model assigns probabilities to possible continuations.
| Possible next token | Probability |
|---|---|
mat | 70% |
floor | 15% |
chair | 5% |
| everything else | 10% |
If the observed token is mat,
cross-entropy loss
measures how much probability the model assigned to it: loss = -ln P(observed token). A
high probability produces a small loss; a low probability produces a large one.
Training onlyTraining only · not inference
How an LLM learns
Walk through the training loop: corpus tokens become predictions, predictions become loss, and loss drives gradient updates that reshape the parameters.
token → embedding
Tokens enter as embeddings
Corpus tokens are looked up in the learned embedding table, producing vectors the model can process.
Parameter changes happen only during training — ordinary inference never backpropagates.
01 / 6
Keyboard shortcuts
- Space play / pause
- ← / → step back / forward
- R reset
- H toggle this help
- The loss function measures the prediction error.
- Backpropagation identifies how parameters contributed to it.
- The optimizer adjusts those parameters to improve future predictions.
context → token probabilities → observed token → loss → backpropagation → updated weights
Repeated across an enormous body of text, this process adjusts billions of parameters. It does not store the corpus as a searchable collection of sentences. It distills statistical regularities into weights that make some continuations more likely than others.
Instruction tuning, preference optimization, and other post-training methods can change which behaviors the model reliably produces. They still work by changing parameters before ordinary inference begins.
04
3. The Model: Learned Weights and Embeddings
The trained model combines an architecture with learned parameters:
- embeddings map token IDs into learned numerical representations;
- attention combines information from different positions in the context;
- feed-forward layers transform each contextualized representation; and
- an output projection turns the final representation into one logit per vocabulary token.
For vocabulary V and embedding width d, a learned table
E ∈ ℝ^{|V|×d} maps token ID xᵢ to vector
eᵢ = E[xᵢ] ∈ ℝᵈ. The rows begin as arbitrary values and acquire predictive structure
during training.
Repeated contexts give the space geometry. Tokens used in similar contexts often develop nearby or directionally related embeddings. The explorer compresses a much larger space into three hand-authored teaching dimensions. Its vector arithmetic is geometric intuition, not a measured identity from a production model. That geometry is useful, but it is not a complete theory of meaning and does not establish that a relationship is true in the world.
An input embedding is only the starting state. Position, attention, and feed-forward layers transform it into a contextual hidden state. The token stable therefore produces different states in stable counting sort and stable employment.
token ID → input embedding → contextual hidden state → output logits → next-token probabilities
Embeddings and weights are learned parameters that persist across requests. Contextual hidden states and probabilities are temporary activations computed for the current prompt.
05
4. Inference: One Token at a Time
At inference time, training has stopped and the model's weights are fixed. A prompt moves through a repeated forward-pass loop: tokenize, embed, contextualize, project to logits, normalize to probabilities, select a token, append it, and run the model again.
Decoder-only model · token-by-token
Watch an LLM generate
Pick an example, press play, and follow every step of inference as the model predicts one token after another — attention, feed-forward, logits, and the decoding choice.
The capital of France is
token + position
The newest context token becomes a vector
The tokenizer's latest piece of context is looked up in the learned embedding table and combined with position information so identical words at different positions can behave differently.
Keyboard shortcuts
- Space play / pause
- ← / → step back / forward
- N next token
- G skip to end of generation
- R reset
- H toggle this help
prompt → tokens → representations → logits → probabilities → selected token → append → repeat
Prediction and selection are different operations. The model produces logits; softmax converts them into a distribution; a decoding strategy decides what to do with that distribution.
Decoder-only model · next-token choice
Decoding strategies
The same logits can produce different next tokens depending on the decoding rule. Compare greedy, temperature, top-k, and top-p on one fixed distribution.
Always select the token with the highest probability. Deterministic, but it can lock the model into repetitive or flat continuations.
Greedy always selects the most probable token, so no sampling is involved.
Greedy decoding selects the most probable token. Temperature, top-k, and top-p can reshape or restrict the available choices before sampling. The same model and prompt can therefore produce different continuations without changing a learned weight. Runtime context and decoding determine which learned patterns are activated and how the resulting probabilities become text.
06
5. What Pattern Prediction Means
Calling an LLM a pattern-prediction machine describes its operating contract; it does not imply that its behavior must be trivial. Learned patterns can include syntax, genre, factual associations, algorithms, explanations, tool-use conventions, and long sequences that resemble deliberate reasoning. All of them are expressed through conditional next-token probabilities.
The model does not retrieve one predetermined answer from its weights. It reconstructs a continuation from the prompt, learned parameters, temporary activations, and decoding rule. Small changes in any of those can change the generated path.
Inference also does not establish that a continuation is correct, meaningful, or worth adopting. As Understanding and Bottlenecks argues, generation produces a candidate artifact; people and institutions still have to ground, interpret, evaluate, and absorb it.
human language → tokens → training examples → prediction error → learned weights and embeddings → prompt-conditioned activations → next-token probabilities → decoded tokens → generated text
The machinery is remarkably capable, but its basic operation remains stable: learn patterns by predicting tokens, then use those learned patterns to predict again.
07
Sources
- Philip Gage, “A New Algorithm for Data Compression” (1994). Introduces byte-pair encoding as a lossless compression technique.
- Rico Sennrich, Barry Haddow, and Alexandra Birch, “Neural Machine Translation of Rare Words with Subword Units” (2016). Adapts byte-pair encoding to subword tokenization.
- Yoshua Bengio, Réjean Ducharme, Pascal Vincent, and Christian Jauvin, “A Neural Probabilistic Language Model” (2003). Connects conditional word prediction with learned distributed representations.
- Tomas Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean, “Efficient Estimation of Word Representations in Vector Space” (2013). Introduces efficient architectures for learning word-vector relationships.
- Ashish Vaswani and colleagues, “Attention Is All You Need” (2017). Introduces the Transformer architecture underlying modern decoder language models.
- Claude E. Shannon, “Prediction and Entropy of Printed English” (1951). Uses next-character prediction to estimate the redundancy of English.