← All writing

Technical primer

Building an LLM

How text becomes tokens, training turns prediction error into learned weights, embeddings organize those patterns, and inference generates one token at a time.

01

Overview

A large language model is a pattern-prediction machine. It receives tokens, uses learned weights to estimate a probability distribution over possible next tokens, selects one, appends it to the context, and repeats.

Building that machine requires four different operations that are easy to blur together:

  1. Input: a tokenizer converts text into token IDs.
  2. Training: prediction error changes the model's weights.
  3. Model and embeddings: the learned weights represent statistical regularities in language.
  4. Inference: fixed weights turn a prompt into one next-token prediction after another.

Core thesis

Training changes the weights. Inference uses the weights.

What the model learns is a structure of conditional patterns: given this context, which token is likely to come next?

This is not a complete recipe for producing a frontier model. It leaves out data acquisition, distributed infrastructure, architecture search, post-training, evaluation, and deployment. Its purpose is narrower: to make the computational loop legible from input to generated text.

02

1. Input: Text Becomes Tokens

The model does not receive words directly. A tokenizer divides text into units from a fixed vocabulary and assigns each unit an integer ID. A token may be a word, part of a word, punctuation, whitespace, or a byte-level fragment.

One influential family uses byte-pair encoding. The tokenizer learns frequently recurring symbol pairs and gives them reusable vocabulary entries. The same tokenizer must encode the prompt and decode generated token IDs back into text.

text → tokenizer → token IDs → tokenizer decoder → text

Tokenization is reversible, but it is not semantically neutral. Different vocabularies divide the same sentence differently, changing sequence length and the units whose relationships the model must learn.

For a sequence of tokens x₁, x₂, …, xₙ, a decoder language model estimates:

P(xₜ | x₁, x₂, …, xₜ₋₁)

Every position asks the same question: given the tokens so far, what token came next in the training text?

03

2. Training: Prediction Error Changes Weights

During pretraining, almost every token becomes the answer to a prediction made from its preceding context. Given The cat sat on the …, the model assigns probabilities to possible continuations.

Possible next tokenProbability
mat70%
floor15%
chair5%
everything else10%

If the observed token is mat, cross-entropy loss measures how much probability the model assigned to it: loss = -ln P(observed token). A high probability produces a small loss; a low probability produces a large one.

Training onlyTraining only · not inference

How an LLM learns

Walk through the training loop: corpus tokens become predictions, predictions become loss, and loss drives gradient updates that reshape the parameters.

Loss 4.31Step 01 / 06
01

token → embedding

Tokens enter as embeddings

Corpus tokens are looked up in the learned embedding table, producing vectors the model can process.

ProducesInput vectors
Token embeddings
Forward pass → prediction
Loss L = −log p(target)
Gradients ∂L/∂θ
Optimizer: θ ← θ − η·∂L/∂θ
Updated parameters · epoch 1

Parameter changes happen only during training — ordinary inference never backpropagates.

01 / 6

e1
Keyboard shortcuts
  • Space play / pause
  • ← / → step back / forward
  • R reset
  • H toggle this help
Simplified explanatory model · deterministic illustrative values · backpropagation belongs to training, never to ordinary inference
  1. The loss function measures the prediction error.
  2. Backpropagation identifies how parameters contributed to it.
  3. The optimizer adjusts those parameters to improve future predictions.

context → token probabilities → observed token → loss → backpropagation → updated weights

Repeated across an enormous body of text, this process adjusts billions of parameters. It does not store the corpus as a searchable collection of sentences. It distills statistical regularities into weights that make some continuations more likely than others.

Instruction tuning, preference optimization, and other post-training methods can change which behaviors the model reliably produces. They still work by changing parameters before ordinary inference begins.

04

3. The Model: Learned Weights and Embeddings

The trained model combines an architecture with learned parameters:

  • embeddings map token IDs into learned numerical representations;
  • attention combines information from different positions in the context;
  • feed-forward layers transform each contextualized representation; and
  • an output projection turns the final representation into one logit per vocabulary token.

For vocabulary V and embedding width d, a learned table E ∈ ℝ^{|V|×d} maps token ID xᵢ to vector eᵢ = E[xᵢ] ∈ ℝᵈ. The rows begin as arbitrary values and acquire predictive structure during training.

Semantic composition

Move through a word space

3D semantic network
=king
Compose up to four terms
The 3D semantic network loads when this explorer enters the viewport.

Repeated contexts give the space geometry. Tokens used in similar contexts often develop nearby or directionally related embeddings. The explorer compresses a much larger space into three hand-authored teaching dimensions. Its vector arithmetic is geometric intuition, not a measured identity from a production model. That geometry is useful, but it is not a complete theory of meaning and does not establish that a relationship is true in the world.

An input embedding is only the starting state. Position, attention, and feed-forward layers transform it into a contextual hidden state. The token stable therefore produces different states in stable counting sort and stable employment.

token ID → input embedding → contextual hidden state → output logits → next-token probabilities

Embeddings and weights are learned parameters that persist across requests. Contextual hidden states and probabilities are temporary activations computed for the current prompt.

05

4. Inference: One Token at a Time

At inference time, training has stopped and the model's weights are fixed. A prompt moves through a repeated forward-pass loop: tokenize, embed, contextualize, project to logits, normalize to probabilities, select a token, append it, and run the model again.

Decoder-only model · token-by-token

Watch an LLM generate

Pick an example, press play, and follow every step of inference as the model predicts one token after another — attention, feed-forward, logits, and the decoding choice.

InferenceToken 01 / 03 · stage 01 / 06
Generated text

The capital of France is

01

token + position

The newest context token becomes a vector

The tokenizer's latest piece of context is looked up in the learned embedding table and combined with position information so identical words at different positions can behave differently.

ProducesEmbedding vector
Stage TokensDeterministic illustrative trace
p0Thep1 capitalp2 ofp3 Francep4 is…next
Keyboard shortcuts
  • Space play / pause
  • ← / → step back / forward
  • N next token
  • G skip to end of generation
  • R reset
  • H toggle this help
Simplified explanatory model · deterministic illustrative values · not a live model trace · decoding shown greedy

prompt → tokens → representations → logits → probabilities → selected token → append → repeat

Prediction and selection are different operations. The model produces logits; softmax converts them into a distribution; a decoding strategy decides what to do with that distribution.

Decoder-only model · next-token choice

Decoding strategies

The same logits can produce different next tokens depending on the decoding rule. Compare greedy, temperature, top-k, and top-p on one fixed distribution.

Always select the token with the highest probability. Deterministic, but it can lock the model into repetitive or flat continuations.

Base logits · same for every strategy
Paris3.42
Lyon2.24
France1.69
world1.37
the-0.60
next-1.30
is-1.90
blue-2.70
After greedy
Paris61%likelyselected
Lyon19%
France11%
world8%
the1%
next1%
is0%
blue0%

Greedy always selects the most probable token, so no sampling is involved.

Sample draw 01
Deterministic illustrative values · sampling draws are seeded, not random · same base logits for every strategy

Greedy decoding selects the most probable token. Temperature, top-k, and top-p can reshape or restrict the available choices before sampling. The same model and prompt can therefore produce different continuations without changing a learned weight. Runtime context and decoding determine which learned patterns are activated and how the resulting probabilities become text.

06

5. What Pattern Prediction Means

Calling an LLM a pattern-prediction machine describes its operating contract; it does not imply that its behavior must be trivial. Learned patterns can include syntax, genre, factual associations, algorithms, explanations, tool-use conventions, and long sequences that resemble deliberate reasoning. All of them are expressed through conditional next-token probabilities.

The model does not retrieve one predetermined answer from its weights. It reconstructs a continuation from the prompt, learned parameters, temporary activations, and decoding rule. Small changes in any of those can change the generated path.

Inference also does not establish that a continuation is correct, meaningful, or worth adopting. As Understanding and Bottlenecks argues, generation produces a candidate artifact; people and institutions still have to ground, interpret, evaluate, and absorb it.

human language → tokens → training examples → prediction error → learned weights and embeddings → prompt-conditioned activations → next-token probabilities → decoded tokens → generated text

The machinery is remarkably capable, but its basic operation remains stable: learn patterns by predicting tokens, then use those learned patterns to predict again.

07

Sources