Explainer / 2026-10-04
From Prompt to Token: How an LLM Generates a Response
A visual walkthrough from raw text to the next generated token, and why inference behaves differently during prefill and decoding.
One minute
A causal language model does not retrieve a finished paragraph. It scores possible next A chunk of text the model treats as one unit. It may be a whole word, part of a word, punctuation, or, in many tokenizers, a word plus the space before it. and then commits to one. That token is appended, and the model scores the next step using what it already computed plus the new token.
The prompt is read first, largely in parallel. That pass is called The forward pass over the prompt. Many prompt tokens can be processed together. This step produces the first next-token distribution and fills the KV cache. Its duration is the main part of time to first token.. It fills a cache of intermediate state. After that, generation is a loop: one new A chunk of text the model treats as one unit. It may be a whole word, part of a word, punctuation, or, in many tokenizers, a word plus the space before it. at a time. The original prompt is not naively run from scratch on every step.
Sampling controls such as A divisor applied to logits before softmax. Lower temperature makes the distribution sharper. Higher temperature makes it flatter. It is not a measure of hardware heat. sit outside the neural network. They change how the next A chunk of text the model treats as one unit. It may be a whole word, part of a word, punctuation, or, in many tokenizers, a word plus the space before it. is chosen from the scores. They do not rewrite the model's weights.
- Text becomes A chunk of text the model treats as one unit. It may be a whole word, part of a word, punctuation, or, in many tokenizers, a word plus the space before it., then vectors.
- Stacked layers mix context (A mechanism that rebuilds each token's representation as a weighted mix of other tokens' representations. It routes information. It is not, by itself, understanding.) and then process each position (The part of a transformer layer that processes each token on its own, after attention has mixed in context. It is where a lot of the model's stored patterns are applied.).
- The last position's final vector is scored against the whole The fixed set of tokens the model can emit. The output projection scores every entry. A typical model vocabulary is tens or hundreds of thousands of tokens, not a dictionary of English words..
- Decoding rules turn those scores into a choice.
- The chosen A chunk of text the model treats as one unit. It may be a whole word, part of a word, punctuation, or, in many tokenizers, a word plus the space before it. is appended, and the cache makes the next step incremental.
What matters
Generation is a loop that adds one token and reuses cached state. It is not a fresh reading of the whole prompt every time, and sampling knobs are not layers of the network.
The pipeline
The same prompt is used at every stage. Play the loop, or move one stage at a time. The content is on the page either way.
Ingest
Input prompt
The model receives raw text. Nothing numerical has happened yet.
Explain more
The string is only characters. Casing, spaces, and punctuation still matter because the tokenizer will not treat them as decorations. This walkthrough uses one sentence the whole way: "The cat sat on the".
The cat sat on theThe cat sat on the
Ingest
Tokenization
A tokenizer cuts the text into the units the model actually knows.
Explain more
The split below is a teaching simplification: The | cat | sat | on | the. Real tokenizers often glue the leading space to the next piece, so "The" and " the" are different A chunk of text the model treats as one unit. It may be a whole word, part of a word, punctuation, or, in many tokenizers, a word plus the space before it.. Subword tokenizers also cut rare words into pieces. The The fixed set of tokens the model can emit. The output projection scores every entry. A typical model vocabulary is tens or hundreds of thousands of tokens, not a dictionary of English words. is a fixed list of these pieces, not an English dictionary.
ThecatsatontheThecatsatontheIngest
Token embeddings and position
Each token id becomes a vector, and the model is told the order of the A chunk of text the model treats as one unit. It may be a whole word, part of a word, punctuation, or, in many tokenizers, a word plus the space before it..
Explain more
An embedding is a learned lookup table from id to vector. A way to tell the model the order of tokens. Some architectures add position vectors. Many current models rotate parts of the representation instead (rotary embeddings). Order is not free. keeps "cat sat" from meaning the same thing as "sat cat". Some models add position vectors. Many current ones apply rotary position A list of numbers that stands in for a token. Similar tokens start with related vectors. Later layers rewrite those vectors using context. inside A mechanism that rebuilds each token's representation as a weighted mix of other tokens' representations. It routes information. It is not, by itself, understanding. instead. Either way, order is represented on purpose. The tiny vectors in the figure are illustrative, not a dump from a trained model.
The184cat9241sat3332on508the627The184cat9241sat3332on508the627The -> [0.12, -0.40, 0.08, ...]
cat -> [0.55, 0.02, -0.31, ...]
Illustrative ids and vectors. The second "the" has a different id on purpose.
Transform
Transformer layers
Each layer lets A chunk of text the model treats as one unit. It may be a whole word, part of a word, punctuation, or, in many tokenizers, a word plus the space before it. exchange context, then processes each token on its own.
Explain more
A layer typically runs self-attention, then a The part of a transformer layer that processes each token on its own, after attention has mixed in context. It is where a lot of the model's stored patterns are applied., with A path that adds a layer's input back onto its output. The original signal stays available while attention and the feed-forward network write updates. and normalization around them. A mechanism that rebuilds each token's representation as a weighted mix of other tokens' representations. It routes information. It is not, by itself, understanding. builds each position's update from other positions. The feed-forward network applies the same learned transform at every position, separately. Stacking layers is what turns a static A list of numbers that stands in for a token. Similar tokens start with related vectors. Later layers rewrite those vectors using context. into a contextual The vector representing a token at a given layer. The final hidden state at the last position is what a standard causal model uses to score the next token.. Attention is a routing mechanism. It is not a transcript of the model's beliefs.
ThecatsatontheSelf-attentionEach position gathers context from the others, under a causal mask.
Feed-forward networkEach position is then transformed on its own.
Residual connections and normalization wrap both. The KV cache stores keys and values from this pass.
Transform
Output projection
A matrix maps that one vector onto a score for every A chunk of text the model treats as one unit. It may be a whole word, part of a word, punctuation, or, in many tokenizers, a word plus the space before it. in the The fixed set of tokens the model can emit. The output projection scores every entry. A typical model vocabulary is tens or hundreds of thousands of tokens, not a dictionary of English words..
Explain more
This matrix is the language-model head. In some models it is tied to the A list of numbers that stands in for a token. Similar tokens start with related vectors. Later layers rewrite those vectors using context. table, so the same vectors used to read A chunk of text the model treats as one unit. It may be a whole word, part of a word, punctuation, or, in many tokenizers, a word plus the space before it. are used to score them. The result is one The raw score for one vocabulary entry before it becomes a probability. Higher means more favored. Logits are not probabilities and do not need to sum to one. per The fixed set of tokens the model can emit. The output projection scores every entry. A typical model vocabulary is tens or hundreds of thousands of tokens, not a dictionary of English words. entry. The vocabulary is large. The figure shows a handful of candidates so the step is visible.
Thecatsatonthefinal vector x vocabulary -> one logit per token
Decode
Logits across the vocabulary
Logits are raw scores. They are not probabilities yet.
Explain more
A higher logit means the model currently favors that A chunk of text the model treats as one unit. It may be a whole word, part of a word, punctuation, or, in many tokenizers, a word plus the space before it. more. The raw score for one vocabulary entry before it becomes a probability. Higher means more favored. Logits are not probabilities and do not need to sum to one. can be negative. They do not sum to one. Any chart that calls them probabilities at this stage is mislabeled. The scores in this figure were chosen for teaching.
Thecatsatonthemat4.2 floor3.1 rug2.4 dog1.1 table0.6 Illustrative logits: mat 4.2, floor 3.1, rug 2.4, dog 1.1, table 0.6.
Decode
Decoding controls
Temperature, A filter that keeps only the k highest-scoring candidates and drops the rest before sampling., Nucleus sampling. Keep the smallest set of highest-probability tokens whose probabilities add up to at least p, then sample inside that set., and sometimes repetition penalties reshape the choice. They are not layers.
Explain more
Temperature divides the The raw score for one vocabulary entry before it becomes a probability. Higher means more favored. Logits are not probabilities and do not need to sum to one. before A function that turns a list of scores into positive numbers that sum to 1, so they can be treated as a probability distribution.. A filter that keeps only the k highest-scoring candidates and drops the rest before sampling. drops all but the k best scores. Nucleus sampling. Keep the smallest set of highest-probability tokens whose probabilities add up to at least p, then sample inside that set. keeps the smallest high-probability set that covers probability mass p. A repetition penalty, where a server uses one, down-weights A chunk of text the model treats as one unit. It may be a whole word, part of a word, punctuation, or, in many tokenizers, a word plus the space before it. that already appeared. Greedy decoding skips sampling and takes the argmax. None of these steps change the trained A learned weight stored in the model. Inference reads parameters. It does not update them..
Thecatsatonthetemperature / top-k / top-p / repetition penalty, where used
Decode
Probability distribution
Softmax turns the surviving scores into positive weights that sum to one.
Explain more
After A divisor applied to logits before softmax. Lower temperature makes the distribution sharper. Higher temperature makes it flatter. It is not a measure of hardware heat. and filters, the remaining The raw score for one vocabulary entry before it becomes a probability. Higher means more favored. Logits are not probabilities and do not need to sum to one. are normalized. That distribution is what sampling draws from. If every candidate but one was filtered, the distribution collapses and sampling becomes a formality. The percentages here are the A function that turns a list of scores into positive numbers that sum to 1, so they can be treated as a probability distribution. of the illustrative logits, shown so the arithmetic is visible.
Thecatsatonthemat63.7% floor21.2% rug10.5% dog2.9% table1.7% Output
Sample or select the next token
The system either takes the top A chunk of text the model treats as one unit. It may be a whole word, part of a word, punctuation, or, in many tokenizers, a word plus the space before it. or draws one from the distribution.
Explain more
In this teaching path the selected A chunk of text the model treats as one unit. It may be a whole word, part of a word, punctuation, or, in many tokenizers, a word plus the space before it. is mat. A real draw can choose a lower-ranked token when sampling is on. That is not an error in the network. It is the decoding policy doing what it was set to do. Greedy selection would have taken the highest bar.
matmatOutput
Append the token
The new A chunk of text the model treats as one unit. It may be a whole word, part of a word, punctuation, or, in many tokenizers, a word plus the space before it. joins the sequence. The text is now "The cat sat on the mat".
Explain more
Appending is the The model generates one token, appends it, and then predicts the next one conditioned on everything so far. The loop is the generation process. step made visible. The model does not revise earlier A chunk of text the model treats as one unit. It may be a whole word, part of a word, punctuation, or, in many tokenizers, a word plus the space before it. in the standard loop. It conditions the future on them. Detokenization, which turns token ids back into text for the person reading, can be slightly delayed in real systems so partial subwords are not shown as broken characters. The idea of the loop is unchanged.
ThecatsatonthematThecatsatonthematOutput
Decode again from cached state
The next step processes the new A chunk of text the model treats as one unit. It may be a whole word, part of a word, punctuation, or, in many tokenizers, a word plus the space before it.. Prior keys and values stay in the Stored key and value vectors from tokens already processed. Later generation steps reuse them instead of recomputing the whole prompt from scratch..
Explain more
This is the correction to a common picture. Generation does not resend the whole prompt through every layer from zero. The cache already holds the key and value vectors for "The cat sat on the". The new step attends against that cache and then appends its own keys and values. The cache grows by one position per generated A chunk of text the model treats as one unit. It may be a whole word, part of a word, punctuation, or, in many tokenizers, a word plus the space before it., per layer. That growth is why long generations cost more memory even though each step adds only one token.
ThecatsatonthematThecatsatontheCache holds the prefix. The new step brings one token.
Autoregressive loop
Append the token. Score again. The sequence is now "The cat sat on the mat", and the next step conditions on that longer prefix.
KV cache, beside the loop
Keys and values for tokens already seen stay put. The new token attends against them. This is the side channel. It is why the prompt is not naively recomputed from scratch.
Attention, as a teaching sketch
Sentence: "The fisherman sat on the bank and watched the water." Select a word. Stronger ties are marked in a heavier style and named in the text. This is not a literal view of every attention head in a production model, and attention is not the same thing as understanding.
Select a word to see the conceptual ties used in this sketch.
Relationship table
| Token | Related token | Strength in the sketch |
|---|---|---|
| fisherman | sat | moderate |
| fisherman | bank | strong |
| fisherman | water | moderate |
| fisherman | watched | moderate |
| sat | fisherman | strong |
| sat | on | moderate |
| bank | fisherman | strong |
| bank | water | strong |
| watched | water | strong |
| watched | fisherman | moderate |
| watched | bank | moderate |
| water | watched | strong |
| water | bank | strong |
| water | fisherman | moderate |
What the diagram leaves in and out
A standard causal layer is A mechanism that rebuilds each token's representation as a weighted mix of other tokens' representations. It routes information. It is not, by itself, understanding., then a The part of a transformer layer that processes each token on its own, after attention has mixed in context. It is where a lot of the model's stored patterns are applied., with A path that adds a layer's input back onto its output. The original signal stays available while attention and the feed-forward network write updates. and normalization. Drawings that show only attention are incomplete. The feed-forward network is where a large share of A learned weight stored in the model. Inference reads parameters. It does not update them. apply a token-wise transform after context has been mixed in.
The causal mask stops a position from attending to future A chunk of text the model treats as one unit. It may be a whole word, part of a word, punctuation, or, in many tokenizers, a word plus the space before it.. That is what makes the model a left-to-right predictor rather than a reader of the whole sentence at once.
Positional information is not one technique. Earlier models added position vectors. Many current models rotate parts of the query and key instead. The page only needs the consequence: order is represented, and it is represented inside the A mechanism that rebuilds each token's representation as a weighted mix of other tokens' representations. It routes information. It is not, by itself, understanding. computation for those models.
The diagram does not cover mixture-of-experts routing, speculative decoding, tool calls, or multimodal encoders. Those change the system around this loop. They do not remove it.
Sampling playground
Illustrative probability distribution for teaching purposes. This browser demonstration is not running a full language model.
To brew a good espresso, you must control the..
Logit table
| Candidate | Illustrative logit |
|---|---|
| pressure | 2.4 |
| temperature | 2.1 |
| grind | 1.8 |
| timing | 1.2 |
| beans | 0.7 |
| gravity | -0.4 |
probs = softmax(logits / temperature) keep the top k scores keep the smallest high-probability set that reaches p draw one token from what remains
Order used here: temperature, then top-k, then top-p, then renormalize, then sample. Greedy decoding would skip the draw and take the largest score.
Two very different workloads
Prefill and decode are often limited by different things. The split depends on the model, the batch, the length, and the hardware. The language below is "commonly" and "often" on purpose.
Many prompt tokens processed in parallel
- High parallelism across the prompt.
- Large matrix operations. Commonly compute-intensive.
- Fills the KV cache and produces the distribution for the first new token.
- Dominates time to first token.
Text alternative: four prompt tokens move through the model together during prefill.
One new token per generation step
- Sequential generation for a single sequence.
- Each step reads weights and a growing KV cache.
- Commonly memory-bandwidth constrained, not universally.
- Sets the latency between tokens.
Text alternative: one new token enters while the prefix remains available in the cache.
A note on the source
A separate prompt-to-response notebook was not in this repository when the page was written. The sequence follows the standard causal language-model path. Every number in the figures is illustrative and labeled as such. Where popular explanations say the entire prompt is recomputed from scratch for every new token, this page does not.
Sources
- Attention Is All You Need. Vaswani et al., 2017. The original encoder-decoder transformer. The explainer uses the later causal language-model form of the same layer ideas.
- Language Models are Unsupervised Multitask Learners. Radford et al., 2019. GPT-2: a causal transformer language model and byte-level byte-pair encoding.
- Efficiently Scaling Transformer Inference. Pope et al., 2022. A clear engineering treatment of prefill versus decode and why their bottlenecks differ.
- The Illustrated Transformer. Jay Alammar. A visual companion for the layer mechanics. Useful beside the papers, not a substitute for them.
Glossary
- Token
- A chunk of text the model treats as one unit. It may be a whole word, part of a word, punctuation, or, in many tokenizers, a word plus the space before it.
- Embedding
- A list of numbers that stands in for a token. Similar tokens start with related vectors. Later layers rewrite those vectors using context.
- Positional information
- A way to tell the model the order of tokens. Some architectures add position vectors. Many current models rotate parts of the representation instead (rotary embeddings). Order is not free.
- Hidden state
- The vector representing a token at a given layer. The final hidden state at the last position is what a standard causal model uses to score the next token.
- Attention
- A mechanism that rebuilds each token's representation as a weighted mix of other tokens' representations. It routes information. It is not, by itself, understanding.
- Residual connection
- A path that adds a layer's input back onto its output. The original signal stays available while attention and the feed-forward network write updates.
- Feed-forward network
- The part of a transformer layer that processes each token on its own, after attention has mixed in context. It is where a lot of the model's stored patterns are applied.
- Parameter
- A learned weight stored in the model. Inference reads parameters. It does not update them.
- Logit
- The raw score for one vocabulary entry before it becomes a probability. Higher means more favored. Logits are not probabilities and do not need to sum to one.
- Softmax
- A function that turns a list of scores into positive numbers that sum to 1, so they can be treated as a probability distribution.
- Temperature
- A divisor applied to logits before softmax. Lower temperature makes the distribution sharper. Higher temperature makes it flatter. It is not a measure of hardware heat.
- Top-k
- A filter that keeps only the k highest-scoring candidates and drops the rest before sampling.
- Top-p
- Nucleus sampling. Keep the smallest set of highest-probability tokens whose probabilities add up to at least p, then sample inside that set.
- KV cache
- Stored key and value vectors from tokens already processed. Later generation steps reuse them instead of recomputing the whole prompt from scratch.
- Prefill
- The forward pass over the prompt. Many prompt tokens can be processed together. This step produces the first next-token distribution and fills the KV cache. Its duration is the main part of time to first token.
- Decode
- Generating later tokens one step at a time for a sequence. Each step adds one token and reads the growing cache. Decode is often limited by memory movement rather than by raw arithmetic, commonly, not always.
- Memory bandwidth
- How much data the hardware can move each second between memory and the compute units. When decode spends its time loading weights and cache, more arithmetic units do not make it proportionally faster.
- FLOPS
- Floating-point operations per second, a measure of arithmetic throughput. Prefill is commonly rich in large matrix multiplications. FLOPS are not the same thing as user-visible speed.
- Autoregressive
- The model generates one token, appends it, and then predicts the next one conditioned on everything so far. The loop is the generation process.
- Vocabulary
- The fixed set of tokens the model can emit. The output projection scores every entry. A typical model vocabulary is tens or hundreds of thousands of tokens, not a dictionary of English words.