Explainer / 2026-10-04

From Prompt to Token: How an LLM Generates a Response

A visual walkthrough from raw text to the next generated token, and why inference behaves differently during prefill and decoding.

One minute

A causal language model does not retrieve a finished paragraph. It scores possible next A chunk of text the model treats as one unit. It may be a whole word, part of a word, punctuation, or, in many tokenizers, a word plus the space before it. and then commits to one. That token is appended, and the model scores the next step using what it already computed plus the new token.

The prompt is read first, largely in parallel. That pass is called The forward pass over the prompt. Many prompt tokens can be processed together. This step produces the first next-token distribution and fills the KV cache. Its duration is the main part of time to first token.. It fills a cache of intermediate state. After that, generation is a loop: one new A chunk of text the model treats as one unit. It may be a whole word, part of a word, punctuation, or, in many tokenizers, a word plus the space before it. at a time. The original prompt is not naively run from scratch on every step.

Sampling controls such as A divisor applied to logits before softmax. Lower temperature makes the distribution sharper. Higher temperature makes it flatter. It is not a measure of hardware heat. sit outside the neural network. They change how the next A chunk of text the model treats as one unit. It may be a whole word, part of a word, punctuation, or, in many tokenizers, a word plus the space before it. is chosen from the scores. They do not rewrite the model's weights.

  1. Text becomes A chunk of text the model treats as one unit. It may be a whole word, part of a word, punctuation, or, in many tokenizers, a word plus the space before it., then vectors.
  2. Stacked layers mix context (A mechanism that rebuilds each token's representation as a weighted mix of other tokens' representations. It routes information. It is not, by itself, understanding.) and then process each position (The part of a transformer layer that processes each token on its own, after attention has mixed in context. It is where a lot of the model's stored patterns are applied.).
  3. The last position's final vector is scored against the whole The fixed set of tokens the model can emit. The output projection scores every entry. A typical model vocabulary is tens or hundreds of thousands of tokens, not a dictionary of English words..
  4. Decoding rules turn those scores into a choice.
  5. The chosen A chunk of text the model treats as one unit. It may be a whole word, part of a word, punctuation, or, in many tokenizers, a word plus the space before it. is appended, and the cache makes the next step incremental.

What matters

Generation is a loop that adds one token and reuses cached state. It is not a fresh reading of the whole prompt every time, and sampling knobs are not layers of the network.

Walk through the pipeline

The pipeline

The same prompt is used at every stage. Play the loop, or move one stage at a time. The content is on the page either way.

Running example

The cat sat on the

Raw text, before any split.

  1. Ingest

    Input prompt

    The model receives raw text. Nothing numerical has happened yet.

    Explain more

    The string is only characters. Casing, spaces, and punctuation still matter because the tokenizer will not treat them as decorations. This walkthrough uses one sentence the whole way: "The cat sat on the".

    The cat sat on the

  2. Ingest

    Tokenization

    A tokenizer cuts the text into the units the model actually knows.

    Explain more

    The split below is a teaching simplification: The | cat | sat | on | the. Real tokenizers often glue the leading space to the next piece, so "The" and " the" are different A chunk of text the model treats as one unit. It may be a whole word, part of a word, punctuation, or, in many tokenizers, a word plus the space before it.. Subword tokenizers also cut rare words into pieces. The The fixed set of tokens the model can emit. The output projection scores every entry. A typical model vocabulary is tens or hundreds of thousands of tokens, not a dictionary of English words. is a fixed list of these pieces, not an English dictionary.

    Thecatsatonthe
  3. Ingest

    Token embeddings and position

    Each token id becomes a vector, and the model is told the order of the A chunk of text the model treats as one unit. It may be a whole word, part of a word, punctuation, or, in many tokenizers, a word plus the space before it..

    Explain more

    An embedding is a learned lookup table from id to vector. A way to tell the model the order of tokens. Some architectures add position vectors. Many current models rotate parts of the representation instead (rotary embeddings). Order is not free. keeps "cat sat" from meaning the same thing as "sat cat". Some models add position vectors. Many current ones apply rotary position A list of numbers that stands in for a token. Similar tokens start with related vectors. Later layers rewrite those vectors using context. inside A mechanism that rebuilds each token's representation as a weighted mix of other tokens' representations. It routes information. It is not, by itself, understanding. instead. Either way, order is represented on purpose. The tiny vectors in the figure are illustrative, not a dump from a trained model.

    The184cat9241sat3332on508the627

    The -> [0.12, -0.40, 0.08, ...]

    cat -> [0.55, 0.02, -0.31, ...]

    Illustrative ids and vectors. The second "the" has a different id on purpose.

  4. Transform

    Transformer layers

    Each layer lets A chunk of text the model treats as one unit. It may be a whole word, part of a word, punctuation, or, in many tokenizers, a word plus the space before it. exchange context, then processes each token on its own.

    Explain more

    A layer typically runs self-attention, then a The part of a transformer layer that processes each token on its own, after attention has mixed in context. It is where a lot of the model's stored patterns are applied., with A path that adds a layer's input back onto its output. The original signal stays available while attention and the feed-forward network write updates. and normalization around them. A mechanism that rebuilds each token's representation as a weighted mix of other tokens' representations. It routes information. It is not, by itself, understanding. builds each position's update from other positions. The feed-forward network applies the same learned transform at every position, separately. Stacking layers is what turns a static A list of numbers that stands in for a token. Similar tokens start with related vectors. Later layers rewrite those vectors using context. into a contextual The vector representing a token at a given layer. The final hidden state at the last position is what a standard causal model uses to score the next token.. Attention is a routing mechanism. It is not a transcript of the model's beliefs.

    Self-attention

    Each position gathers context from the others, under a causal mask.

    Feed-forward network

    Each position is then transformed on its own.

    Residual connections and normalization wrap both. The KV cache stores keys and values from this pass.

  5. Transform

    Final hidden state

    The vector at the last position, after the last layer, is what gets scored.

    Explain more

    A causal model predicts the A chunk of text the model treats as one unit. It may be a whole word, part of a word, punctuation, or, in many tokenizers, a word plus the space before it. that follows the prefix. It uses the final The vector representing a token at a given layer. The final hidden state at the last position is what a standard causal model uses to score the next token. at the final position, not an average of every token. Earlier positions still matter because A mechanism that rebuilds each token's representation as a weighted mix of other tokens' representations. It routes information. It is not, by itself, understanding. and residuals have already written their influence into that last vector. During The forward pass over the prompt. Many prompt tokens can be processed together. This step produces the first next-token distribution and fills the KV cache. Its duration is the main part of time to first token., that last position is the end of the prompt. During later steps, it is the token you just appended.

    Thecatsatonthe

    The last position is the one that gets scored.

  6. Transform

    Output projection

    A matrix maps that one vector onto a score for every A chunk of text the model treats as one unit. It may be a whole word, part of a word, punctuation, or, in many tokenizers, a word plus the space before it. in the The fixed set of tokens the model can emit. The output projection scores every entry. A typical model vocabulary is tens or hundreds of thousands of tokens, not a dictionary of English words..

    Explain more

    This matrix is the language-model head. In some models it is tied to the A list of numbers that stands in for a token. Similar tokens start with related vectors. Later layers rewrite those vectors using context. table, so the same vectors used to read A chunk of text the model treats as one unit. It may be a whole word, part of a word, punctuation, or, in many tokenizers, a word plus the space before it. are used to score them. The result is one The raw score for one vocabulary entry before it becomes a probability. Higher means more favored. Logits are not probabilities and do not need to sum to one. per The fixed set of tokens the model can emit. The output projection scores every entry. A typical model vocabulary is tens or hundreds of thousands of tokens, not a dictionary of English words. entry. The vocabulary is large. The figure shows a handful of candidates so the step is visible.

    final vector x vocabulary -> one logit per token

  7. Decode

    Logits across the vocabulary

    Logits are raw scores. They are not probabilities yet.

    Explain more

    A higher logit means the model currently favors that A chunk of text the model treats as one unit. It may be a whole word, part of a word, punctuation, or, in many tokenizers, a word plus the space before it. more. The raw score for one vocabulary entry before it becomes a probability. Higher means more favored. Logits are not probabilities and do not need to sum to one. can be negative. They do not sum to one. Any chart that calls them probabilities at this stage is mislabeled. The scores in this figure were chosen for teaching.

    mat4.2
    floor3.1
    rug2.4
    dog1.1
    table0.6

    Illustrative logits: mat 4.2, floor 3.1, rug 2.4, dog 1.1, table 0.6.

  8. Decode

    Decoding controls

    Temperature, A filter that keeps only the k highest-scoring candidates and drops the rest before sampling., Nucleus sampling. Keep the smallest set of highest-probability tokens whose probabilities add up to at least p, then sample inside that set., and sometimes repetition penalties reshape the choice. They are not layers.

    Explain more

    Temperature divides the The raw score for one vocabulary entry before it becomes a probability. Higher means more favored. Logits are not probabilities and do not need to sum to one. before A function that turns a list of scores into positive numbers that sum to 1, so they can be treated as a probability distribution.. A filter that keeps only the k highest-scoring candidates and drops the rest before sampling. drops all but the k best scores. Nucleus sampling. Keep the smallest set of highest-probability tokens whose probabilities add up to at least p, then sample inside that set. keeps the smallest high-probability set that covers probability mass p. A repetition penalty, where a server uses one, down-weights A chunk of text the model treats as one unit. It may be a whole word, part of a word, punctuation, or, in many tokenizers, a word plus the space before it. that already appeared. Greedy decoding skips sampling and takes the argmax. None of these steps change the trained A learned weight stored in the model. Inference reads parameters. It does not update them..

    temperature / top-k / top-p / repetition penalty, where used

  9. Decode

    Probability distribution

    Softmax turns the surviving scores into positive weights that sum to one.

    Explain more

    After A divisor applied to logits before softmax. Lower temperature makes the distribution sharper. Higher temperature makes it flatter. It is not a measure of hardware heat. and filters, the remaining The raw score for one vocabulary entry before it becomes a probability. Higher means more favored. Logits are not probabilities and do not need to sum to one. are normalized. That distribution is what sampling draws from. If every candidate but one was filtered, the distribution collapses and sampling becomes a formality. The percentages here are the A function that turns a list of scores into positive numbers that sum to 1, so they can be treated as a probability distribution. of the illustrative logits, shown so the arithmetic is visible.

    mat63.7%
    floor21.2%
    rug10.5%
    dog2.9%
    table1.7%
  10. Output

    Sample or select the next token

    The system either takes the top A chunk of text the model treats as one unit. It may be a whole word, part of a word, punctuation, or, in many tokenizers, a word plus the space before it. or draws one from the distribution.

    Explain more

    In this teaching path the selected A chunk of text the model treats as one unit. It may be a whole word, part of a word, punctuation, or, in many tokenizers, a word plus the space before it. is mat. A real draw can choose a lower-ranked token when sampling is on. That is not an error in the network. It is the decoding policy doing what it was set to do. Greedy selection would have taken the highest bar.

    mat
  11. Output

    Append the token

    The new A chunk of text the model treats as one unit. It may be a whole word, part of a word, punctuation, or, in many tokenizers, a word plus the space before it. joins the sequence. The text is now "The cat sat on the mat".

    Explain more

    Appending is the The model generates one token, appends it, and then predicts the next one conditioned on everything so far. The loop is the generation process. step made visible. The model does not revise earlier A chunk of text the model treats as one unit. It may be a whole word, part of a word, punctuation, or, in many tokenizers, a word plus the space before it. in the standard loop. It conditions the future on them. Detokenization, which turns token ids back into text for the person reading, can be slightly delayed in real systems so partial subwords are not shown as broken characters. The idea of the loop is unchanged.

    Thecatsatonthemat
  12. Output

    Decode again from cached state

    The next step processes the new A chunk of text the model treats as one unit. It may be a whole word, part of a word, punctuation, or, in many tokenizers, a word plus the space before it.. Prior keys and values stay in the Stored key and value vectors from tokens already processed. Later generation steps reuse them instead of recomputing the whole prompt from scratch..

    Explain more

    This is the correction to a common picture. Generation does not resend the whole prompt through every layer from zero. The cache already holds the key and value vectors for "The cat sat on the". The new step attends against that cache and then appends its own keys and values. The cache grows by one position per generated A chunk of text the model treats as one unit. It may be a whole word, part of a word, punctuation, or, in many tokenizers, a word plus the space before it., per layer. That growth is why long generations cost more memory even though each step adds only one token.

    Thecatsatonthe

    Cache holds the prefix. The new step brings one token.

Autoregressive loop

Append the token. Score again. The sequence is now "The cat sat on the mat", and the next step conditions on that longer prefix.

KV cache, beside the loop

Keys and values for tokens already seen stay put. The new token attends against them. This is the side channel. It is why the prompt is not naively recomputed from scratch.

Attention, as a teaching sketch

Sentence: "The fisherman sat on the bank and watched the water." Select a word. Stronger ties are marked in a heavier style and named in the text. This is not a literal view of every attention head in a production model, and attention is not the same thing as understanding.

Select a word to see the conceptual ties used in this sketch.

Relationship table
Conceptual relationships in the teaching sketch. Not measured attention weights.
TokenRelated tokenStrength in the sketch
fishermansatmoderate
fishermanbankstrong
fishermanwatermoderate
fishermanwatchedmoderate
satfishermanstrong
satonmoderate
bankfishermanstrong
bankwaterstrong
watchedwaterstrong
watchedfishermanmoderate
watchedbankmoderate
waterwatchedstrong
waterbankstrong
waterfishermanmoderate

What the diagram leaves in and out

A standard causal layer is A mechanism that rebuilds each token's representation as a weighted mix of other tokens' representations. It routes information. It is not, by itself, understanding., then a The part of a transformer layer that processes each token on its own, after attention has mixed in context. It is where a lot of the model's stored patterns are applied., with A path that adds a layer's input back onto its output. The original signal stays available while attention and the feed-forward network write updates. and normalization. Drawings that show only attention are incomplete. The feed-forward network is where a large share of A learned weight stored in the model. Inference reads parameters. It does not update them. apply a token-wise transform after context has been mixed in.

The causal mask stops a position from attending to future A chunk of text the model treats as one unit. It may be a whole word, part of a word, punctuation, or, in many tokenizers, a word plus the space before it.. That is what makes the model a left-to-right predictor rather than a reader of the whole sentence at once.

Positional information is not one technique. Earlier models added position vectors. Many current models rotate parts of the query and key instead. The page only needs the consequence: order is represented, and it is represented inside the A mechanism that rebuilds each token's representation as a weighted mix of other tokens' representations. It routes information. It is not, by itself, understanding. computation for those models.

The diagram does not cover mixture-of-experts routing, speculative decoding, tool calls, or multimodal encoders. Those change the system around this loop. They do not remove it.

Sampling playground

Illustrative probability distribution for teaching purposes. This browser demonstration is not running a full language model.

To brew a good espresso, you must control the..

1.0
6
1.00

Logit table
Illustrative logits chosen for this demo. They were not measured from a language model.
CandidateIllustrative logit
pressure2.4
temperature2.1
grind1.8
timing1.2
beans0.7
gravity-0.4
probs = softmax(logits / temperature)
keep the top k scores
keep the smallest high-probability set that reaches p
draw one token from what remains

Order used here: temperature, then top-k, then top-p, then renormalize, then sample. Greedy decoding would skip the draw and take the largest score.

Two very different workloads

Prefill and decode are often limited by different things. The split depends on the model, the batch, the length, and the hardware. The language below is "commonly" and "often" on purpose.

Many prompt tokens processed in parallel

  • High parallelism across the prompt.
  • Large matrix operations. Commonly compute-intensive.
  • Fills the KV cache and produces the distribution for the first new token.
  • Dominates time to first token.

Text alternative: four prompt tokens move through the model together during prefill.

One new token per generation step

Thecatsatonthe
  • Sequential generation for a single sequence.
  • Each step reads weights and a growing KV cache.
  • Commonly memory-bandwidth constrained, not universally.
  • Sets the latency between tokens.

Text alternative: one new token enters while the prefix remains available in the cache.

A note on the source

A separate prompt-to-response notebook was not in this repository when the page was written. The sequence follows the standard causal language-model path. Every number in the figures is illustrative and labeled as such. Where popular explanations say the entire prompt is recomputed from scratch for every new token, this page does not.

Sources

  1. Attention Is All You Need. Vaswani et al., 2017. The original encoder-decoder transformer. The explainer uses the later causal language-model form of the same layer ideas.
  2. Language Models are Unsupervised Multitask Learners. Radford et al., 2019. GPT-2: a causal transformer language model and byte-level byte-pair encoding.
  3. Efficiently Scaling Transformer Inference. Pope et al., 2022. A clear engineering treatment of prefill versus decode and why their bottlenecks differ.
  4. The Illustrated Transformer. Jay Alammar. A visual companion for the layer mechanics. Useful beside the papers, not a substitute for them.
Glossary
Token
A chunk of text the model treats as one unit. It may be a whole word, part of a word, punctuation, or, in many tokenizers, a word plus the space before it.
Embedding
A list of numbers that stands in for a token. Similar tokens start with related vectors. Later layers rewrite those vectors using context.
Positional information
A way to tell the model the order of tokens. Some architectures add position vectors. Many current models rotate parts of the representation instead (rotary embeddings). Order is not free.
Hidden state
The vector representing a token at a given layer. The final hidden state at the last position is what a standard causal model uses to score the next token.
Attention
A mechanism that rebuilds each token's representation as a weighted mix of other tokens' representations. It routes information. It is not, by itself, understanding.
Residual connection
A path that adds a layer's input back onto its output. The original signal stays available while attention and the feed-forward network write updates.
Feed-forward network
The part of a transformer layer that processes each token on its own, after attention has mixed in context. It is where a lot of the model's stored patterns are applied.
Parameter
A learned weight stored in the model. Inference reads parameters. It does not update them.
Logit
The raw score for one vocabulary entry before it becomes a probability. Higher means more favored. Logits are not probabilities and do not need to sum to one.
Softmax
A function that turns a list of scores into positive numbers that sum to 1, so they can be treated as a probability distribution.
Temperature
A divisor applied to logits before softmax. Lower temperature makes the distribution sharper. Higher temperature makes it flatter. It is not a measure of hardware heat.
Top-k
A filter that keeps only the k highest-scoring candidates and drops the rest before sampling.
Top-p
Nucleus sampling. Keep the smallest set of highest-probability tokens whose probabilities add up to at least p, then sample inside that set.
KV cache
Stored key and value vectors from tokens already processed. Later generation steps reuse them instead of recomputing the whole prompt from scratch.
Prefill
The forward pass over the prompt. Many prompt tokens can be processed together. This step produces the first next-token distribution and fills the KV cache. Its duration is the main part of time to first token.
Decode
Generating later tokens one step at a time for a sequence. Each step adds one token and reads the growing cache. Decode is often limited by memory movement rather than by raw arithmetic, commonly, not always.
Memory bandwidth
How much data the hardware can move each second between memory and the compute units. When decode spends its time loading weights and cache, more arithmetic units do not make it proportionally faster.
FLOPS
Floating-point operations per second, a measure of arithmetic throughput. Prefill is commonly rich in large matrix multiplications. FLOPS are not the same thing as user-visible speed.
Autoregressive
The model generates one token, appends it, and then predicts the next one conditioned on everything so far. The loop is the generation process.
Vocabulary
The fixed set of tokens the model can emit. The output projection scores every entry. A typical model vocabulary is tens or hundreds of thousands of tokens, not a dictionary of English words.