Independent study
From prompt to token, taught in public
A visual walkthrough of how a causal language model turns text into the next token, built to be accurate enough for a technical reader and short enough for everyone else.
Context
I have been studying how transformer language models actually produce a response: tokenization, attention, the KV cache, and the difference between prefill and decode. The public artifact is the research note. A separate source notebook was not in this repository when the page was written; the page follows the standard causal-model inference path and labels every figure that is not a real model trace.
Problem
Teams are making decisions about tools whose internal loop they have only met as a product demo. A metaphor ("it predicts the next word") is true and insufficient. A research paper is sufficient and, for many readers, unusable on a first pass.
My role
I researched the mechanism, checked the common failure points in popular explanations, and designed the page. I did not train a model for this piece, and the sampling playground is not running one.
System
A prompt, a tokenizer, a stack of transformer layers, a vocabulary projection, decoding rules, and the cache that makes the next step cheap enough to repeat. Around that loop: a reader who may be an executive, an analyst, a student, or an engineer. The page has to serve all four without pretending they need the same depth at the same time.
Approach
Put the safest conceptual sequence on one path. Separate the neural forward pass from the sampling controls so temperature does not look like a layer of the network. Show prefill and decode as different workloads. Say "commonly" where a bottleneck depends on model, batch, and hardware. Keep a glossary on the page.
Output
The explainer at /research/llm-prompt-to-response/, with an attention sketch, a sampling playground, and a prefill/decode comparison. Pseudocode appears only where it teaches the sampling step.
Outcome
The outcome I will claim is the artifact and the corrections it makes to sloppy explanations. I do not claim a measured learning gain.
What this demonstrates
The ability to learn a technical system from the mechanism and then teach it without theater, including the discipline to mark illustrative numbers as illustrative.