Expertise

Artificial intelligence

Transformer inference, studied from the mechanism and taught at more than one depth.

Transformer fundamentals

Self-attention, feed-forward networks, residual connections, and normalization, stacked so each layer rewrites token representations. The public walkthrough stays with the causal language-model case rather than attempting every variant.

Inference

Inference is the forward pass plus the decoding loop. Prefill and decode are different workloads. The KV cache is why the prompt is not naively recomputed for every new token.

LLM systems

A deployed system is the model plus a tokenizer, decoding policy, cache, and serving constraints. Product behavior that looks like "the model" is often the decoding policy or the prompt wrapper.

Model behavior

Sampling temperature, top-k, and top-p change the distribution you draw from. They do not change the trained weights. I keep that distinction visible because it changes how people debug a strange answer.

Model architecture

I focus on the standard dense causal transformer well enough to teach it. Mixture-of-experts routing, speculative decoding, and multimodal encoders are acknowledged on the deep dive as things the main diagram does not cover.

Technical research

The prompt-to-token note is the research artifact. It is an explainer with a technical-accuracy pass, not a paper claiming a new result.

Open-source experimentation

Experimentation belongs in the lab, labeled with its real maturity. I will not present a notebook sketch as a product, and I will not present illustrative probabilities as measured model output.