The laboratory

Learn by doing

Four short experiments, followed by readings to understand what the results actually support.

The model in motion

How does an LLM predict the next token?

Follow a sentence through tokens, vectors, attention, Transformer layers and next-token selection.

48 seconds · no audio · 1080p

Educational illustration: simplified word-level tokenization, fictional vectors and weights. Probabilities are calculated over four candidates, not a model’s full vocabulary. The film selects the maximum; models can also sample.

Read the film description

From text to numbers

The sentence “The capital of France is” splits into five words, treated here as tokens. Each token becomes a vector; twelve dimensions are illustrated in a schematic space.

From relations to context

Connections between positions become a matrix of attention weights. Each row attends to earlier positions and its own position. Future positions are masked.

Through the layers

The representation passes through the Transformer layers. Attention, the MLP and the residual stream contribute to its transformation. The journey through the layers is schematic.

From scores to the next token

The final position produces logits for four candidates: Paris 8.7; Lyon 3.1; Marseille 1.8; London 0.9. Softmax at temperature 1 converts them to probabilities totaling 100%. Paris receives about 99.49% in this reduced example. The most probable token joins the sentence: “The capital of France is Paris”. The process can repeat.

Download the MP4

Build an assistant you can rely on

Imagine an assistant helping a team handle customer requests. It must understand the problem, extract useful information and consult the right policies. At each stage, a plausible answer can hide an error. This learning path helps you identify those errors and build the missing checks.

  1. Understand the request

    Distinguish success on a dataset from the ability to recognize the problem that matters.

  2. Extract the right information

    Understand why usable formatting and correct information are separate requirements.

  3. Find the applicable policy

    Distinguish an apparently relevant document from sufficient evidence for a specific situation.

  4. Understand variation in answers

    Distinguish the probability of selecting a token from the probability that an answer is true.

  5. Evaluate the service provided

    Read precision from the perspective of received alerts and recall from the perspective of actual incidents.