Yes. But only some of what you see is the raw data. The rest is either derived from it by a standard technique, or drawn to make it readable. This page separates the three.
The one-paragraph version
Every number on screen comes from the published weights of the model you picked, running for real in your browser: the built-in GPT-2, DistilGPT2 or MDLM, or an open model loaded from Hugging Face. The point cloud is that model's own vocabulary, its first latent space: every token is a vector of hundreds or thousands of numbers (50,257 tokens of 768 numbers for GPT-2, 151,936 tokens of 1,536 for Qwen2.5-1.5B). You can't see that many dimensions, so the cloud is a 3-D map of them, and like any map it distorts. When you type a query the model genuinely runs; every intermediate result is recorded and replayed slowly. The playback is slowed down, not the model.
Where the numbers come from
The model runs once per token at full speed. The visualizer records everything it computed, then plays it back layer by layer at whatever speed you choose.
Checked against the standard reference implementations on the same weights. The built-in GPT-2 matches to 4 decimal places, and MDLM to about 0.00001. The Hugging Face models marked Verified in the model browser pick the same next word, share at least 9 of the top 10 guesses, and give probabilities within 0.0013 at worst. Models marked Should load use a supported design but have not been checked one by one.
Three kinds of thing on screen
Read straight out of the computation. No interpretation added.
Established techniques applied to the real numbers. Useful, but they add assumptions.
Drawings that make the numbers readable. They are not objects inside the model.
The most important caveat
The real space has hundreds to thousands of dimensions: 768 for GPT-2, 1,536 for Qwen2.5-1.5B. The cloud squeezes it into 3 so you can see it, in one of two layouts (switch with C). Clusters (UMAP) groups tokens with some of their neighbours, so points close together tend to be related, but most true neighbours end up elsewhere: for GPT-2, only about 7% of each token's 15 nearest neighbours in the real space are also among its 15 nearest in the cloud. Distances between clusters, directions and the overall shape mean very little; the layout is one of many equally valid arrangements, and UMAP started from a different random seed would place the clusters differently. PCA shows the three directions along which the vectors vary most. Those directions are real, but for GPT-2 they hold only 3.4% of the variation, so most tokens pile into one blob and fewer than 1% of true neighbours stay together. To see the real geometry, open the Embeddings view and look at a token: the cyan neighbour lines are measured in the full space, which is why they often jump across the cloud. That jump is the map's distortion made visible.
Piece by piece
The actual attention weights each head assigns to each earlier token, for the token you are looking at. One row per layer, one column per head: 12 × 12 for GPT-2 and MDLM, 6 × 12 for DistilGPT2, 28 × 12 for Qwen2.5-1.5B. In Llama, SmolLM and Qwen models, several heads share one set of keys and values (grouped-query attention), but each head still has its own attention pattern, and that is what is shown. For MDLM, attention runs both ways, so a slot can attend to words on its right.
How much each head adds to the model's running state (the residual stream) at this token: the length of what it writes. A head can attend strongly and still write very little.
The actual activation of every neuron in every layer: the number that feeds the layer's output projection. In GPT-2, DistilGPT2 and Pythia it is one projection passed through GELU. In Llama, SmolLM and Qwen it is a gated product (SwiGLU): one projection passed through SiLU, times another. Either way it is the model's own number. MDLM's replay stores attention and neurons at 8-bit precision to save memory, which is enough for colour.
In the Embeddings view: the token's actual coordinates (768 for GPT-2, 576 for SmolLM2-135M, 1,536 for Qwen2.5-1.5B), and the tokens closest to it in the full space by cosine similarity. For the built-in models the neighbours were worked out exactly in advance. For models built in your browser they are worked out when you look: a quick 8-bit pass shortlists 64 candidates, which are then rescored from the stored vectors, so they are exact unless a true neighbour missed the shortlist. The most faithful geometry on the page.
The model's working state partway through, pushed through its final normalisation and output layer to ask "what would you say if you stopped here?". A standard interpretability technique, not something the model does itself. Trustworthy near the top, rough near the bottom. For models with very large vocabularies it is worked out at a spread of depths rather than at every layer, to fit in memory (6 of 29 for Qwen2.5-1.5B).
Clusters (UMAP) or PCA of the vocabulary; switch with C. For the built-in models and popular Hugging Face models the cluster map is computed in advance. For other models your browser shows PCA first and computes the clusters in the background, with the same recipe. Local neighbourhoods are roughly right; global layout is arbitrary.
At each layer, the average map position of the model's top 10 guesses, weighted by how likely each is. A picture of what the model is leaning toward. It is not the model's internal state moving through space. That state is a vector as wide as the token vectors (768 numbers for GPT-2) at every layer; it is computed but not drawn as a point, because it doesn't live on the vocabulary map. The fly-through (T) moves the camera along this path, so you travel through the picture of what the model is leaning toward, not through its internal state. The guesses labelled along the way are the real logit-lens readouts for that layer.
Real weights, but averaged over all the heads of the current layer unless you pin one, and anything under 2% is hidden.
A masked slot has no real position; it is drawn at the average of its guesses until it locks onto a word.
Word, subword, number, punctuation: a rule based on how the token is spelled. The model knows nothing about these categories.
Which model you are looking at
Every model here is a transformer. Token vectors go in, pass through a stack of layers that each gather information from other tokens (attention) and transform it (the MLP), and come out as guesses for the next word. The view is the same for all of them; these are the differences worth knowing.
Position is a learned vector added to each token at the start. LayerNorm and GELU neurons. The same table reads tokens in and picks the output word.
Position is not added as a vector: it rotates each head's queries and keys (rotary embeddings), so attention depends on how far apart two tokens are. RMSNorm, gated SwiGLU neurons and grouped-query attention. Qwen3 also normalises each head's queries and keys. These small models share one table for input and output.
Rotary positions on part of each head, and attention and the MLP read the same input side by side instead of one after the other. It has a separate output table.
Not a next-word model. The whole answer starts as [MASK] and every slot is filled in over denoising steps, in any order, with attention running both ways. It has a separate output table.
Your text goes in exactly as typed, without the chat template these models were fine-tuned on. So they continue your text rather than reply like a chatbot, and what you see is the model's raw response to that text.
Where people get confused
Like a VR headset, the visualizer loads full detail only where you are looking. The model's attention is a separate, real thing, shown in the heads grid and the arcs. They share a metaphor; keep them apart.
The cloud is the token embedding space: a fixed table and the model's entry point. In GPT-2 and most models here the same table also picks the output word, which is why plotting predictions on it makes sense. MDLM and Pythia use separate input and output tables, so placing their guesses on the input map is an approximation. The hidden states inside the model are latent spaces too, arguably the more interesting ones. They are fully computed and shown indirectly, through their strength, the logit lens and attention, but not as a cloud of their own.
If someone asks "so is it real?"
The model and every number it produces are real. The cloud is an honest but distorted map of the model's vocabulary space. The glowing path is a picture of what it is leaning toward, not a literal trajectory.