How to Trace the Thoughts of a Large Language Model
Tracing the thoughts of a large language model means inspecting the internal computations that produce its output — the activations, attention patterns, and learned features inside the network — rather than reading the text it generates. The output is the symptom; the trace is the mechanism.
The concrete problem: a model gives you a wrong answer, and the answer looks exactly like a right one. Same confident tone, same formatting, same fluent sentence. You cannot tell from the outside whether it retrieved a fact, pattern-matched a plausible-sounding string, or followed a chain of reasoning that broke at step three. So you want to open the hood. This article walks through what that actually involves for an open-weight model, which tools do the work, and — importantly — where the whole exercise stops being reliable.
Why the output alone can't tell you what happened
A language model generates one token at a time. At each step it produces a probability distribution over its entire vocabulary — a set of scores called logits, converted into probabilities. The token with the highest probability usually gets picked.
That distribution is the closest thing to a "thought" you can observe at the output layer. If a model answers a question about a date and the correct year has a probability of 0.31 while an incorrect year sits at 0.29, the model was genuinely uncertain — it just happened to sample the wrong one. If the correct year sits at 0.001, the model never had the right answer available at that position at all. Those are completely different failures, and the generated sentence looks identical in both cases.
So the first practical move is not exotic. Log the top-k token probabilities at each generation step. That single habit separates "the model knew and picked wrong" from "the model never knew."
A worked example: tracing a wrong year in GPT-2 Small
Take an open-weight model small enough to run on a laptop. GPT-2 Small is the standard starting point — 12 layers, roughly 124 million parameters, weights publicly available, and small enough that you can hold the whole activation history in memory.
Suppose you prompt it with a factual completion where the correct answer is a specific year and the model produces the wrong one. Here is the trace, step by step:
- Capture the residual stream. Using TransformerLens, you can cache the residual stream — the running vector that each layer reads from and writes to — at every layer for every token position. That gives you a layer-by-layer record of how the representation evolved.
- Read the logit lens at each layer. Apply the model's final unembedding matrix to the residual stream at intermediate layers. This is the "logit lens" trick: it shows what the model would have predicted if generation had stopped at layer 6, layer 8, layer 10. Often the correct year appears in the top predictions around the middle layers and then gets overwritten by the wrong one near the end. That is a specific, inspectable event.
- Attribute the shift. Once you know which layer flipped the prediction, you can look at which attention heads and MLP sublayers wrote into the residual stream at that layer. TransformerLens exposes per-head and per-MLP activation values, so you can zero out one component at a time and see whether the correct answer returns to the top of the distribution.
That last step is the whole point. You are not guessing why the model got it wrong. You are locating the component that changed the answer and confirming it by ablation — removing it and watching the prediction change back.
What "features" actually means inside the network
Individual neurons in a language model are mostly polysemantic — one neuron fires for several unrelated concepts, which makes reading them like reading a sentence in a language where every word means five things. Sparse autoencoders are the current workaround: they learn a much larger set of directions in activation space that are individually more interpretable.
Neuronpedia hosts browsable feature sets for several open models, so you can look up what a given feature responds to without training your own autoencoder. This is where interpretability gets genuinely useful for debugging: instead of "layer 9 head 4 matters," you get "there's a feature that fires on mentions of a specific country, and it's active at the position where the wrong year was produced."
The honest caveat: features found this way are correlational. A feature being active when the model produces a wrong answer does not mean the feature caused the wrong answer. You still need the ablation step to make a causal claim.
The tools, and what they don't do
Three names cover most of what a practitioner actually needs:
- TransformerLens — a Python library that reimplements popular open models in a consistent interface so you can hook, cache, and ablate any activation. It supports GPT-2, several Pythia sizes, and other open-weight families. This is the workhorse.
- Neuronpedia — a hosted browser for sparse-autoencoder features across multiple models. Good for lookup, not for custom analysis.
- Captum — PyTorch's attribution library, useful for input-attribution methods that tell you which input tokens mattered most to a given output.
None of these work on closed API models. If you are calling a hosted model through an API, you get tokens and logprobs at best — no activations, no attention, no residual stream. That is the single biggest constraint on this entire practice, and it is not a tooling gap you can engineer around. Interpretability requires weight access.
There is also a scale problem. Techniques that work cleanly on GPT-2 Small get expensive fast. Caching residual streams across every layer for every token grows with sequence length and model depth, and sparse autoencoders have to be trained per model per layer. A 12-layer model is a weekend project. A frontier-scale model is a research program.
Where tracing breaks down
Three limitations worth stating plainly, because the field's own papers are candid about them.
Faithfulness. A feature attribution can be a plausible story that does not reflect the actual computation. The only defense is causal testing — ablation, activation patching, intervention — and even that has edge cases where removing a component causes the network to route around it.
Composition. Circuits in language models are not modular. A component identified as "the date head" in one prompt may do something unrelated in another. Findings tend not to generalize as cleanly as the writeups suggest.
Coverage. You trace the specific behavior you prompted for. You do not get a general audit of the model. Tracing one wrong answer tells you about that answer.
For anyone weighing whether this is worth the effort: the answer depends on whether you need to fix a specific failure or understand the model broadly. For the first, tracing works well. For the second, it is still early.
One practical note on tool selection, since the landscape moves quickly: this site maintains an internal database of 360 AI tools with pricing and capability snapshots recorded at verification time, most recently updated 2026-09-18. It covers tooling categories and pricing bands — it is not a source of interpretability methods or research findings, and it should not be treated as one.
Key Takeaways
- Tracing model thoughts means reading internal activations and logits, not the generated text.
- Log top-k token probabilities first — it separates "knew but picked wrong" from "never knew."
- TransformerLens caches and ablates activations in open-weight models; Neuronpedia hosts browsable features.
- Interpretability requires weight access. Closed API models give you logprobs at best.
- Feature findings are correlational until you confirm them with ablation.
The short version
If you want to start today: pick GPT-2 Small, install TransformerLens, and cache the residual stream for one prompt where the model gets a fact wrong. Run the logit lens at each layer and find the layer where the correct answer drops out of the top predictions. Then ablate the attention heads at that layer one at a time until the correct answer comes back. That loop — observe, hypothesize, ablate, confirm — is the whole method. Everything else is scale.
What you will learn quickly is that the model's "reasoning" is not a chain of propositions you can read off. It is a sequence of vector transformations, and your job is to find the transformation that mattered. That is harder than it sounds and more tractable than it looks.
Sources
- AI Tool Database, Internal tool and pricing snapshot, 2026. Internal record of 360 AI tools with pricing and capability data verified as of 2026-09-18.
- TransformerLens documentation, TransformerLens: A Library for Mechanistic Interpretability. Python library for hooking, caching, and ablating activations in open-weight language models.
- Neuronpedia, Neuronpedia Feature Browser. Hosted browser for sparse-autoencoder features across multiple open models.
Frequently Asked Questions
Can I trace the thoughts of a model I only access through an API?
Not meaningfully. API access typically returns generated tokens and, at most, log probabilities for those tokens. You cannot see activations, attention patterns, or the residual stream, which is where the actual computation lives. Interpretability work requires direct weight access, which in practice means open-weight models you can run or host yourself. Closed models are a black box by design, and no tooling changes that.
What is the logit lens and why does it matter?
The logit lens applies a model's final unembedding matrix to intermediate-layer representations, showing what the model would have predicted if generation stopped early. It matters because it turns a single output into a layer-by-layer record of how the prediction evolved. You can see the correct answer appear in mid-layers and get overwritten later, which tells you exactly which layers to investigate further.
Do sparse autoencoders actually make model internals readable?
They help, but they are not a clean readout. Individual neurons are polysemantic, firing for several unrelated concepts, so autoencoders learn a larger set of more separable directions. The catch is that these features are correlational. A feature being active during a wrong answer does not prove it caused the wrong answer. You still need ablation or activation patching to establish causation, and findings often do not generalize across prompts.