Beneath the Surface of Chains-of-Thought: A Mechanistic Interpretation of Reasoning Operations in LLMsExplained for Beginners
Seogyeong Jeong, Jaehui Hwang, Dongyoon Han +3 more
Abstract
Reasoning in large language models unfolds through diverse functional operations, such as problem formulation, goal decomposition, and deduction. Although these operations are explicitly distinguished in text, little is known about how they are geometrically organized in representation spaces. To this end, we investigate whether distinct reasoning operations exhibit corresponding geometric structure in hidden representations. We find that operations are separable in held-out representations, with separability peaking in middle layers, and verify that this structure is not explained by lexical or positional confounds. Across layers, token-wise operation-alignment becomes more distributed over spans, while identical surface tokens are represented differently depending on the operation of its surrounding chunk. Attention-masking interventions further show that operation-aligned representations at chunk onset depend on preceding reasoning context. Consequently, our work demonstrates that language models maintain representational correspondence between linguistic reasoning expressions and their internal geometric structures. Code and project materials are available at https://github.com/naver-ai/beneath-cot.
Here is the structured explanation based on the research paper "Beneath the Surface of Chains-of-Thought: A Mechanistic Interpretation of Reasoning Operations in LLMs."
1. The Problem
The core issue this paper addresses is a gap in our understanding of how Large Language Models (LLMs) handle multi-step reasoning. We know that when an LLM solves a complex problem, it generates a "chain-of-thought" (CoT) trace—a sequence of intermediate steps leading to the final answer. We can read these traces and see distinct stages, like "decomposing a problem" or "recalling a formula."
However, the paper asks: Do these textual stages correspond to actual, distinguishable patterns inside the model's neural network? Or, is the model just predicting the next likely word based on surface statistics?
The authors argue that while reasoning operations are explicitly distinguished in the generated text, little was known about how they are "geometrically organized" in the model's representation spaces. The central question is whether the model’s internal activations actually encode which reasoning operation is being performed, beyond just the words being used.
2. How It Works (The Technical Mechanics)
The authors investigate this by treating the CoT trace not as a random string of tokens, but as a sequence of "reasoning operation chunks." They categorize these chunks using Polya’s problem-solving framework (a classic math pedagogy), identifying eight frequent operations: Extraction, Direct mapping, Decomposition, Recall, Deduction, Algebraic manipulation, Arithmetic computation, and Final answer.
To understand the internal geometry, they use a technique called representation probing. Essentially, they train a simple classifier (Linear Discriminant Analysis) to look at the model’s hidden activations and ask: "Can you tell which of the eight reasoning operations is happening just by looking at the brain activity (hidden states)?"
Here is the concrete analogy a software engineer or product manager would grasp:
Think of the model's hidden layers like a complex, multi-dimensional spreadsheet. Different reasoning operations might use different columns or rows of this spreadsheet. The authors trained "probes" to check if, for a given chunk of text labeled "Decomposition," the model's activation pattern consistently lights up a specific set of coordinates in that spreadsheet, regardless of the specific math problem being solved. If the probe can reliably distinguish "Decomposition" from "Recall" based solely on the activations, it means the model has internally segregated these operations.
Key findings on how it works:
- Separability Peaks in the Middle: The probes found that these operation clusters are most distinct in the "middle layers" of the model. Early layers are noisy and often rely on local word cues; late layers are more abstract. The "sweet spot" for distinguishing the type of reasoning is in the middle.
- Distribution Across Spans: In early layers, the signal for a reasoning operation might be concentrated on just one or two "cue" tokens (e.g., the word "therefore" might trigger a classification). But as you move to middle layers, the signal becomes distributed across the entire span of the operation. The model starts "thinking" with the whole chunk, not just a keyword.
- Identical Tokens, Different Meanings: Perhaps the most striking result is that the model represents the same word differently depending on the surrounding reasoning context. For example, the word "is" might appear in both an "Extraction" step and a "Decomposition" step, but the model's internal activation for that "is" will be different depending on which operation is active. The text identity is fixed, but the representation is contextualized.
3. Key Results & Benchmarks
The paper presents quantitative results using AUROC (Area Under the Receiver Operating Characteristic Curve) and AUPRC (Area Under the Precision-Recall Curve). In plain language, these numbers measure how well the probe can distinguish one operation from all others.
The Bottom Line Numbers:
- Strong Separability: Across all three models tested (Qwen2.5-7B, Qwen3-8B, and Gemma4-31B), the probes achieved high AUROC scores, often exceeding 0.90 (where 0.5 is random chance). This means the model's hidden states are quite good at signaling which reasoning operation is in play.
- The Peak: Separability is strongest in the middle layers, confirming the mechanistic finding.
- Context is King: When the authors looked at "identical surface tokens" (e.g., the word "x" appearing in different operations), the probes could still distinguish the operations. This suggests that simply knowing the word isn't enough; the model's internal state carries the operation label.
Plain-Language Translation of Benchmarks:
"The model's internal activations can correctly identify whether a chunk of text is performing 'Algebraic Manipulation' versus other operations about 92% of the time on the standard test set used by the field. Furthermore, as the reasoning progresses into the middle layers of the model, the signal becomes more evenly spread across the tokens of the chunk, meaning the model isn't relying on a single 'trick word' to identify the operation."
4. Why It Matters (Key Takeaways)
The authors distill their findings into four key takeaways regarding broader significance:
- Representational Correspondence: Language models maintain a faithful correspondence between the reasoning operations expressed in text and their internal geometric organization. The "thought" you read on the screen has a matching "shape" in the model's neural activations.
- Context-Dependent Processing: Reasoning operations are not locally self-contained. The representation of a new operation depends on the preceding context. If you mask (hide) the previous reasoning steps, the model's ability to correctly identify the current operation drops. This implies that the model "builds" its reasoning step-by-step, using the previous step as input for the next.
- Robustness to Errors: The internal geometry of operations persists even when the model makes factual errors. If the model messes up a calculation, it still activates the "Arithmetic Computation" representation, even if the result is wrong. This suggests a dissociation between how the model reasons (the operation) and whether its facts are correct.
- Pathways for Intervention: By identifying where and how operations are represented, this work opens the door to "latent space interventions." In the future, researchers might be able to intervene on specific activation directions to steer the model's reasoning process, potentially improving reliability or detecting reasoning failures without changing the model's weights.
Summary
In essence, this paper provides the first systematic evidence that LLMs develop a kind of "internal syntax" for reasoning. Just as a human mathematician might have distinct mental modules for "decomposing a problem" versus "checking a calculation," the LLM develops distinguishable geometric patterns in its neural network for these functional steps. These patterns are strongest in the middle layers, distributed across token spans, and crucially dependent on the preceding reasoning context.
Want to understand AI papers like this from scratch?
Follow the free AI Learning Roadmap →