What AstroPT knows about galaxies, and what that can teach us about LLMsExplained for Beginners
UniverseTBD, Kshitij Duraphe, Aman Kumar +2 more
Abstract
Interpretability research increasingly asks when concepts emerge during training and whether linear probes recover real structure, but in language models these claims are hard to validate because language offers little ground-truth ordering of concepts or relationships among them. We propose the use of astronomical ground truth through AstroPT, a transformer trained on millions of galaxy images, as a calibration testbed. AstroPT is an LLM-like model trained within a domain where the difficulty ordering of concepts and the relations among them are known in advance. Probing frozen representations across checkpoints, layers, model sizes, and objective choices, we find that galaxy properties emerge in a fixed order that tracks their known difficulty---quantities written almost directly into the pixels (band magnitude) become decodable early in training and shallow in the network, while multiband/spectra based and inferred quantities (such as redshift and specific star formation rate) emerge later and deeper. This order is invariant to our tested training objectives, and scales in magnitude but not in sequence with capacity. Our linear probe directions further recover the known physical structure among galaxy properties. Our findings suggest that astronomy offers a controlled sandbox for calibrating mechanistic interpretability methods we otherwise apply to LLMs blind.
What AstroPT Knows About Galaxies, and What That Teaches Us About LLMs
The Problem
If you are a product manager or engineer working with large language models (LLMs), you have likely heard the buzz around “mechanistic interpretability.” This is the effort to reverse-engineer what a model is thinking, layer by layer, to understand if it is relying on logic, hallucinating, or developing some strange internal geometry. But there is a fundamental problem: language is a messy, open-ended playground. Unlike physics or chemistry, there is no “ground truth” answer key for what a model should be learning. We can’t easily say, “Yes, this neuron is definitely tracking the concept of ‘cause and effect’” because there is no agreed-upon dictionary of concepts to check against.
This paper steps into that gap. The authors argue that we need a calibration testbed—a domain where we already know the answers, so we can verify our interpretability tools before pointing them at the opaque black box of an LLM. They propose using astronomy. Specifically, they trained AstroPT, a transformer model that looks and learns like a text-based LLM, but processes galaxy images instead of words.
The real-world gap this addresses is the inability to validate interpretability claims in LLMs. By creating a domain with a built-in “answer key”—the known physics of galaxies—they can finally test when and how concepts emerge during training. For the rest of us, this matters because any insights gained here could improve our ability to trust, debug, and steer the LLMs we use every day.
How It Works (The Technical Mechanics)
The mechanics of AstroPT are designed to mimic an LLM as closely as possible. Imagine taking a beautiful image of a galaxy and slicing it into a grid of small patches, much like a jigsaw puzzle. These patches are then treated as “tokens”—the basic units of information—just as words are tokens in a text model.
AstroPT is trained in two familiar modes:
- Autoregressive (AR) mode: Predicting the next patch in the sequence (like a GPT model).
- Masked Autoencoding (MAE) mode: Reconstructing missing patches (like a BERT model).
The model learns by trying to rebuild the original image from these patches.
Now, here is the clever part. In astronomy, we already know a "difficulty ladder" of galaxy properties. Some are written right into the image pixels, while others require complex physics to deduce.
- Easy (written in pixels): r-band magnitude (brightness). This is essentially just measuring the total light coming from the galaxy. It is the astronomical equivalent of seeing a word’s spelling.
- Medium: Redshift. This tells us how fast a galaxy is moving away from us, which requires looking at specific colors or patterns across multiple bands of light.
- Hard (inferred): Specific Star Formation Rate (sSFR). This is a measure of how quickly a galaxy is making new stars relative to its size. You can’t just look at a pretty picture to get this; you need complex models of stellar evolution.
The authors probe the frozen model at various stages—different training steps, different network layers, and different model sizes—to see when these properties become "decodable" (predictable) from the model's internal representations.
The Analogy
Think of this like a software engineer onboarding a new hire. The "easy" properties (brightness/magnitude) are like skills the hire picks up on day one—obvious, surface-level things. The "medium" properties (redshift) require a bit more training and experience to access. The "hard" properties (sSFR) are like specialized knowledge that only emerges after deep immersion in the codebase and perhaps some mentorship. The key takeaway is that this learning order is predictable and consistent, regardless of how big the model is or which training recipe is used.
Key Results & Benchmarks
The results are striking and beautifully illustrate the "difficulty ordering" the authors predicted.
- The Emergence Order: As training progresses, the model becomes able to decode properties in a specific sequence. First, it learns r-band magnitude. This makes sense because the brightness is right there in the pixels. Next, it picks up redshift. Finally, and most stubbornly, it struggles to decode sSFR within the scope of a single training epoch.
- The Depth Order: This emergence isn't just happening across time (training steps); it's also happening across the network's depth (layers). r-band magnitude is decodable early on and across many layers (shallow learning). Redshift becomes decodable deeper in the network. sSFR remains weak even in the deepest layers after one epoch of training.
- Scaling: Interestingly, making the model bigger (going from 1M to 100M parameters) doesn't rearrange this order. A larger model just gets better at decoding everything, but it still decodes magnitude before redshift, and redshift before sSFR. The sequence is stable; the proficiency improves with scale.
Furthermore, the "probe directions"—the specific pathways the model uses to represent these properties—recover the known physical structure of the universe. The model aligns luminosity with stellar mass (big galaxies have more light), anti-aligns sSFR with mass (as galaxies grow, they often stop forming stars), and aligns redshift with mass (brighter, more massive galaxies are visible across greater distances).
Why It Matters (Key Takeaways)
This work is a significant "proof of concept" for the AI interpretability field. Here are the key takeaways:
- A Controlled Sandbox: Astronomy provides a rare "ground truth" that language lacks. Because the physical relationships between galaxy properties are well-understood, we can tell if a model is "getting it right" or just memorizing statistical coincidences.
- Predictable Emergence: The order in which concepts emerge is robust. It tracks the physical complexity of the data, not the whims of the training algorithm. This suggests that when we see certain behaviors in LLMs, there might be an underlying structural logic tied to the data's inherent difficulty.
- Probe Geometry is Meaningful: The fact that the model's internal "directions" for properties like mass and luminosity align with known physics suggests that these models are learning something resembling a coherent world model, rather than just pattern-matching.
- What to Watch For: The main limitation is that this is correlational and limited to one epoch of training. We don't yet know if the order persists indefinitely or if it shifts with different data regimes. However, the invariance to model size and training objective is a strong signal that these are fundamental properties of the learning task, not quirks of a specific setup.
In essence, this paper gives us a "control group" for AI interpretability. By showing that we can reliably track concept emergence in a domain we understand, it paves the way for more rigorous auditing of the large models that power everything from search engines to code assistants. We now have a way to ask: "Is the model learning this concept early because it's easy, or because it's privileged in the architecture?" and answer it with data.
Want to understand AI papers like this from scratch?
Follow the free AI Learning Roadmap →