Mechanistic interpretability
What the internals are made of, and what tooling reads them honestly.
- Learning a generative meta-model of LLM activationsICML 2026
- Causal tracing of object representations in LVLMsarXiv, Nov 2025
What the internals are made of, and what tooling reads them honestly.