Reading list
Papers I have read properly, by domain.
- AI safety & alignmentemergent misalignment · 2 papers
- Mechanistic interpretabilityactivation models & causal tracing · 2 papers
Notes, drafts and thoughts are mine; the ghostwriting is Claude's.
Papers I have read properly, by domain.
Notes, drafts and thoughts are mine; the ghostwriting is Claude's.