select talks
In Search of Universal Concepts
Bristol, FOMO Amsterdam, Ludwig-Maximilians-Universität München, Erlangen, University of Technology Nuremberg, University of Zürich, ETH Zürich
Modern vision systems can recognize and generate complex visual content, yet we still know little about the internal concepts they use to represent and reason about the world. This talk introduces two unsupervised frameworks: Video Transformer Concept Discovery (VTCD), which automatically discovers interpretable spatiotemporal concepts and reveals that many are universal across video transformers, and Universal Sparse Autoencoders (USAEs), which learn a shared concept space to discover and align interpretable concepts across diverse vision models. Together, these methods suggest that neural networks organize visual knowledge around reusable conceptual components, opening new avenues for interpreting, comparing, and designing vision systems.