One of the pivotal recent challenges in neural network interpretability is polysemanticity, where a single neuron is activated by multiple, often unrelated concepts, hindering clear functional understanding.
Although prior work has explored this phenomenon, existing approaches remain architecture-specific and depend on manual heuristics such as a fixed number of concept clusters ($K$), limiting their generality and scalability--especially for modern Transformer-based models.
To address these limitations, we introduce SPICE (Simple Polysemantic Feature Interpretation via Clustering-based Explanation), a generalizable framework for analyzing polysemanticity in deep vision architectures.
SPICE avoids architecture-dependent propagation rules, enabling the first systematic comparison of polysemanticity across both CNNs and Transformers, and automatically determines the number of concept clusters per neuron, eliminating reliance on a preset $K$ and supporting scalable analysis for large models.
Using SPICE, we conduct a comprehensive investigation into how polysemanticity emerges, varies across depth and architecture, and forms through distinct computational pathways.
SPICE (Simple Polysemantic Feature Interpretation via Clustering-based Explanation)
SPICE
Simple
Polysemantic
Interpretation
Clustering-based
Explanation