首页 > AI前沿 > SPICE: Simple Polysemantic Feature Interpretation via Clustering-based Explanation

SPICE: Simple Polysemantic Feature Interpretation via Clustering-based Explanation

arXiv机器学习 2026-08-14 09:04 7 阅读 查看原文

One of the pivotal recent challenges in neural network interpretability is polysemanticity, where a single neuron is activated by multiple, often unrelated concepts, hindering clear functional understanding.

Although prior work has explored this phenomenon, existing approaches remain architecture-specific and depend on manual heuristics such as a fixed number of concept clusters ($K$), limiting their generality and scalability--especially for modern Transformer-based models.

To address these limitations, we introduce SPICE (Simple Polysemantic Feature Interpretation via Clustering-based Explanation), a generalizable framework for analyzing polysemanticity in deep vision architectures.

SPICE avoids architecture-dependent propagation rules, enabling the first systematic comparison of polysemanticity across both CNNs and Transformers, and automatically determines the number of concept clusters per neuron, eliminating reliance on a preset $K$ and supporting scalable analysis for large models.

Using SPICE, we conduct a comprehensive investigation into how polysemanticity emerges, varies across depth and architecture, and forms through distinct computational pathways.

SPICE (Simple Polysemantic Feature Interpretation via Clustering-based Explanation)
SPICE

Simple

Polysemantic

Interpretation

Clustering-based

Explanation