首页 > AI前沿 > TwinMark: A Unified Watermark for Provable Survival Under Feature and Logit Distillation

TwinMark: A Unified Watermark for Provable Survival Under Feature and Logit Distillation

arXiv机器学习 2026-07-19 04:42 3 阅读 查看原文

We propose TwinMark, a watermarking scheme that reads a single SHAKE128 secret through two complementary linear functionals of model-output summaries:

1. a covariance projector against the carrier-set covariance (cov-Feat)

2. a class-conditional Fisher-aligned linear carrier decoded from class-mean logits (cc-FALC)

The two readouts share one bit vector and cover the two extraction surfaces of a deployed vision model:

1. a classifier API attacked by KL knowledge distillation (KD) (Std. KL-KD)

2. a representation-only host attacked by feature-matching KD (FM-KD)

Each readout admits a teacher-measurable a posteriori certificate that lower-bounds post-distillation detection power, and the two channels combine under a regime-restricted OR rule whose test statistic (calibrated null or bit vote) is selected by the exposed surface.

cov-Feat admits a rank-blind operator-norm certificate, cc-FALC admits a centered-logit-gap certificate that decouples bit capacity from class count:

at K=1024 in m=100 classes (a 10.24x over-encoding), the bit-vote attains z=23.0 sigma at a teacher-accuracy cost of +0.9+-0.2%p.

Across 13 attacks on CIFAR-10, CIFAR-100, and Mini-ImageNet, TwinMark verifies on every cell whose post-attack model retains task utility, survives cross-architecture distillation onto ResNet-18/50, VGG-16, and MobileNet-V3, and ports to GNSS few-shot, VOC detection, ISIC segmentation, and STL-10 SimCLR.