一种自监督视觉学习方法,通过对比同一图像不同裁剪视图训练模型捕捉语义一致的特征,无需人工标注标签
The choice of pretraining (ImageNet-1K vs ImageNet-21K, supervised vs self-supervised like MAE/DINO) significantly impacts downstream performance, potentially causing several percentage points of difference