TempatGunting

Vision Transformers: From ViT to DINOv2 and Modern Visual Foundation Models

Images as Sequences of Patches

The Vision Transformer (ViT) paper from Dosovitskiy et al. asked a provocative question: what if we just apply a standard NLP transformer to images, with no convolutional layers at all? The answer turned out to be remarkably simple. Split the image into non-overlapping patches, flatten each patch into a vector, project it into the model's embedding dimension, and treat the resulting sequence exactly like a sequence of word tokens in BERT or GPT.

A 224×224 image with 16×16 patches produces a sequence of 196 patch tokens. Each patch is flattened from a 16×16×3 tensor (768 values) into a 1D vector and linearly projected to the model's hidden dimension. Learnable positional embeddings are added to each patch token to encode its spatial position. A prepended CLS token (borrowed directly from BERT) aggregates global information through self-attention and serves as the input to the classification head.

import torch
import torch.nn as nn

class PatchEmbedding(nn.Module):
    def __init__(self, img_size=224, patch_size=16, in_channels=3, embed_dim=768):
        super().__init__()
        self.num_patches = (img_size // patch_size) ** 2
        self.proj = nn.Conv2d(
            in_channels, embed_dim,
            kernel_size=patch_size, stride=patch_size
        )
        self.cls_token = nn.Parameter(torch.randn(1, 1, embed_dim))
        self.pos_embed = nn.Parameter(
            torch.randn(1, self.num_patches + 1, embed_dim)
        )

    def forward(self, x):
        B = x.shape[0]
        x = self.proj(x).flatten(2).transpose(1, 2)  # [B, num_patches, embed_dim]
        cls = self.cls_token.expand(B, -1, -1)
        x = torch.cat([cls, x], dim=1)
        return x + self.pos_embed

The decision to use a convolution for patch projection — technically a single conv2d with kernel_size=stride=16 — is a pragmatic shortcut. It's mathematically equivalent to the paper's linear projection but runs faster on GPUs. This minor implementation detail matters because it means ViT isn't purely attention-based: the patch projection is a convolutional operation, even if the rest of the architecture is pure transformer.

The Data Hunger Problem

The original ViT had a critical limitation: it required enormous pretraining datasets to match CNN performance. Trained on ImageNet-1K (1.2M images), ViT-B/16 underperformed a ResNet-50. Trained on JFT-300M (300M images), the same architecture exceeded every CNN on the same benchmarks. The reason is that CNNs encode strong inductive biases — translation equivariance and locality — that help them learn efficiently from limited data. Transformers learn these patterns from data instead, requiring much more of it.

DeiT (Data-efficient Image Transformers) addressed this with three techniques. First, aggressive data augmentation including RandAugment, random erasing, and Mixup provided the regularization that the missing inductive biases would otherwise supply. Second, knowledge distillation from a pretrained CNN teacher helped the transformer learn spatial features faster. Third, a careful training recipe with stochastic depth, repeated augmentation, and label smoothing stabilized training at smaller data scales.

The result was a ViT that matched CNN accuracy on ImageNet-1K without external data. This was an important proof point: vision transformers don't inherently need hundreds of millions of images — they need the right training recipe. The data augmentation strategies that enabled this breakthrough have since become standard practice for all vision transformer training.

Swin Transformer: Hierarchical Vision with Windowed Attention

Standard ViT computes global self-attention — every patch attends to every other patch. For a 224×224 image with 16×16 patches, that's 196 tokens, and the quadratic attention cost is manageable. But for high-resolution images (1024×1024 or larger) or dense prediction tasks (segmentation, detection), the patch count explodes to thousands, and global attention becomes prohibitively expensive.

Swin Transformer solves this with two innovations. First, self-attention is computed within local windows of 7×7 patches, reducing the attention complexity per token from O(N) to O(49) — constant regardless of image size. Second, consecutive layers alternate between regular and shifted windows, where the shift is half a window size in both dimensions. This shifting mechanism allows information to flow across window boundaries without paying the cost of global attention.

The second innovation is the hierarchical architecture. Swin starts with small 4×4 patches and progressively merges them — concatenating 2×2 groups of patches and projecting to a larger dimension — creating a feature pyramid with resolutions at 1/4, 1/8, 1/16, and 1/32 of the input size. This multi-scale representation is exactly what dense prediction tasks need, and it directly replaces the CNN backbone in frameworks like Mask R-CNN and UPerNet.

Stage 1 56×56 patches dim=96 W-MSA × 2 Stage 2 28×28 patches dim=192 Stage 3 14×14, dim=384 Stage 4 7×7, dim=768 merge merge merge Each stage: Window attention → Shifted window attention → Patch merging Resolution halves, channels double at each merge

Self-Supervised Pretraining: DINO and Masked Image Modeling

The next breakthrough wasn't an architecture change — it was a training methodology change. Instead of pretraining on labeled datasets (supervised learning), self-supervised approaches learn visual features from raw images without any human annotations. Two families of self-supervised methods have proven dominant for vision transformers.

DINO (Self-DIstillation with NO labels) uses a student-teacher framework where both networks are vision transformers processing different augmented views of the same image. The teacher is an exponential moving average of the student's weights. The student learns to match the teacher's output distribution using a cross-entropy loss over the CLS token representations. The key insight is centering and sharpening the teacher's outputs to prevent mode collapse (where both networks output the same constant vector regardless of input).

Masked image modeling (MAE, BEiT) follows the BERT playbook: mask 75% of image patches, encode only the visible patches, then decode and reconstruct the masked patches. MAE's aggressive 75% masking ratio is crucial — it forces the encoder to learn holistic visual understanding rather than just interpolating nearby pixels. The decoder is lightweight and discarded after pretraining, making the approach compute-efficient. These pretraining approaches benefit from the same distributed training infrastructure used for language models.

DINOv2: The Visual Foundation Model

DINOv2 from Meta AI combined the best ideas from both self-supervised families: the student-teacher distillation from DINO and the masked image modeling from iBOT, trained on a curated dataset of 142 million images. The result is a vision foundation model whose frozen features — without any fine-tuning — match or exceed supervised models on classification, segmentation, depth estimation, and retrieval.

The curation pipeline is as important as the architecture. DINOv2's training data was assembled from web-crawled images filtered through multiple quality stages: deduplication using copy detection embeddings, NSFW filtering, resolution and aspect ratio filtering, and retrieval-based balancing to ensure coverage across visual concepts. This automatic curation produced a training set comparable in quality to manually curated datasets but at far larger scale.

import torch
from transformers import AutoModel, AutoImageProcessor

processor = AutoImageProcessor.from_pretrained("facebook/dinov2-base")
model = AutoModel.from_pretrained("facebook/dinov2-base")

# Extract features (no fine-tuning needed)
inputs = processor(images=image, return_tensors="pt")
with torch.no_grad():
    outputs = model(**inputs)

# CLS token for global features
global_features = outputs.last_hidden_state[:, 0]  # [B, 768]

# Patch tokens for dense prediction
patch_features = outputs.last_hidden_state[:, 1:]  # [B, 196, 768]

DINOv2's strength is its versatility. The same frozen backbone produces features that work for object detection, semantic segmentation (by upsampling patch features), monocular depth estimation, image retrieval, and classification — all without task-specific training. This "frozen backbone" paradigm changes how teams build vision applications: instead of training end-to-end models, they extract DINOv2 features and train lightweight task-specific heads.

Positional Encoding: The Spatial Awareness Problem

Self-attention is permutation-invariant — without positional information, the transformer treats its input as a set, not a sequence. For images, spatial position is critical: a patch in the top-left corner should be processed differently than the same patch in the bottom-right corner. Vision transformers encode position through several mechanisms.

Absolute learnable embeddings, used by the original ViT, add a learned vector to each patch position. These embeddings encode both position and proximity: nearby patches have similar positional embeddings. The limitation is that they're fixed to the training resolution — a model trained on 224×224 images (196 patches) can't directly process 384×384 images (576 patches) without interpolating the positional embeddings.

Relative positional bias, used by Swin Transformer, encodes the relative distance between patches rather than absolute positions. Each attention head learns a bias matrix indexed by the relative position between query and key patches. This approach generalizes better to different resolutions because the relative distances remain meaningful even when the total number of patches changes.

Rotary position embeddings (RoPE), borrowed from language models like LLaMA, have recently been adapted for vision transformers. RoPE encodes position by rotating the query and key vectors, with the rotation angle proportional to position. The dot product between rotated vectors naturally decays with distance, providing a smooth positional signal. RoPE supports arbitrary sequence lengths without interpolation, making it attractive for variable-resolution vision applications where efficient attention is critical.

Efficiency Techniques for Production Vision Transformers

Deploying vision transformers in production requires managing their computational cost. Several techniques reduce latency and memory without significant accuracy loss.

Token pruning removes uninformative patch tokens during the forward pass. An attention-based scoring mechanism identifies patches that contribute little to the output (background regions, uniform textures) and drops them after the first few transformer layers. Aggressive pruning can remove 50-70% of tokens, reducing compute by a similar factor with less than 1% accuracy loss on most tasks.

Knowledge distillation compresses a large ViT into a smaller student model. The student learns from both the ground truth labels and the teacher's soft predictions. DeiT demonstrated that distilling from a CNN teacher is particularly effective for vision transformers, as the CNN's inductive biases provide complementary supervision. The connection to LLM distillation techniques is direct — the core methodology transfers across modalities.

Quantization reduces weight and activation precision from FP32 to INT8 or lower. Vision transformers are more sensitive to quantization than CNNs due to the softmax in self-attention, which amplifies quantization errors in the attention logits. Post-training quantization with careful calibration achieves INT8 with under 0.5% accuracy loss. For INT4, quantization-aware training is typically necessary. These trade-offs mirror the broader quantization landscape across model architectures.

Vision Transformers for Dense Prediction

Classification was the first success, but dense prediction tasks — segmentation, detection, depth estimation — are where vision transformers have had the largest practical impact. The challenge is that these tasks require high-resolution feature maps, while ViT produces features at a single resolution (1/16 of the input).

ViTDet uses the plain ViT backbone and adds simple upsampling with deconvolution layers to create a feature pyramid from a single-scale feature map. Despite its simplicity, ViTDet with a ViT-H backbone set state-of-the-art results on COCO object detection, demonstrating that the strong features from a large transformer backbone compensate for the lack of a built-in multi-scale architecture.

For semantic segmentation, the standard approach attaches an MLP decoder head to the patch-level features. Segmenter uses a mask transformer decoder that produces per-class masks from patch embeddings. DINOv2's patch features are strong enough that a simple linear probe (one linear layer per class) achieves competitive segmentation results without any decoder architecture at all.

The Road Ahead: Scaling Laws and Multimodal Fusion

Vision transformer scaling laws follow a similar pattern to language model scaling: performance improves predictably with compute, data, and model size. DINOv2 demonstrated that a ViT-g (1.1B parameters) significantly outperforms a ViT-L (304M parameters) on virtually every benchmark, and the scaling curve shows no sign of saturation. The question isn't whether bigger vision transformers are better — they are — but whether the improvement justifies the cost.

The more transformative trend is multimodal fusion. Models like CLIP, SigLIP, and Florence use contrastive learning to align vision transformer features with text embeddings, creating shared representation spaces where images and text are directly comparable. These vision-language models enable zero-shot classification, text-based image retrieval, and visual question answering without task-specific training. The architecture typically pairs a vision transformer with a language model, connected by a lightweight projection layer.

The vision transformer's dominance in computer vision is now as established as the transformer's dominance in NLP. The remaining questions are about efficiency (how to make them practical for edge deployment), scale (how far the scaling laws extend), and unification (whether a single architecture can handle images, video, 3D, and text as a unified modality). Research in sparse architectures may provide answers to the efficiency question.