Vision Transformer (ViT)

Vision Transformer (ViT) applies transformer attention mechanisms to image patches for classification and representation learning. It is widely used in multimodal stacks with CLIP and in segmentation systems like Segment Anything Model (SAM).

Related terms

Related terms

Your next idea starts here

GPT 5.6 Terra

Create personal portfolio

Build startup site

Launch landing page

Start company blog

Framer UI showing Pages, Layers and Assets panels titles
Framer UI showing Pages, Layers and Assets panels titles