s01-p08: Vision Transformer (ViT) | CS4CV
Computer Science for Computer Vision
0:00 / 0:00
s01-p08: Vision Transformer (ViT) | CS4CV
12 просмотров · 8 дней назад
Computer Science for Computer Vision
9 подписчиков
12 просмотров · 8 дней назад
Course: Computer Science for Computer Vision
Channel: @CS4CV ( / @cs4cv )
Playlist: Section 1 – Attention Mechanisms & Transformers ( • CS4CV Section 01: Attention Mechanisms & T... )
In this lecture from the Computer Science for Computer Vision course, we introduce the Vision Transformer (ViT) and explore how Transformer architectures can be applied to image recognition.
Topics covered in this lecture:
• Vision Transformer (ViT)
• Why Transformers can be used for images
• Image patches and patch embeddings
• Converting an image into a sequence of tokens
• Positional embeddings
• Transformer Encoder for image classification
• ViT architecture
• Patch Embedding implementation with PyTorch
• Classification with ViT
We also discuss the motivation behind dividing an image into fixed-size patches and how these patches are converted into vector representations before being processed by a Transformer encoder.
#VisionTransformer #ViT #ComputerVision #Transformers #DeepLearning #AttentionMechanism #MachineLearning #ArtificialIntelligence #PyTorch #ComputerVisionCourse #AI #NeuralNetworks