s01-p09: Implementing Vision Transformer (ViT) from Scratch in PyTorch | CS4CV
Computer Science for Computer Vision
0:00 / 0:00
s01-p09: Implementing Vision Transformer (ViT) from Scratch in PyTorch | CS4CV
26 просмотров · 9 дней назад
Computer Science for Computer Vision
10 подписчиков
26 просмотров · 9 дней назад
Course: Computer Science for Computer Vision
Channel: @CS4CV ( / @cs4cv )
Playlist: Section 1 – Attention Mechanisms & Transformers ( • CS4CV Section 01: Attention Mechanisms & T... )
In this lecture from the Computer Science for Computer Vision course, we implement a Vision Transformer (ViT) from scratch using PyTorch.
In this hands-on lecture, we focus on the implementation of the main components of a Vision Transformer and see how an image can be transformed into a sequence of patch embeddings and processed by a Transformer-based architecture.
Topics covered include:
• Vision Transformer implementation
• Patch Embeddings
• Image-to-patch representation
• Positional Embeddings
• Transformer Encoder
• ViT architecture in PyTorch
• Building ViT components from scratch
• Image classification with Vision Transformer
This lecture is designed to complement the theoretical discussion of Vision Transformers with a practical implementation in PyTorch.
#VisionTransformer #ViT #PyTorch #ComputerVision #DeepLearning #Transformers #MachineLearning #ArtificialIntelligence #AI #Coding #FromScratch #NeuralNetworks