Embeddings: teaching a machine what words mean
Jimmy's Tech Deep Dive
0:00 / 0:00
Embeddings: teaching a machine what words mean
6 просмотров · 7 дней назад
Jimmy's Tech Deep Dive
9 подписчиков
6 просмотров · 7 дней назад
A computer has never read a word. It only sees numbers.
So how does it know that "cancel my plan" means "end your subscription"?
This is a full walkthrough of text embeddings, built from the ground up. We start with the basic constraint that models are just matrix multiplication, show why one-hot encoding destroys meaning, and use the distributional hypothesis to turn words into positions in space. From there we cover how Word2Vec and GloVe are actually trained, why one fixed vector per word eventually broke, how ELMo and BERT made vectors depend on the sentence, and how all of this ends up powering semantic search, RAG, recommendations and clustering. We finish with the 2026 model landscape, what MTEB scores do and don't tell you, and how to choose a model for your own data.
• Why one-hot encoding makes cat and dog as unrelated as cat and democracy
• Skip-gram, CBOW, negative sampling, and why the predictions are thrown away
• GloVe's count ratios, FastText's character chunks, and the frozen-vector problem
• ELMo, BERT, masking and attention: the shift to contextual embeddings
• Pooling, contrastive training, vector databases and two-stage retrieval with reranking
• RAG, chunking pitfalls, and reading MTEB leaderboards without being misled
Chapters
0:00 Machines only see numbers
0:37 The one-hot encoding dead end
1:16 The distributional hypothesis
1:51 Words as points in space
2:56 Cosine similarity and analogies
3:34 How Word2Vec is trained
5:15 GloVe and count ratios
6:26 One word, one vector
7:00 ELMo, BERT and context
8:03 Sentence embeddings and pooling
8:58 Vector databases
9:33 RAG and chunking
10:40 Recommendations and multimodal
11:14 The 2026 model landscape
12:12 Benchmarks are a shortlist
12:50 The whole chain in one idea
Sources
What Are Word Embeddings? | IBM - https://www.ibm.com/think/topics/word...
Word Embeddings: Word2Vec, GloVe and Vector Semantics - https://mbrenndoerfer.com/writing/wor...
Understanding the difference between Contextual and Static Embeddings - / understanding-the-difference-between-conte...
Embeddings for Semantic Search and RAG | Snowflake - https://www.snowflake.com/en/artifici...
Best Embedding Models in 2026: OpenAI vs Open Source by MTEB | Cognee - https://www.cognee.ai/best-embedding-...
MTEB 2026: State of the Embeddings Benchmark - https://app.ailog.fr/en/blog/news/rag...
#Embeddings #RAG #VectorSearch #MachineLearning #NLP