Перейти к содержимому

Embeddings: teaching a machine what words mean

Jimmy's Tech Deep Dive

0:00 / 0:00

Embeddings: teaching a machine what words mean

6 просмотров · 7 дней назад
Jimmy's Tech Deep Dive
9 подписчиков
6 просмотров · 7 дней назад
A computer has never read a word. It only sees numbers. So how does it know that "cancel my plan" means "end your subscription"? This is a full walkthrough of text embeddings, built from the ground up. We start with the basic constraint that models are just matrix multiplication, show why one-hot encoding destroys meaning, and use the distributional hypothesis to turn words into positions in space. From there we cover how Word2Vec and GloVe are actually trained, why one fixed vector per word eventually broke, how ELMo and BERT made vectors depend on the sentence, and how all of this ends up powering semantic search, RAG, recommendations and clustering. We finish with the 2026 model landscape, what MTEB scores do and don't tell you, and how to choose a model for your own data. • Why one-hot encoding makes cat and dog as unrelated as cat and democracy • Skip-gram, CBOW, negative sampling, and why the predictions are thrown away • GloVe's count ratios, FastText's character chunks, and the frozen-vector problem • ELMo, BERT, masking and attention: the shift to contextual embeddings • Pooling, contrastive training, vector databases and two-stage retrieval with reranking • RAG, chunking pitfalls, and reading MTEB leaderboards without being misled Chapters 0:00 Machines only see numbers 0:37 The one-hot encoding dead end 1:16 The distributional hypothesis 1:51 Words as points in space 2:56 Cosine similarity and analogies 3:34 How Word2Vec is trained 5:15 GloVe and count ratios 6:26 One word, one vector 7:00 ELMo, BERT and context 8:03 Sentence embeddings and pooling 8:58 Vector databases 9:33 RAG and chunking 10:40 Recommendations and multimodal 11:14 The 2026 model landscape 12:12 Benchmarks are a shortlist 12:50 The whole chain in one idea Sources What Are Word Embeddings? | IBM - https://www.ibm.com/think/topics/word... Word Embeddings: Word2Vec, GloVe and Vector Semantics - https://mbrenndoerfer.com/writing/wor... Understanding the difference between Contextual and Static Embeddings -   / understanding-the-difference-between-conte...   Embeddings for Semantic Search and RAG | Snowflake - https://www.snowflake.com/en/artifici... Best Embedding Models in 2026: OpenAI vs Open Source by MTEB | Cognee - https://www.cognee.ai/best-embedding-... MTEB 2026: State of the Embeddings Benchmark - https://app.ailog.fr/en/blog/news/rag... #Embeddings #RAG #VectorSearch #MachineLearning #NLP