Перейти к содержимому

Ep 15 | Message Queues & Event-Driven Design: Kafka vs SQS vs Pub/Sub | System Design

Bits To Billions

0:00 / 0:00

Ep 15 | Message Queues & Event-Driven Design: Kafka vs SQS vs Pub/Sub | System Design

20 просмотров · 1 день назад
Bits To Billions
19 подписчиков
20 просмотров · 1 день назад
Message queues explained from zero, in 84 minutes - queues, pub/sub, logs, Kafka, SQS, Pub/Sub, delivery guarantees, idempotency, the outbox, partitions, consumer groups, dead letter queues, backpressure, sagas and event sourcing. No prior knowledge assumed. Every term is defined before it is used. Somebody says "just put it on a queue", and they are right. It is the correct fix. But you have not removed a problem. You have traded it for a different set of problems, and nobody tells you what they are. Your messages will now arrive twice. Your events will arrive out of order. One bad message will stop everything behind it. And your failure has not gone away - it has moved somewhere you are not looking. This is the long one, on purpose. We start from two programs trying to talk to each other, and we finish at sagas, event sourcing and schema registries - with a full section on when NOT to use a queue at all, which is the part that actually separates people in interviews. WHAT YOU'LL LEARN The three shapes: a queue for work, pub/sub for announcements, a log for both What a log really is, and why a bookmark you can move backwards changes everything At-most-once vs at-least-once - and that the whole choice is WHEN you acknowledge Why exactly-once DELIVERY is impossible: the Two Generals problem, from 1975 And why exactly-once PROCESSING is achievable anyway, using idempotency Idempotency from zero, idempotency keys, and how Stripe actually does it The dual-write problem, and the outbox pattern - the one good answer to it Partitions, keys and ordering: the same idea as sharding, with the same trap Consumer groups, rebalancing, and the rebalance loop that never ends Consumer lag: why you watch the MAXIMUM, and the age of the oldest message The poison pill, head-of-line blocking, and dead letter queues with alerts on them Retries done properly: exponential backoff, jitter, and what never to retry Backpressure, and why a queue defers a capacity problem rather than solving one A real incident: a stale consumer position, a replay, and duplicate emails Events, not instructions - and why the name you choose decides your coupling Choreography vs orchestration, and Netflix's precise complaint about the first one Sagas, compensating actions, and why a compensation is NOT a rollback (1987) Event sourcing: the real advantages, and the costs people have publicly regretted Schema evolution: backward, forward, full - and why a registry fails your deploy CHAPTERS 00:00 Intro 02:25 Two programs talking 04:23 Put something in the middle 06:00 The words 07:28 Shape one: the queue 09:31 Shape two: pub/sub 11:31 Shape three: the log 13:26 Inside the log 15:35 The three shapes, side by side 16:58 What you would actually use 18:53 What it costs 20:01 Is it actually safe? 21:46 How many times? 23:18 The most important choice 25:13 Why exactly-once is contentious 27:05 What Kafka actually sells 29:11 Idempotency, from zero 30:35 Idempotency keys 32:25 The deduplication table 34:24 The dual-write problem 35:42 The outbox pattern 37:37 What "in order" really means 39:33 Partitions and the key 41:36 Consumer groups 43:12 Bookmarks and how to lose them 45:06 Rebalancing 46:39 The loop that never ends 48:18 Ordering in the other systems 49:48 The number to watch 51:09 The poison pill 53:13 The dead letter queue 54:29 The graveyard nobody visits 56:14 Retrying properly 58:20 Backpressure 1:00:28 When it actually happened 1:02:08 Replay, carefully 1:03:24 What to actually watch 1:04:31 Events, not instructions 1:06:05 Who is in charge? 1:08:09 The saga 1:09:20 Compensation is not rollback 1:11:10 Event sourcing is a different thing 1:12:32 And what it costs 1:14:15 The messages change 1:16:14 When not to use a queue 1:18:05 Or just use your database 1:19:20 What do we notice? 1:20:38 The traps 1:21:43 Your checklist 1:22:43 Recap SOURCES Akkoyunlu, Ekanadham & Huber, 1975 - the Two Generals result Garcia-Molina & Salem - Sagas, 1987 (compensation is semantic, not a rollback) Apache Kafka documentation: transactions, exactly-once, rebalancing, KRaft AWS SQS / SNS docs: visibility timeout, redrive policy, FIFO limits and quotas Google Cloud Pub/Sub docs: ordering keys, dead letter topics, delivery attempts THE SERIES A complete system design course, zero to senior-interview level. Every episode assumes no prior knowledge. 1 What Is System Design | 2 The 5-Step Framework | 3 Estimation | 4 Scaling to 10M Users | 5 Load Balancers & API Gateways | 6 Caching | 7 Typing a URL | 8 SQL vs NoSQL | 9 Why Is Your Database Slow? | 10 Indexes & B-Trees | 11 LSM-Trees | 12 Replication, Lag & Quorums | 13 Sharding & Partitioning | 14 Consistent Hashing 15 Message Queues & Event-Driven Design - you are here Subscribe so you don't miss the rest of the series. New episodes regularly. #systemdesign #messagequeue #kafka #eventdriven #distributedsystems #backend #softwareengineering #techinterview