Ep 15 | Message Queues & Event-Driven Design: Kafka vs SQS vs Pub/Sub | System Design
Bits To Billions
0:00 / 0:00
Ep 15 | Message Queues & Event-Driven Design: Kafka vs SQS vs Pub/Sub | System Design
20 просмотров · 1 день назад
Bits To Billions
19 подписчиков
20 просмотров · 1 день назад
Message queues explained from zero, in 84 minutes - queues, pub/sub, logs,
Kafka, SQS, Pub/Sub, delivery guarantees, idempotency, the outbox, partitions,
consumer groups, dead letter queues, backpressure, sagas and event sourcing.
No prior knowledge assumed. Every term is defined before it is used.
Somebody says "just put it on a queue", and they are right. It is the correct fix.
But you have not removed a problem. You have traded it for a different set of
problems, and nobody tells you what they are. Your messages will now arrive twice.
Your events will arrive out of order. One bad message will stop everything behind
it. And your failure has not gone away - it has moved somewhere you are not looking.
This is the long one, on purpose. We start from two programs trying to talk to each
other, and we finish at sagas, event sourcing and schema registries - with a full
section on when NOT to use a queue at all, which is the part that actually separates
people in interviews.
WHAT YOU'LL LEARN
The three shapes: a queue for work, pub/sub for announcements, a log for both
What a log really is, and why a bookmark you can move backwards changes everything
At-most-once vs at-least-once - and that the whole choice is WHEN you acknowledge
Why exactly-once DELIVERY is impossible: the Two Generals problem, from 1975
And why exactly-once PROCESSING is achievable anyway, using idempotency
Idempotency from zero, idempotency keys, and how Stripe actually does it
The dual-write problem, and the outbox pattern - the one good answer to it
Partitions, keys and ordering: the same idea as sharding, with the same trap
Consumer groups, rebalancing, and the rebalance loop that never ends
Consumer lag: why you watch the MAXIMUM, and the age of the oldest message
The poison pill, head-of-line blocking, and dead letter queues with alerts on them
Retries done properly: exponential backoff, jitter, and what never to retry
Backpressure, and why a queue defers a capacity problem rather than solving one
A real incident: a stale consumer position, a replay, and duplicate emails
Events, not instructions - and why the name you choose decides your coupling
Choreography vs orchestration, and Netflix's precise complaint about the first one
Sagas, compensating actions, and why a compensation is NOT a rollback (1987)
Event sourcing: the real advantages, and the costs people have publicly regretted
Schema evolution: backward, forward, full - and why a registry fails your deploy
CHAPTERS
00:00 Intro
02:25 Two programs talking
04:23 Put something in the middle
06:00 The words
07:28 Shape one: the queue
09:31 Shape two: pub/sub
11:31 Shape three: the log
13:26 Inside the log
15:35 The three shapes, side by side
16:58 What you would actually use
18:53 What it costs
20:01 Is it actually safe?
21:46 How many times?
23:18 The most important choice
25:13 Why exactly-once is contentious
27:05 What Kafka actually sells
29:11 Idempotency, from zero
30:35 Idempotency keys
32:25 The deduplication table
34:24 The dual-write problem
35:42 The outbox pattern
37:37 What "in order" really means
39:33 Partitions and the key
41:36 Consumer groups
43:12 Bookmarks and how to lose them
45:06 Rebalancing
46:39 The loop that never ends
48:18 Ordering in the other systems
49:48 The number to watch
51:09 The poison pill
53:13 The dead letter queue
54:29 The graveyard nobody visits
56:14 Retrying properly
58:20 Backpressure
1:00:28 When it actually happened
1:02:08 Replay, carefully
1:03:24 What to actually watch
1:04:31 Events, not instructions
1:06:05 Who is in charge?
1:08:09 The saga
1:09:20 Compensation is not rollback
1:11:10 Event sourcing is a different thing
1:12:32 And what it costs
1:14:15 The messages change
1:16:14 When not to use a queue
1:18:05 Or just use your database
1:19:20 What do we notice?
1:20:38 The traps
1:21:43 Your checklist
1:22:43 Recap
SOURCES
Akkoyunlu, Ekanadham & Huber, 1975 - the Two Generals result
Garcia-Molina & Salem - Sagas, 1987 (compensation is semantic, not a rollback)
Apache Kafka documentation: transactions, exactly-once, rebalancing, KRaft
AWS SQS / SNS docs: visibility timeout, redrive policy, FIFO limits and quotas
Google Cloud Pub/Sub docs: ordering keys, dead letter topics, delivery attempts
THE SERIES
A complete system design course, zero to senior-interview level. Every episode assumes no prior knowledge.
1 What Is System Design | 2 The 5-Step Framework | 3 Estimation | 4 Scaling to 10M Users | 5 Load Balancers & API Gateways | 6 Caching | 7 Typing a URL | 8 SQL vs NoSQL | 9 Why Is Your Database Slow? | 10 Indexes & B-Trees | 11 LSM-Trees | 12 Replication, Lag & Quorums | 13 Sharding & Partitioning | 14 Consistent Hashing
15 Message Queues & Event-Driven Design - you are here
Subscribe so you don't miss the rest of the series. New episodes regularly.
#systemdesign #messagequeue #kafka #eventdriven #distributedsystems #backend #softwareengineering #techinterview