Visual deep-dives on messaging, data structures, and interview topic comparisons.
When 30k req/s hit a box that can only do 10k, extra servers take time and the database may not scale with you. Drop the overflow — keep checkout and payment, shed recommendations — via concurrency limits, bounded queues, health checks, or priority.
Retrying is not enough: 10,000 clients waiting 1s then 5s then 10s retry in lockstep and stampede the dependency again. Exponential backoff spaces attempts; jitter (1.2s, 5.5s, 10.1s) desynchronizes them.
Stop calling a failing dependency. Closed is normal traffic, Open fails fast with no calls, Half-Open lets a few requests through to see if the service recovered — Resilience4j, PyBreaker, gobreaker.
When someone posts, do you push the update into every follower feed (fan-out on write) or assemble the feed when they open the app (fan-out on read)? Same idea as Twitter timelines — write is fast to read, read is cheaper to write.
Write the business row and an outbox event in the same database transaction, then a worker publishes pending events to Kafka. That way payment and analytics never miss an order that already committed.
When a hot cache key expires, thousands of requests miss at once and stampede the database. One request should rebuild the cache (often with a lock); the rest wait or keep serving stale data.