How Uber Matches Riders and Drivers in Real Time

You open the Uber app and request a ride. Within seconds, the app tells you about nearby drivers. You confirm. A driver accepts. A car appears on the map moving toward you. The whole process takes less than 30 seconds. Behind that 30 seconds is a system that must track the location of millions of drivers, match riders to drivers based on proximity and preference, recalculate routes in real time, adjust pricing dynamically, and handle payments across dozens of countries. All while dealing with network latency, GPS inaccuracy, and the inherent unpredictability of human behavior. ...

March 23, 2026 · 8 min · 1530 words · Ahmad Hassan

Rate Limiting

Your API handles 100 requests per second comfortably. A user writes a script that sends 10,000 requests per second. Your servers crawl. Legitimate users get timeouts. Your database connection pool exhausts. The entire system degrades because of one bad actor. This is why rate limiting exists. It’s not just about preventing abuse. It’s about protecting the system from itself. Every resource is finite. Rate limiting ensures no single consumer consumes more than their fair share. ...

March 20, 2026 · 6 min · 1134 words · Ahmad Hassan

Message Queues

A user places an order. Your application needs to process payment, update inventory, send a confirmation email, update analytics, and trigger a fraud check. If the application calls each service directly and synchronously, what happens when the email service is slow? The user waits. What happens when analytics goes down? The whole chain breaks. What happens when traffic spikes on Black Friday? Every service in the chain must handle peak load simultaneously. ...

March 17, 2026 · 5 min · 1011 words · Ahmad Hassan

Consistent Hashing

You have 5 cache servers. You hash each user ID modulo 5 to decide which server holds their data. User 42 goes to server 2. User 87 goes to server 2 as well. Everything works. Then traffic grows. You add a 6th server. Now you hash modulo 6. User 42 hashes to server 0. User 87 hashes to server 3. Almost every user’s data is now on the wrong server. Your cache hit rate drops to near zero. Every request that would have been a cache hit becomes a cache miss. Your database drowns. ...

March 14, 2026 · 5 min · 1029 words · Ahmad Hassan

How WhatsApp Handles Billions of Messages

WhatsApp serves over 2 billion users with fewer than 100 engineers. That ratio is absurd. A company the size of a small startup powering the world’s largest messaging platform. The architecture that makes this possible is worth understanding because the design decisions are deliberately different from what most teams would choose. The foundation of WhatsApp’s backend is Erlang and the BEAM virtual machine. Erlang was designed at Ericsson in the 1980s for telephone switches. These systems had specific requirements. They must never go down. They must handle millions of simultaneous connections. They must update without restarting. And they must process messages with microsecond latency. ...

March 11, 2026 · 7 min · 1481 words · Ahmad Hassan

Consensus Algorithms

Three generals are surrounding a city. They must attack at dawn or retreat. If they all attack, they win. If they all retreat, they live to fight another day. If some attack and some retreat, they are destroyed. They can only communicate by messenger. One general might be a traitor. How do they reach agreement? This is the Byzantine Generals Problem. And it cuts to the heart of distributed systems. ...

March 8, 2026 · 5 min · 970 words · Ahmad Hassan

API Gateway

You have ten microservices. Each has its own URL, its own authentication logic, its own rate limiting. Your frontend team needs to know all ten URLs. Your mobile app needs to call ten different endpoints. When a service changes its address, every client breaks. Now imagine your system grows to fifty services. Then a hundred. The frontend is now managing a web of direct connections. Authentication is duplicated across every service. Rate limiting is inconsistent. Logging is scattered. Monitoring is a nightmare. ...

March 5, 2026 · 4 min · 839 words · Ahmad Hassan

Database Replication

Your database holds all your data on one machine. That machine dies. Your entire application goes down. Every user sees errors. Revenue stops. And there is nothing you can do until that machine comes back online. This is the availability problem. The fix is replication. Keep multiple copies of the same data on different machines. If one fails, another takes over. But replication is not just about Copies. It raises real design questions. Who accepts writes? How do copies stay in sync? What happens when they fall behind? How do you handle conflicts when two copies disagree? ...

March 2, 2026 · 5 min · 883 words · Ahmad Hassan

How YouTube Serves Billions of Videos

Over 500 hours of video are uploaded to YouTube every minute. Over a billion hours of video are watched every day. The scale is hard to comprehend. Let’s break down how a system this large actually works. When you upload a video to YouTube, nothing about the experience suggests what’s happening behind the scenes. The upload finishes, you see a progress bar, and eventually the video is live. But in that gap, the system does an enormous amount of work. ...

February 28, 2026 · 7 min · 1367 words · Ahmad Hassan

Database Sharding

Your database has one server. It holds all your data. It works fine until it doesn’t. The CPU maxes out. Disk I/O crawls. Queries that took 5ms now take 500ms. You add more RAM. You upgrade the CPU. You get a bigger machine. But eventually, one machine cannot keep up. This is the vertical scaling ceiling. You make the box bigger until you cannot make it any bigger. The alternative is horizontal scaling. Instead of one giant database, you spread the data across multiple smaller databases. Each one holds a subset of the data. Each one handles a subset of the traffic. Together, they behave like one logical database. ...

February 26, 2026 · 5 min · 859 words · Ahmad Hassan
ESC