Learn systems by using them.

27 visual, hands-on guides to the computer science ideas behind reliable software. Pick a topic, change the inputs, and see what happens.

Follow the path

Track 01Foundations

Core ideas that show up in almost every large system.

  1. Bloom filters

    A probabilistic set that can tell you “definitely not” or “maybe”. Add words, watch the bits flip, and trip a false positive yourself.

    3 ch
  2. Hashing

    From hash tables to consistent rings to hot key meltdowns - three chapters in one lab. Watch collisions form, see why plain hash % N breaks when a server leaves, and watch a single viral key overwhelm one server while its peers sit idle.

    3 ch
  3. Load balancing

    Follow traffic from anycast and global routing through L4 and L7 balancers, then experiment with IP hashing, sticky sessions, health failures, and backend selection algorithms.

    5 ch
  4. Big O notation

    Drag n and watch how each complexity class grows. The gap between O(log n) and O(n²) stops being abstract when you can see it explode.

    3 ch
  5. System design math

    A back-of-the-envelope calculator. Set daily active users and watch it cascade into requests/sec, servers, storage, and a monthly bill.

    4 ch
  6. Caching strategies

    Cache-aside, read-through, write-through, write-behind. Pick one, run traffic through it, and see the read path light up - and the database load drop.

    3 ch
  7. Local vs distributed caching

    Every pod keeps its own in-memory cache - fast, but they drift apart. Update the database, then watch local caches go stale while a shared Redis stays consistent. Invalidation, made visible.

    3 ch
  8. Queueing

    Requests pile up faster than a server can drain them. Send traffic into FIFO, LIFO and priority queues, watch the line grow, and see tail latency explode as load approaches capacity.

    3 ch
  9. Retries

    Retrying a failed request seems harmless - until every client retries at once and buries a struggling service. Compare naive retries, fixed delay, exponential backoff and jitter, and watch the retry storm form (or not).

    3 ch
  10. System evolution

    Step a product from a single VM to a sharded, replicated architecture - splitting the DB, measuring before scaling, load-balancing, caching, moving slow work to queues, then replicas and shards. Each stage shows the bottleneck it fixes and the trade-off it brings.

    3 ch

Track 02System design

Build familiar products one architectural decision at a time.

  1. Design a rate limiter

    Throttle a flood of requests four different ways - token bucket, leaking bucket, fixed window and sliding window. Open the tap, watch requests get accepted or 429'd in real time, and feel exactly where each algorithm leaks or bursts.

    4 ch
  2. Design a unique ID generator

    Sortable, 64-bit IDs across thousands of machines with no coordination - the Snowflake layout of timestamp, machine ID and sequence bits.

    3 ch
  3. Design a URL shortener

    Turn a long URL into a tiny one - base-62 encoding, hash collisions, and the read-heavy cache that makes the redirect instant.

    3 ch
  4. Design a key-value store

    Build a distributed hash map: consistent hashing for placement, replication for durability, and the quorum dial between consistency and availability.

    3 ch
  5. Design a web crawler

    A BFS frontier, politeness delays per host, and dedup with a bloom filter - crawl a tiny web without hammering any one domain.

    3 ch
  6. Design a notification system

    Fan a single event out to push, SMS and email through queues and workers, with retries and rate limits at each provider.

    3 ch
  7. Design a news feed

    Fan-out on write vs on read - the timeline trade-off that decides whether a celebrity post melts your database.

    3 ch
  8. Design a chat system

    WebSocket sessions, presence, and message ordering - deliver a message exactly once across a fleet of stateful chat servers.

    3 ch
  9. Design search autocomplete

    A trie of top queries served in milliseconds - prefix lookups, cached suggestions, and ranking by popularity as you type.

    3 ch
  10. Design a video platform

    Upload, transcode into multiple bitrates, and stream from a CDN - adaptive bitrate that scales from one viewer to millions.

    3 ch
  11. Design a file store

    Sync files across devices with block-level dedup, deltas and a metadata service - only the changed chunks ever cross the wire.

    3 ch

Track 03Advanced systems

Harder problems involving real-time and location-based data.

  1. Design a proximity service

    Find every business inside a radius without scanning the planet. Watch a naive lat/lng scan crawl, build a geohash bit by bit, hit the boundary problem head-on, then let a quadtree carve the map exactly where the density is.

    4 ch
  2. Design nearby friends

    Stream a moving user's location to only the friends within a 5-mile radius. Watch a peer-to-peer mesh collapse, push updates through WebSockets and Redis Pub/Sub, filter by distance, and shard the channels across a ring as load explodes.

    4 ch

Track 04GPU

The hardware every model runs on, and the limits it imposes.

  1. How a GPU runs your model

    Warps, occupancy and the memory wall, made playable. Starve an SM of registers until it has nothing to switch to, watch stalls stop being hidden, then move a kernel along the roofline and see why one user generating one token uses a fraction of a percent of the card.

    3 ch

Track 05Inference engineering

What a serving engine does between your request and the first token.

  1. Inside LLM inference

    Drive a live serving engine. Fire prompts at one GPU and watch prefill chunks, decode streams and KV blocks fight over the same step budget - then break it: turn off continuous batching, shrink the KV pool until requests get preempted, and see tokens/sec collapse.

    6 ch

Track 06Agentic AI

The loops and graphs that turn a model into something that finishes work.

  1. The agent loop

    Think, call a tool, observe, repeat. Watch an agent work a task turn by turn while its context window fills, then choose what happens when it runs out - truncate, compact, or retrieve - and watch the agent forget a result it still needed.

    4 ch
  2. Agent graphs

    Wire a multi-agent graph and run it. Fan out workers, add a critic loop, flip parallel execution on and off, and watch the Gantt chart redraw as the critical path - not the total work - decides how long the run takes.

    4 ch