Design a video platform

Upload, transcode into multiple bitrates, and stream from a CDN - adaptive bitrate that scales from one viewer to millions.

9 min read

A video platform is two systems that barely know each other: a batch pipeline that runs once per upload, and a streaming path that runs millions of times a second. Both exist because of one fact: you cannot ship the file the creator gave you.

Serve the original and almost nobody can watch it

A creator uploads a 4K master. It is enormous, it is in a format many devices cannot decode, and it assumes a network nobody on a train has. Serving it directly is not a performance compromise; it is a product that does not work.

Smart TV on fibre
80 Mbps available
plays the 50 Mbps master
Laptop on home wifi
25 Mbps available
stalls: needs 50 Mbps
Tablet in an office
6 Mbps available
stalls: needs 50 Mbps
Phone on a train
1.2 Mbps available
stalls: needs 50 Mbps
Serve only the 4K master and three of four viewers stall. Transcode once and everyone plays.

A player cannot invent a lower bitrate from a stream it receives too slowly: the bits must arrive before they can be decoded. At 4 Mbps against a stream that needs 50, the buffer fills once and then drains forever. Converting at request time for every viewer is impossible, so the conversion happens once, up front, for every rung of quality you will ever need.

Do the conversions once, in parallel, before anyone presses play

On upload the raw file goes through a pipeline: inspect it, split audio from video, encode each rendition, generate thumbnails and captions, package the results for the HLS and DASH streaming formats, and push them to the CDN. Most of those tasks do not depend on each other, so the pipeline is a graph, not a line.

UploadInspectSplit audio and videoEncode, in parallelPackage for HLS and DASHPush to CDN
One worker: 38 minutes. 4 workers: 16 minutes. That is the length of the slowest single encode, the floor no number of workers can beat.
4
Every rendition encodes from the same source, so the encodes run side by side. The stage takes as long as the slowest one.

Every rendition reads the same source and writes its own output, so nothing forces them into an order. On enough workers the encode stage takes as long as the slowest single encode rather than the sum of all of them, which is why a long upload is watchable in minutes. Transcoding is expensive and embarrassingly parallel, and it happens once per upload rather than once per view. That is what makes the economics work.

The player, not the server, decides your quality

Video is delivered in short segments of a few seconds, and every segment exists at every rung of the ladder. The player watches how fast segments arrive and how full its buffer is, and picks the quality of the next segment accordingly. That is why quality can change mid-video without anything reconnecting.

240p360p480p720p1080pstalled
No stalls: the player stepped the quality down while the connection was poor and back up afterwards.
1 Mbps
8 s
With adaptive bitrate the picture softens during the dip and nothing stops. Turn it off and the same dip becomes a spinner.

The buffer buys the player a few seconds to notice a slow connection, and adaptive bitrate spends them stepping down rather than stalling. It degrades the picture to protect continuity, because viewers forgive softness and abandon spinners. The segments themselves come from CDN edges near the viewer, so the origin serves each segment once per edge rather than once per viewer.

The short version

  • The source file is too big and too exotic to stream; transcode it into a ladder of renditions.
  • Encodes run in parallel from one source, so the stage costs the slowest encode, not the sum.
  • Video ships as short segments at every rung; the player picks each segment's quality.
  • Adaptive bitrate trades sharpness for continuity, served from CDN edges.