Design a file store

Sync files across devices with block-level dedup, deltas and a metadata service - only the changed chunks ever cross the wire.

9 min read

A file store like Google Drive or Dropbox looks like a folder that happens to be everywhere. Underneath it is three separate problems: moving bytes efficiently, knowing which bytes changed, and deciding whose version is the real one when two devices disagree.

You changed one line. It sent the whole file

The simplest sync client uploads the file whenever it changes. That is fine for a document you touch twice a day and ruinous for one you are actively working in, because editors save far more often than you think.

Whole file every save
3.75 GB
Only the changed 4 MB block
480 MB
What actually changed
240 KB
Log scale
120 autosaves in 1 hour. Whole-file sync sends 3.75 GB to express 240 KB of change.
32 MB
30 s
1
An hour on a 32 MB file with 30-second autosaves: 3.75 GB uploaded to express 240 KB of change.

The client cannot tell “the file changed” from “this part of the file changed”, so it re-sends the only unit it understands: all of it. Whole-file sync scales with how often you save, not how much you change. Every improvement from here comes from making the unit of transfer smaller than the file.

Split the file, hash each piece, send only what moved

Drive-style storage splits every file into fixed 4 MB blocks, each named by a hash of its contents. On save, the client compares hashes with the server and uploads only the blocks that differ. That is delta sync.

2 of 8 blocks have a new hash. The next sync sends 8 MB of 32 MB.
Two blocks changed, so the next sync is 8 MB, not 32. Edit every block and delta sync saves nothing.

The block is the unit, so a few changed bytes still cost a whole 4 MB block. Smaller blocks waste less per edit but mean more hashes to track and more round trips; Dropbox settled on 4 MB. Content-addressed blocks pay off three ways at once: delta sync, deduplication (two files sharing a block store it once, across every user), and cheap versioning (a version is just a list of block hashes). Blocks are compressed and encrypted before upload, with the hash taken over the plaintext so deduplication still works.

Two devices edited the same file. Now what?

Blocks solve bandwidth. They do not solve disagreement. The metadata database, not the block store, is the source of truth for which version is current, and it decides with nothing more sophisticated than a version number: every update says which version it was based on.

Metadata service: the file is at v2, “budget: 45k”
Laptopbased on v2, current
Phonebased on v1, stale
Laptop pushed v2, based on v1: accepted.
The laptop already pushed v2. The phone's edit is based on v1. Push it and see what the server does.

The phone says “here is my update, based on version 1”, but the server is already at version 2. It cannot merge arbitrary binary content, and it will not silently throw away either edit, so it does the only safe thing: keep the first write and hand the second back as a conflicted copy for a person to resolve. Real-time editors like Google Docs avoid this by merging individual operations, but a general file store cannot know what the bytes mean.

The short version

  • Whole-file sync costs the file size on every autosave.
  • Split files into hashed 4 MB blocks and upload only the blocks whose hash changed.
  • Content hashes also give deduplication and cheap versions.
  • The metadata service decides conflicts by version number; losers become conflicted copies.