Design a file store
Sync files across devices with block-level dedup, deltas and a metadata service - only the changed chunks ever cross the wire.
9 min read
A file store like Google Drive or Dropbox looks like a folder that happens to be everywhere. Underneath it is three separate problems: moving bytes efficiently, knowing which bytes changed, and deciding whose version is the real one when two devices disagree.
You changed one line. It sent the whole file
The simplest sync client uploads the file whenever it changes. That is fine for a document you touch twice a day and ruinous for one you are actively working in, because editors save far more often than you think.
The client cannot tell “the file changed” from “this part of the file changed”, so it re-sends the only unit it understands: all of it. Whole-file sync scales with how often you save, not how much you change. Every improvement from here comes from making the unit of transfer smaller than the file.
Split the file, hash each piece, send only what moved
Drive-style storage splits every file into fixed 4 MB blocks, each named by a hash of its contents. On save, the client compares hashes with the server and uploads only the blocks that differ. That is delta sync.
The block is the unit, so a few changed bytes still cost a whole 4 MB block. Smaller blocks waste less per edit but mean more hashes to track and more round trips; Dropbox settled on 4 MB. Content-addressed blocks pay off three ways at once: delta sync, deduplication (two files sharing a block store it once, across every user), and cheap versioning (a version is just a list of block hashes). Blocks are compressed and encrypted before upload, with the hash taken over the plaintext so deduplication still works.
Two devices edited the same file. Now what?
Blocks solve bandwidth. They do not solve disagreement. The metadata database, not the block store, is the source of truth for which version is current, and it decides with nothing more sophisticated than a version number: every update says which version it was based on.
The phone says “here is my update, based on version 1”, but the server is already at version 2. It cannot merge arbitrary binary content, and it will not silently throw away either edit, so it does the only safe thing: keep the first write and hand the second back as a conflicted copy for a person to resolve. Real-time editors like Google Docs avoid this by merging individual operations, but a general file store cannot know what the bytes mean.
The short version
- Whole-file sync costs the file size on every autosave.
- Split files into hashed 4 MB blocks and upload only the blocks whose hash changed.
- Content hashes also give deduplication and cheap versions.
- The metadata service decides conflicts by version number; losers become conflicted copies.