Designing Dropbox / Cloud File Storage and Sync
Difficulty: Advanced Topics: Chunking, Deduplication, Delta Sync, Conflict Resolution, Object Storage, Metadata Service Asked at: Google, Amazon, Microsoft, Dropbox, PhonePe, Flipkart Prerequisites:Object Storage, Consistent Hashing, and Message Queues
1. Understanding the Problem
A cloud file storage service lets users upload files, sync them across devices, share them with others, and access version history. The hard engineering problems: syncing a 4GB video edit without re-uploading the entire file (delta sync), handling two people editing the same file offline simultaneously (conflict resolution), and storing petabytes of files cost-efficiently while keeping metadata lookups fast.
Real examples: Dropbox, Google Drive, OneDrive, iCloud Drive, Box.
2. Naive First Cut
flowchart LR
CLIENT["Desktop Client"]:::client
API["API Server"]:::service
STORE[("Single File Server<br/>stores full files")]:::data
CLIENT -->|"Upload entire file"| API
API --> STORE
classDef client fill:#4c3a5e,stroke:#818cf8,color:#e2e8f0
classDef service fill:#1a3a2a,stroke:#4ade80,color:#e2e8f0
classDef data fill:#3b3520,stroke:#fbbf24,color:#e2e8f0
User uploads the entire file to a server; server stores it as-is on disk. On change, re-upload the whole file.
Why this breaks:
- Re-uploading a 2GB file because one byte changed wastes bandwidth and takes minutes
- No deduplication - 100 users with the same PDF = 100 copies stored
- Single file server has finite disk; canβt scale to petabytes
- No versioning - overwriting loses history
- No sync - other devices donβt know a file changed until they poll
- Concurrent edits from two devices silently overwrite each other (last-write-wins data loss)
The rest of the doc evolves this into a chunked, content-addressed, event-driven sync system.
3. Prior Art Weβre Drawing From
- Dropbox Block-Level Sync - Files are split into 4MB chunks, each identified by SHA-256 hash. Only changed chunks are uploaded/downloaded. This single optimization reduced Dropboxβs bandwidth usage by 60%+ for typical edit patterns. (Dropbox Tech Blog)
- Google Drive Conflict Resolution - Uses operational-transform-style conflict detection: when two clients edit the same file offline, the second to sync gets a βconflict copyβ rather than silent overwrite. User explicitly resolves. (Google Workspace Blog)
- Rsync Rolling Checksum - The rsync algorithm uses a rolling hash (Adler-32) to detect which blocks of a file have changed without comparing the entire file byte-by-byte. Dropboxβs delta sync is conceptually derived from this. (Rsync Technical Report)
- Content-Addressable Storage (CAS) - Git, IPFS, and Dropbox all use content-addressed storage (hash of content = storage key). This gives deduplication for free β identical blocks map to the same key regardless of which user uploaded them.
4. Functional Requirements
Core (Top 3)
- Upload and download files - users can store files up to 50GB and retrieve them from any device
- Sync across devices - changes on one device propagate to all other devices within seconds
- Version history - users can view and restore previous versions of any file (last 30 days)
Below the Line
- File and folder sharing with permissions (view/edit)
- Collaborative real-time editing (Google Docs territory)
- Offline editing with eventual sync
- Storage quota management
- Trash / soft-delete with recovery
5. Non-Functional Requirements
Core
- Reliability: Zero data loss - files stored with 99.999999999% (11 nines) durability
- Sync latency: Changes propagate to other devices within 5-10 seconds (after upload completes)
- Upload efficiency: Only transfer changed bytes, not entire files (delta sync)
- Scale: 500M files per workspace, 100K concurrent sync sessions globally
Below the Line
- Eventual consistency for metadata across regions (but strong within a region)
- Bandwidth-aware sync (pause on metered connections)
- Cost-efficient storage tiering (hot/warm/cold)
6. Core Entities
- Workspace - a userβs or teamβs logical container for all files (the root of their file tree)
- FileMetadata - path, size, content hash, chunk list, version history, permissions
- Chunk - a fixed-size (4MB) block of a file, identified by its SHA-256 hash (content-addressed)
- Version - a snapshot of a fileβs chunk list at a point in time
- SyncEvent - a notification that a file changed (created/modified/deleted/moved)
- EditSession - tracks an active deviceβs sync state (cursor position in the event stream)
7. API / System Interface
POST /v1/files/upload/init
Body: {"path": "/docs/report.pdf", "size": 52428800, "chunks": ["sha256_a", "sha256_b", ...]}
Response: {"upload_id": "u123", "chunks_needed": ["sha256_b"]}
// Server already has sha256_a (dedup) - only upload the new chunk
PUT /v1/files/upload/{upload_id}/chunk/{chunk_hash}
Body: <binary chunk data>
Response: 200 OK
POST /v1/files/upload/{upload_id}/complete
Response: {"file_id": "f456", "version": 3}
GET /v1/files/{file_id}/download?version=latest
Response: 302 Redirect to signed S3 URL (or chunked download URLs)
GET /v1/sync/events?cursor={last_event_id}&limit=100
Response: {"events": [...], "cursor": "evt_789", "has_more": false}
// Long-poll: server holds connection open until new events arrive or timeout (30s)
Security notes: all chunk uploads go to pre-signed URLs (client uploads directly to object store, not through API server). Download URLs are time-limited (15 min). Access checks happen at the metadata layer before issuing signed URLs.
8. High-Level Design
FR1: Upload and download files
Read the requirement plainly: a file the user put on one machine can be got back from another. That needs the bytes kept somewhere durable, and a record of what the file is called and where it sits in their tree.
Build that and nothing else. The bytes go to object storage as one object; the name, path and version go in a row in Postgres. No chunking, no hashing, no dedup check. Every one of those answers a non-functional requirement β the requirements list βonly transfer changed bytesβ under non-functional, which is precisely the signal that it is a deep dive and not part of the base design.
flowchart LR
CLIENT["Desktop Client"]:::client
API["API Gateway"]:::edge
META["Metadata Service"]:::service
OBJSTORE[("Object Storage<br/>one object per version")]:::data
METADB[("Postgres<br/>files and versions")]:::data
CLIENT -->|"1. Request an upload slot"| API
API -->|"2. Reserve a version row"| META
META -->|"3. Write path and object key"| METADB
CLIENT -->|"4. PUT the whole file"| OBJSTORE
CLIENT -->|"5. Mark upload complete"| API
classDef client fill:#4c3a5e,stroke:#818cf8,color:#e2e8f0
classDef edge fill:#1e3a5f,stroke:#38bdf8,color:#e2e8f0
classDef service fill:#1a3a2a,stroke:#4ade80,color:#e2e8f0
classDef data fill:#3b3520,stroke:#fbbf24,color:#e2e8f0
| Color | Meaning |
|---|---|
| Purple | Client |
| Blue | Edge / Gateway |
| Green | Application Service |
| Yellow | Data Store |
New components:
- Metadata Service: Owns the file tree. Given a path it can say which object key holds the current bytes, and it records a new version row on every upload.
- Object Storage: Holds the file bytes, one object per uploaded version, keyed by
workspace_id/file_id/version. Immutable once written, which is what makes 11-nines durability somebody elseβs problem. - Metadata DB (Postgres): Paths, versions, permissions. Sharded by
workspace_id, because no query in this design ever needs to cross workspaces.
Flow:
- Client asks to upload
~/report.pdf; the Metadata Service allocates the next version number and an object key for it - It writes the row with
status = uploading, so a crashed upload leaves an obvious orphan rather than a corrupt file - Client PUTs the entire file to that object key with a pre-signed URL, so the bytes never pass through our servers
- Client calls complete; the Metadata Service flips the row to
status = currentand demotes the previous version - Download reverses it: read the row for the path, hand back a pre-signed GET for that object key
Why send bytes straight to object storage instead of through the API? A 50GB file through our gateway means paying for that bandwidth twice and holding a connection open for hours on a process that also has to serve everyone else. Pre-signed URLs let the client talk to storage directly while the Metadata Service stays a small, fast, purely-metadata service.
What we have deliberately left broken. This stores and retrieves files correctly, which is what the requirement asked. Three things about it are indefensible at scale, and all three are non-functional:
- Every edit re-uploads everything. Change one byte of a 2GB video and we transfer 2GB. On a home connection that is hours, and it happens again on the next save. That is Deep Dive 1.
- Identical files are stored once per user. A 50MB installer that a thousand people upload is stored a thousand times, 50GB, for one fileβs worth of information. Nothing here notices the bytes are the same. That is Deep Dive 2.
- Two devices uploading at once, one wins silently. Step 4 has no notion of what version the client started from, so the later complete simply demotes the earlier one and that userβs work is gone with no error and no copy. That is Deep Dive 4.
FR2: Sync across devices
The other devices have to find out something changed. The requirement is βwithin secondsβ, and the target in the non-functional list is 5-10 seconds, which is loose enough that the simplest mechanism clears it: let each device ask.
New components we need: none. The version numbers FR1 is already writing are enough to answer βwhat changed since I last looked?β, so this is one more read on the Metadata Service rather than new infrastructure.
flowchart LR
CLIENT["Device A"]:::client
API["API Gateway"]:::edge
META["Metadata Service"]:::service
METADB[("Postgres<br/>files and versions")]:::data
OBJSTORE[("Object Storage")]:::data
CLIENT2["Device B"]:::client
CLIENT -->|"1. Upload completes"| API
API -->|"2. Bump workspace version"| META
META -->|"3. Write new version row"| METADB
CLIENT2 -->|"4. Poll changes since cursor"| META
META -->|"5. Rows newer than cursor"| CLIENT2
CLIENT2 -->|"6. Download the file"| OBJSTORE
classDef client fill:#4c3a5e,stroke:#818cf8,color:#e2e8f0
classDef edge fill:#1e3a5f,stroke:#38bdf8,color:#e2e8f0
classDef service fill:#1a3a2a,stroke:#4ade80,color:#e2e8f0
classDef data fill:#3b3520,stroke:#fbbf24,color:#e2e8f0
Flow:
- Device A finishes an upload as in FR1, which writes a version row
- Each workspace carries a monotonically increasing counter, bumped on every committed change
- Device B polls
GET /changes?since=<its last counter>every 5 seconds - The Metadata Service returns the rows above that counter β usually none
- For each change, Device B downloads the file and advances its stored cursor
Why a cursor rather than a timestamp? Clocks on client machines are wrong, sometimes by minutes, and two changes can share a millisecond. A counter the server owns gives every device an unambiguous βI have seen everything up to hereβ, which is also exactly what makes a device that has been offline for a week recoverable: it asks from an old number and gets the backlog.
What we have deliberately left broken. Polling satisfies the requirement, and it is wasteful in two specific ways:
- Almost every request is a wasted round trip. 100K concurrent sync sessions polling every 5 seconds is 20K requests/second, and in a quiet workspace essentially all of them answer βnothingβ. We are paying for a constant load whose whole purpose is to discover that nothing happened.
- It puts a floor under the latency we promised. A 5-second poll interval means average detection lag of 2.5 seconds and worst case 5, before any download starts, against a 5-10 second end-to-end target. There is no headroom left for a slow download, and tightening the interval multiplies the request rate above.
Both are Deep Dive 3, which is also where a device that missed events while offline gets a protocol rather than a hopeful cursor.
FR3: Version history
This one comes almost free, because of a decision already made in FR1: each upload writes a new object rather than overwriting the old one. The history is therefore already on disk. All that is missing is a way to list it and a way to go back.
New components we need:
- Retention Job: the requirement is a 30-day window, so something has to delete what falls outside it. A daily pass finds versions older than 30 days that are not the current one and deletes their objects.
flowchart LR
CLIENT["Desktop Client"]:::client
META["Metadata Service"]:::service
METADB[("Postgres<br/>versions table")]:::data
OBJSTORE[("Object Storage<br/>one object per version")]:::data
GC["Retention Job<br/>daily"]:::async
CLIENT -->|"1. List versions for a path"| META
META -->|"2. Read version rows"| METADB
CLIENT -->|"3. Restore version 3"| META
META -->|"4. Add v5 pointing at v3 object"| METADB
GC -->|"5. Delete objects past 30 days"| OBJSTORE
classDef client fill:#4c3a5e,stroke:#818cf8,color:#e2e8f0
classDef service fill:#1a3a2a,stroke:#4ade80,color:#e2e8f0
classDef data fill:#3b3520,stroke:#fbbf24,color:#e2e8f0
classDef async fill:#3a2a4c,stroke:#c084fc,color:#e2e8f0
Flow:
- Client asks for the history of a path; the Metadata Service returns the version rows with timestamps and sizes
- Client restores version 3
- The Metadata Service does not move any bytes. It writes a new row, v5, whose object key is the one v3 already points at, and marks it current
- Restore is therefore a single metadata write and is instant regardless of file size
- The Retention Job deletes objects belonging to versions that have aged out, skipping any key still referenced by a live row
Why create a new version on restore instead of rolling back to v3? Because rolling back destroys v4, and the user who restored may well have wanted it. Making restore an append-only operation means every action in this system is recoverable from the version log, including the mistakes.
What we have deliberately left broken. Correct, and the storage bill is absurd:
- Every version is a full copy. A 2GB video saved ten times is 20GB, even if each save changed a title frame. This is the same wound as Deep Dive 1 seen from the storage side rather than the network side, and Deep Dives 1 and 2 together are what shrink it.
- Everything sits in the most expensive storage class forever. A version from 29 days ago that nobody will ever open costs exactly what the file someone is editing right now costs. Nothing here distinguishes them. That is Deep Dive 5.
9. Technology Choices
| Tier | Purpose | Stores | Access Pattern | Primary Pick | Alternatives |
|---|---|---|---|---|---|
| Object store | File chunk storage | Raw binary chunks (4MB each) | PUT/GET by content hash | S3 / GCS / Azure Blob | MinIO (self-hosted) / Ceph |
| Metadata DB | File tree and versions | File paths, chunk lists, versions, sharing | OLTP reads and writes | Postgres (strong consistency) | CockroachDB / TiDB / Spanner |
| Sync queue | Change notifications | File change events | Pub/Sub per user | Kafka / Redis Streams | RabbitMQ / SNS+SQS |
| Cache | Hot metadata | Recently accessed file metadata | Key-value lookup | Redis / Memcached | - |
| CDN | Download acceleration | Popular shared files | Read-heavy, geo-distributed | CloudFront / Cloudflare | Fastly / Akamai |
| Block diff engine | Delta computation | Rolling hash state | CPU-intensive, stateless | Custom service (rsync-like) | librsync library |
Why Postgres for metadata, not a NoSQL store? File systems have strong hierarchy and referential constraints (folders contain files, files have ordered chunk lists, sharing permissions reference users). Relational integrity + ACID transactions prevent orphaned chunks and inconsistent states. At Dropbox-scale, shard Postgres by workspace_id.
10. Data Modeling
Postgres (Metadata DB β file tree and versions, sharded by workspace_id):
CREATE TABLE files (
file_id UUID PRIMARY KEY,
workspace_id UUID NOT NULL,
parent_folder_id UUID,
name VARCHAR(500) NOT NULL,
size_bytes BIGINT,
content_hash VARCHAR(64), -- SHA-256 of full file
version INTEGER NOT NULL DEFAULT 1,
is_deleted BOOLEAN DEFAULT false,
created_at TIMESTAMP NOT NULL,
updated_at TIMESTAMP NOT NULL
);
CREATE INDEX idx_files_parent ON files(workspace_id, parent_folder_id, name);
CREATE TABLE file_chunks (
file_id UUID NOT NULL,
chunk_index INTEGER NOT NULL,
chunk_hash VARCHAR(64) NOT NULL, -- SHA-256, also the S3 key
chunk_size INTEGER NOT NULL,
PRIMARY KEY (file_id, chunk_index)
);
S3 (Object Store β content-addressed chunks):
Key: "chunks/{sha256_hash}" β Raw binary (4MB chunk)
Deduplication: identical chunks across all users share the same S3 object
Durability: 11 nines (99.999999999%)
Kafka / Redis Streams (Sync Queue β change notifications):
Topic: file-changes (partitioned by workspace_id)
Key: workspace_id
Value: { file_id, event_type (created/modified/deleted/moved), version, chunk_list, timestamp }
Per-user cursor: each device tracks last_synced_version per workspace
Access Patterns:
| Query | Data Source | How |
|---|---|---|
| List folder contents | Postgres | SELECT * FROM files WHERE workspace_id = ? AND parent_folder_id = ? AND is_deleted = false |
| Upload file (chunked) | S3 + Postgres | Upload each chunk to S3 by hash β write file_chunks rows β update files metadata |
| Download file | Postgres β S3 | Read chunk list from file_chunks β parallel GET from S3 β reassemble |
| Sync changes to other devices | Kafka | Device subscribes to file-changes topic, pulls events after its cursor |
| Deduplicate chunks | S3 key check | Before upload: HEAD chunks/{hash} β if exists, skip upload (chunk already stored) |
How Delta Sync Minimizes Upload Bandwidth:
- User modifies a 100MB file β client splits file into 4MB chunks (25 chunks)
- For each chunk: compute SHA-256 hash. Compare with stored chunk list from metadata
- Only 2 chunks changed (positions 5 and 12) β upload only 8MB instead of 100MB
PUT chunks/{new_hash_5}andPUT chunks/{new_hash_12}to S3- Update
file_chunksin Postgres: replace chunk_hash at indices 5 and 12 - Increment file version, publish
file-changesevent to Kafka - Other devices receive the event, download only the 2 changed chunks, patch their local copy
11. Deep Dives
1) How do we stop a one-byte edit to a 2GB video re-uploading all 2GB?
Bad: What FR1 built β one object per version, so every save transfers the whole file. A 2GB video with a corrected title frame costs 2GB on the wire, which on a 20Mbps home upload is a little over two hours during which the userβs other work is competing for the same pipe. Autosave makes it pathological: ten saves in an evening is 20GB uploaded to communicate a few kilobytes of actual change. The wasted bytes are not an efficiency footnote either, because FR2βs sync means every other device then downloads that same 2GB.
Good: Fixed-size chunking (4MB blocks). Only changed chunks are uploaded. If the user edits byte 5,000,000, only the chunk containing that byte (chunk 2) gets re-uploaded. Savings: 99.8% for a single edit to a large file.
Great: Content-defined chunking (CDC) using a rolling hash (Rabin fingerprint). Instead of fixed 4MB boundaries, chunk boundaries are determined by the content itself. This means inserting bytes at the start of a file doesnβt shift ALL chunk boundaries (which fixed-size would) β only the chunks near the insertion point change. Dropbox uses this to reduce unnecessary re-uploads from ~40% to ~5% for insert-heavy workloads (documents, code files).
2) A thousand people upload the same 50MB installer. How do we store it once?
Bad: What FR1 built, and Deep Dive 1 does not fix it. Chunking makes uploads smaller but still stores whatever arrives: object keys are assigned per file version, so two identical files land under two different keys and nothing ever compares their contents. A 50MB installer that a thousand people sync is 50GB for one fileβs worth of information. The same pattern repeats inside a single account β a folder duplicated as a backup doubles instantly β and it repeats across versions, because Deep Dive 1βs chunk boundaries mean an unchanged chunk is re-stored under a new versionβs key rather than recognised as the chunk we already hold.
Good: Content-addressed storage β chunk hash IS the storage key. Before uploading, check if the hash exists. Global dedup across all users. Dropbox reported 60%+ storage savings from cross-user dedup.
Great: Add a Bloom filter in front of the dedup check. With 10B chunks, checking existence in Postgres for every chunk hash on every upload is expensive. A Bloom filter (in memory, ~10 bytes per entry = ~100GB for 10B chunks) gives a fast βdefinitely not storedβ answer for new chunks, and only hits the DB for probable matches. False positive rate of 0.1% means 99.9% of DB lookups are eliminated.
3) 20K polls a second just to learn that nothing changed. What replaces them?
Bad: What FR2 built β every device asks every 5 seconds. At the 100K concurrent sync sessions in the requirements that is 20K requests/second, and in a workspace where nobody is currently editing, close to all of them return an empty list. That is a permanent floor of load whose entire function is to discover the absence of news, and it scales with the number of idle users rather than with activity. The latency arithmetic is the sharper problem: a 5-second interval means 2.5 seconds of average detection lag and 5 in the worst case, before the download even begins, against a 5-10 second end-to-end promise. Halving the interval to buy headroom doubles the request rate.
Good: Long-polling β client holds a connection open for 30 seconds. Server responds immediately when an event occurs, or times out with βno changes.β 99% fewer requests than polling.
Great: WebSocket server-push with cursor-based catch-up. Holding the socket open means a change reaches the other device in the time it takes to write a frame, which retires the 5-second floor entirely. Two pieces make it work at scale:
- A pub/sub hop between the writer and the socket holder. The Metadata Service has no idea which server is holding a given deviceβs connection, and should not have to. It publishes the change to an event bus (Kafka / Redis Streams / Pub/Sub) keyed by workspace; the Sync Service instances subscribe and push to whichever devices they happen to be holding. This is also what lets the two scale independently β connection count drives Sync Service capacity, write rate drives Metadata Service capacity, and neither has to care about the other.
- The cursor survives disconnection. On reconnect the client sends the last event id it processed and the server replays forward from there, so a device that was offline for a week is the same code path as one that blinked. Nothing is missed and nothing is applied twice.
Add exponential backoff with jitter on reconnect. When a Sync Service instance restarts, every device it was holding reconnects at once; with 100K concurrent sessions spread over a modest fleet that is a large simultaneous burst against whichever instances are still up, and without jitter the retries stay synchronised and keep arriving in waves.
4) Two devices edit the same file offline. Who wins, and does anyone find out?
Bad: What FR1 built β last write wins, and it wins silently. Step 4 of that flow commits a version without ever asking which version the client started from, so if a laptop and a desktop both edit a document offline and then reconnect, the one that happens to call complete second becomes current and the other personβs afternoon is simply gone. Nothing errors, nothing is flagged, and the losing device cheerfully downloads the winner over its own work, so the local copy that held the lost edits is destroyed too. The requirements open with zero data loss and 11 nines of durability; storing the bytes perfectly and then overwriting them on a race is still data loss, just ours rather than the diskβs.
Good: Optimistic concurrency with version checks (as shown in Core Flows). Detect conflicts and create conflict copies. User resolves manually.
Great: For collaborative scenarios, use vector clocks or Lamport timestamps to track causality. If two edits are causally independent (neither βsawβ the other), itβs a true conflict. If one causally follows the other (Device B saw Aβs version before editing), itβs a clean update. This reduces false conflicts β only truly divergent edits require user resolution. Google Drive uses a version-vector approach internally to minimize unnecessary conflict copies.
5) Why does a version nobody will ever open cost the same as the file being edited?
Bad: What FR3 built β every version sits in the standard storage class until the retention job deletes it, and access frequency never enters the decision. At S3 Standardβs $0.023 per GB per month, 10PB is about $241K a month, so roughly $2.9M a year, and Deep Dive 1βs version history is actively inflating that number since each retained version is charged in full. The waste is heavily skewed: the overwhelming majority of those bytes are versions from three weeks ago that will never be read again, and they are billed at the rate we are paying for the file somebody has open right now.
Good: Move versions older than 30 days to S3 Infrequent Access ($0.0125/GB). Move versions older than 90 days to S3 Glacier ($0.004/GB). 3-5x cost reduction for archival data.
Great: Intelligent tiering based on access patterns, not just age. Track per-chunk access frequency. A 2-year-old chunk thatβs part of an actively-used shared folder stays in Standard. A 1-week-old chunk from a file nobodyβs opened since upload moves to IA immediately. Combine with regional replication only for active workspaces β archive workspaces replicate to one region only.
12. Design Self-Audit
- Stale reads after writes? No β Metadata Service uses Postgres with read-after-write consistency within the same workspace shard. Other devices get notified via event stream within seconds.
- Single points of failure? Metadata DB is the critical path β mitigated with Postgres streaming replication (sync replica for zero data loss). Object store (S3) has built-in 11-nines durability. Kafka is multi-broker.
- Dead-letter / reconciliation? If Sync Service fails to deliver an event, the clientβs next long-poll catch-up (cursor-based) will replay missed events. No silent data loss.
- Cost? The biggest cost is object storage. Dedup + tiering reduce it by 70-80% vs naive storage. At 1PB, expect ~$15K-25K/month (vs $100K+ without optimization).
- Hot partition? Workspaces with very large shared folders (10K+ files, many editors) could hot-spot the metadata shard. Mitigated by folder-level read caching and async event batching.
13. Core Flows
Flow 1: File Upload with Delta Sync
sequenceDiagram
participant C as Desktop Client
participant A as API Gateway
participant M as Metadata Service
participant DB as Metadata DB
participant S3 as Object Store
participant K as Kafka
participant SS as Sync Service
participant C2 as Other Device
C->>C: Detect file change via filesystem watcher
C->>C: Chunk file into 4MB blocks and compute SHA-256 per chunk
C->>A: POST /upload/init with chunk hash list
A->>M: Check which chunks exist
M->>DB: SELECT existing chunk hashes
DB-->>M: chunks_existing = [hash_1, hash_3]
M-->>A: chunks_needed = [hash_2] (only the new one)
A-->>C: Upload only hash_2
C->>S3: PUT chunk hash_2 via pre-signed URL
S3-->>C: 200 OK
C->>A: POST /upload/complete
A->>M: Create new version
M->>DB: INSERT version record
M->>K: Publish SyncEvent
K->>SS: Consume event
SS->>C2: Push via long-poll
C2->>S3: Download chunk hash_2
C2->>C2: Reconstruct file from local cache + new chunk
- Filesystem watcher detects the change instantly (no polling)
- Client-side chunking means the expensive hashing happens locally, not on the server
- Dedup check at init means only genuinely new data traverses the network
- Pre-signed URL upload bypasses the API server (direct client-to-S3)
- Version creation is atomic in Postgres (transaction)
- Sync notification is fire-and-forget from Metadata Serviceβs perspective
Non-obvious failure path: If the client crashes mid-upload (some chunks uploaded, complete never called), the upload_id times out after 24 hours. Orphaned chunks in S3 are cleaned by the Garbage Collector (theyβre unreferenced by any version).
Flow 2: Conflict Resolution (Two Offline Edits)
sequenceDiagram
participant A as Device A (offline)
participant B as Device B (offline)
participant M as Metadata Service
participant DB as Metadata DB
Note over A,B: Both devices edit the same file while offline
A->>A: Edit file and queue local version (base = v3)
B->>B: Edit file and queue local version (base = v3)
Note over A,B: Both come online
A->>M: Upload complete (base_version = v3)
M->>DB: v3 is current -> accept as v4
M-->>A: Success - you are v4
B->>M: Upload complete (base_version = v3)
M->>DB: Current is v4 not v3 -> CONFLICT
M-->>B: Conflict detected
B->>B: Save as "report (conflict copy - Device B).pdf"
B->>M: Upload conflict copy as separate file
- Each client tracks the base_version it last synced
- On upload-complete, Metadata Service checks: is base_version still current?
- If yes, accept as next version (optimistic concurrency)
- If no, reject with CONFLICT β the client creates a conflict copy
- User manually resolves by picking one or merging
Non-obvious failure: Network partition during upload-complete. Client retries with idempotency key β Metadata Service deduplicates and never creates duplicate versions for the same upload.
14. Final Architecture
flowchart TD
CLIENT["Desktop and Mobile Clients"]:::client
CDN["CDN<br/>(download acceleration)"]:::edge
API["API Gateway"]:::edge
META["Metadata Service"]:::service
SYNC["Sync Service<br/>(long-poll and WebSocket)"]:::service
METADB[("Metadata DB<br/>Postgres sharded")]:::data
OBJSTORE[("Object Store<br/>S3 with tiering")]:::data
QUEUE["Kafka<br/>(sync events)"]:::async
CACHE["Redis<br/>(metadata cache)"]:::data
GC["Garbage Collector"]:::async
BLOOM["Bloom Filter<br/>(dedup check)"]:::service
CLIENT -->|"Upload or sync file"| API
CLIENT -->|"Download via CDN"| CDN
CDN -->|"Fetch origin"| OBJSTORE
API -->|"Write file metadata"| META
META -->|"Read chunk tree"| METADB
META -->|"Lookup cached metadata"| CACHE
META -->|"Dedup check"| BLOOM
BLOOM -->|"Store new chunks"| OBJSTORE
CLIENT -->|"Upload chunk data"| OBJSTORE
META -->|"Publish sync event"| QUEUE
QUEUE -->|"Notify synced devices"| SYNC
SYNC -->|"Push to other devices"| CLIENT
GC -->|"Delete orphan chunks"| OBJSTORE
GC -->|"Remove stale metadata"| METADB
classDef client fill:#4c3a5e,stroke:#818cf8,color:#e2e8f0
classDef edge fill:#1e3a5f,stroke:#38bdf8,color:#e2e8f0
classDef service fill:#1a3a2a,stroke:#4ade80,color:#e2e8f0
classDef data fill:#3b3520,stroke:#fbbf24,color:#e2e8f0
classDef async fill:#3a2a4c,stroke:#c084fc,color:#e2e8f0
How it works end-to-end:
- Client uploads file β chunked upload sent directly to Object Store (S3); metadata request goes to API Gateway
- Metadata Service processes β checks Redis cache, queries Postgres for file tree, runs Bloom Filter dedup check
- Dedup result β if content hash already exists in S3, skip upload and just link metadata; otherwise persist new chunks
- Sync event emitted β Metadata Service publishes change event to Kafka
- Sync Service notifies other devices β consumes from Kafka, pushes update via long-poll or WebSocket to the clientβs other devices
- Other devices download β fetch changed chunks from CDN (cache hit) or S3 origin (cache miss)
- Garbage Collector reclaims storage β periodically scans for orphaned chunks (unreferenced after delete/overwrite) and removes from S3 and Metadata DB
Discussion
Newest first