Limited time: AI code review, hints, mock interviews, whiteboard analysis, and all Pro features are unlocked. Enroll
⏱️ 10 min read

Distributed Locking - Complete Deep Dive

Prerequisites: Database Concepts, Caching Used in: Digital Wallet, BookMyShow, Job Scheduler, any system with shared mutable state


What is Distributed Locking?

A distributed lock coordinates multiple servers so that, in the normal case, only ONE of them acts on a shared resource at a time.

Read that β€œin the normal case” carefully, because it is the whole subject. What a lock with a TTL actually hands you is a lease β€” exclusivity for a bounded period. If the holder stalls longer than the TTL (a GC pause, a lost packet, a slow disk), the lease expires, someone else acquires it, and the original holder is still running and still convinced it holds the lock. Nothing about the lock itself stops it from writing. That is why a lock is not, on its own, a correctness guarantee for anything that must happen once β€” see Problem 2: Lock Expires While Still Working below, and treat the lock as a way to make conflicts rare while the store that owns the data enforces the invariant.

Real-world analogy: A bathroom with a lock. One person enters, locks it. Others wait. When done, they unlock, and the next person enters. In distributed systems, the β€œbathroom” is a shared resource (database row, file, external API) and the β€œpeople” are different servers. The analogy breaks in the way that matters: a real bolt cannot un-bolt itself after 30 seconds while you are still inside.

Why it’s needed:

Without lock:
  Server A: reads balance=100, deducts 80 β†’ writes 20
  Server B: reads balance=100 (same time!), deducts 80 β†’ writes 20
  Expected final: 100 - 80 - 80 = -60 (impossible!) or 20 (lost update)

With lock:
  Server A: acquires lock β†’ reads 100 β†’ deducts 80 β†’ writes 20 β†’ releases lock
  Server B: tries lock β†’ WAITS β†’ acquires lock β†’ reads 20 β†’ deducts 80 β†’ FAILS (insufficient)
  Correct!

When Do You Need Distributed Locks?

Scenario Why
Prevent double-booking Two users try to book the last seat at the same time
Prevent double-spending Two transfers from the same wallet concurrently
Exactly-once job execution Cron job running on 5 servers β€” only one should execute
Rate limit writes to external API Only one server should call the API at a time
Resource ownership Only one worker processes a specific task

Approaches

1. Redis SETNX (Simplest, Most Common)

Use Redis SET key value NX EX ttl β€” sets the key only if it doesn’t exist (NX), with expiry (EX).

ACQUIRE LOCK:
  SET "lock:order_123" "server_A_uuid" NX EX 30
  β†’ returns OK if acquired (key didn't exist)
  β†’ returns nil if someone else holds it

RELEASE LOCK (only if you own it):
  if GET "lock:order_123" == "server_A_uuid":
      DEL "lock:order_123"

Why UUID as value? To ensure only the lock OWNER can release it. Without it, Server B could accidentally delete Server A’s lock.

Why TTL (EX 30)? If the lock holder crashes, the lock auto-expires after 30s. Without TTL, a crashed holder keeps the lock forever (deadlock).

// Java implementation
public boolean acquireLock(String resource, String owner, int ttlSeconds) {
    String key = "lock:" + resource;
    String result = redis.set(key, owner, SetParams.setParams().nx().ex(ttlSeconds));
    return "OK".equals(result);
}

public boolean releaseLock(String resource, String owner) {
    String key = "lock:" + resource;
    // Lua script for atomic check-and-delete
    String lua = "if redis.call('get', KEYS[1]) == ARGV[1] then " +
                 "return redis.call('del', KEYS[1]) else return 0 end";
    return redis.eval(lua, List.of(key), List.of(owner)) == 1;
}

Why Lua for release? GET and DEL must be atomic. Without Lua:

Server A: GET lock β†’ "server_A" (correct owner)
  --- A's TTL expires here ---
Server B: SET lock β†’ "server_B" (acquires lock)
Server A: DEL lock β†’ deletes B's lock!  ← BUG

Pros: Simple, fast (~1ms), widely used. Cons: Single Redis instance is a SPOF. If Redis crashes, lock is gone.


2. Redlock (Multi-Node Redis)

For higher reliability, acquire locks on a MAJORITY of Redis nodes (e.g., 3 out of 5).

5 independent Redis instances.

ACQUIRE:
  Try SET NX EX on all 5 instances
  If successful on β‰₯ 3 (majority) β†’ lock acquired
  If < 3 β†’ release all, retry after random delay

RELEASE:
  DEL key on ALL 5 instances

Pros: Tolerates up to 2 Redis nodes dying. Cons: Complex. Debated in the community (Martin Kleppmann’s critique). Clock skew issues.

For interviews: Mention Redlock exists but say β€œfor most applications, single Redis SETNX with TTL is sufficient. Redlock adds complexity.”


3. DynamoDB Conditional Write

Use PutItem with a condition expression.

Put lock item:
  PK: "lock:order_123"
  owner: "server_A"
  expiresAt: now + 30s
  ConditionExpression: "attribute_not_exists(PK) OR expiresAt < :now"

If condition passes β†’ lock acquired
If condition fails β†’ someone else holds it β†’ retry or fail

Pros: No extra infrastructure (already using DynamoDB). Serverless-friendly. TTL via DynamoDB TTL feature. Cons: ~10ms latency (vs ~1ms Redis). Eventually consistent reads might miss a lock (use consistent reads).


4. Database Row Locking (SELECT FOR UPDATE)

Use your existing relational database.

BEGIN;
SELECT * FROM locks WHERE resource = 'order_123' FOR UPDATE;
-- If row exists and not expired β†’ someone holds it β†’ rollback
-- If row doesn't exist or expired β†’ insert/update with our owner
INSERT INTO locks (resource, owner, expires_at) 
VALUES ('order_123', 'server_A', NOW() + INTERVAL '30 seconds')
ON CONFLICT (resource) DO UPDATE SET owner = 'server_A', expires_at = ...
WHERE locks.expires_at < NOW();
COMMIT;

Pros: ACID guaranteed. No extra infrastructure. Cons: Database becomes the bottleneck. Locks are slow (disk I/O). Connection pool pressure.


5. ZooKeeper / etcd (Heavyweight)

Create an ephemeral node. When the holder disconnects, the node auto-deletes (lock released).

ACQUIRE: create("/locks/order_123", ephemeral=true)
  β†’ success: you have the lock
  β†’ already exists: set a watch, wait for deletion

RELEASE: node auto-deletes when session ends (client disconnects/crashes)

Pros: Battle-tested. Auto-cleanup on crash. No TTL guessing. Cons: Heavy infrastructure (ZooKeeper cluster). Higher latency. Overkill for most use cases.


Comparison

Method Latency Reliability Complexity Best For
Redis SETNX ~1ms Single Redis SPOF Low Most applications
Redlock ~5ms Tolerates minority failure High Critical locks needing HA
DynamoDB ~10ms Highly available (managed) Medium AWS/serverless apps
DB Row Lock ~20-50ms As reliable as your DB Low Small scale, already have DB
ZooKeeper ~10-20ms Very high (consensus) High Distributed coordination

Common Problems and Solutions

Problem 1: Lock Holder Crashes (Deadlock)

Lock is acquired, holder crashes, lock never released.

Solution: TTL/expiry on every lock. After TTL, lock auto-releases.

Trade-off: TTL too short β†’ lock expires while holder is still working (split-brain). TTL too long β†’ long wait if holder crashes.

Best practice: TTL = 3x expected operation time. If operation takes 5s, set TTL = 15s.


Problem 2: Lock Expires While Still Working

Server A acquires lock (TTL=30s)
Server A's operation takes 35s (slow network/GC pause)
Lock expires at 30s β†’ Server B acquires lock
Both A and B are now operating on the same resource! ← DANGEROUS

Solution: Fencing Tokens

Every lock acquisition returns an incrementing token number.

Server A gets lock with token=33
Server B gets lock with token=34 (after A's expires)

Storage checks: "reject writes with token < current max"
Server A tries to write with token=33 β†’ REJECTED (current max is 34)
Server B writes with token=34 β†’ ACCEPTED

Problem 3: Lock Ordering (Deadlock Between Two Locks)

Server A: acquires lock on Wallet_1, then tries lock on Wallet_2
Server B: acquires lock on Wallet_2, then tries lock on Wallet_1
β†’ Both wait forever (deadlock)

Solution: Always acquire locks in the same order (sorted by ID).

// Transfer from wallet A to wallet B
Wallet first = walletA.id < walletB.id ? walletA : walletB;
Wallet second = walletA.id < walletB.id ? walletB : walletA;

acquireLock(first.id);   // Both threads try this first
acquireLock(second.id);  // Then this
// No circular wait β†’ no deadlock

Common Interview Questions

Q: β€œHow do you prevent double-booking?” A: β€œUse a distributed lock on the resource (seat_id). Before booking, acquire lock with Redis SETNX + TTL. If acquired, proceed with booking. If not, return β€˜already booked’. Release lock after commit.”

Q: β€œWhat if the lock holder crashes?” A: β€œTTL auto-releases the lock. The next server to try will acquire it. The crashed server’s partial work should be rolled back (use DB transactions) or designed to be idempotent.”

Q: β€œRedis SETNX vs database locking?” A: β€œRedis is faster (~1ms vs ~20ms) and doesn’t pressure your DB connection pool. Use Redis if you have it. Use DB locking only if you don’t want to add Redis infrastructure.”

Q: β€œHow do you handle lock contention (too many waiting)?” A: β€œ1) Exponential backoff on retry. 2) Timeout β€” if can’t acquire within 5s, fail fast. 3) Redesign to reduce lock scope (lock individual items, not entire collections).”


← Back to Fundamentals ← Database Sharding Next: Event Sourcing & CQRS β†’

Free system design + DSA prep. If it helped you crack an interview, consider supporting.

SensAI SensAI
Beta
Listening...
Tap mic to stop voice mode

Shape what we build next

Every piece of feedback is read by the team and directly influences our roadmap.

What type of feedback?

Install SystemCraft

Add to your home screen for instant access, offline reading, and a distraction-free experience.

Offline reading Faster loads No browser tabs App-like feel

Unlock AI Features

One click to activate - no payment, no credit card. Just sign in and you're in.

AI code review and hints
SensAI chat assistant
AI mock interviews
Whiteboard analysis
100% free during early access