Limited time: AI code review, hints, mock interviews, whiteboard analysis, and all Pro features are unlocked. Enroll
⏱️ 19 min read

Image Processing Microservice - HLD

Difficulty: Intermediate–Advanced
Prerequisites: Message Queues, Object Storage, and Scalability
Asked at: Amazon, Google, Cloudinary, imgix, Shopify, Flipkart


TL;DR

An async image processing service that accepts image transformation jobs (resize, compress, format convert, watermark, thumbnail), executes them through a scalable worker pool, and delivers results via webhook or polling. Presigned URLs handle large uploads without proxying bytes through the API.

flowchart LR
    CLIENT["Client<br/>submits job"]:::client
    API["Processing API"]:::service
    S3[("Object Storage<br/>S3")]:::data
    QUEUE["Job Queue<br/>SQS or Kafka"]:::async
    WORKERS["Worker Pool<br/>libvips"]:::service
    WEBHOOK["Webhook<br/>callback"]:::client

    CLIENT -->|"1. Get presigned URL"| API
    CLIENT -->|"2. Upload image"| S3
    CLIENT -->|"3. POST job with operations"| API
    API -->|"4. Enqueue job"| QUEUE
    QUEUE -->|"5. Pull and process"| WORKERS
    WORKERS -->|"6. Read source"| S3
    WORKERS -->|"7. Write output"| S3
    WORKERS -->|"8. Notify completion"| WEBHOOK

    classDef client fill:#4c3a5e,stroke:#818cf8,color:#e2e8f0
    classDef service fill:#1a3a2a,stroke:#4ade80,color:#e2e8f0
    classDef async fill:#AB47BC,stroke:#4A148C,color:#fff
    classDef data fill:#3b3520,stroke:#fbbf24,color:#e2e8f0
Color Role
Purple Client
Green Service
Magenta Async / broker
Gold Data store

In 3 sentences: Clients upload images to S3 via presigned URLs, then submit a job describing which operations to chain (resize β†’ compress β†’ convert). The API enqueues to SQS/Kafka; stateless workers pull jobs, stream images through libvips, write outputs back to S3, and fire a webhook. Auto-scaling on queue depth keeps latency low without wasting compute.


Functional Requirements

Core (top 3)

  1. Submit a processing job β€” client uploads an image and specifies a pipeline of operations (resize, compress, format convert, watermark, thumbnail generation).
  2. Execute the pipeline β€” workers process operations in sequence, streaming the image through each step, and store the result in object storage.
  3. Retrieve results β€” client polls job status or receives a webhook callback with the output URL on completion.

Below the line


Non-Functional Requirements

Core

Below the line


Core Entities


API / System Interface

POST   /v1/upload-url              -> { presigned_url, source_key }
POST   /v1/jobs                    -> Job
GET    /v1/jobs/{id}               -> Job + output URL

POST /v1/upload-url

// Request
{ "filename": "product-hero.jpg", "content_type": "image/jpeg", "size_bytes": 2048000 }

// Response
{ "presigned_url": "https://bucket.s3.amazonaws.com/uploads/abc123?X-Amz-...",
  "source_key": "uploads/abc123/product-hero.jpg",
  "expires_in_seconds": 300 }

POST /v1/jobs

// Request
{ "source_key": "uploads/abc123/product-hero.jpg",
  "operations": [
    { "type": "resize", "width": 800, "height": 600, "fit": "cover" },
    { "type": "compress", "quality": 80 },
    { "type": "convert", "format": "webp" }
  ],
  "webhook_url": "https://myapp.com/hooks/image-done" }

// Response
{ "job_id": "job_7f2a9c", "status": "PENDING", "created_at": "2026-08-17T10:00:00Z" }

GET /v1/jobs/{id}

{ "job_id": "job_7f2a9c", "status": "COMPLETED",
  "output_url": "https://cdn.example.com/outputs/job_7f2a9c.webp",
  "completed_at": "2026-08-17T10:00:03Z" }

Security notes:


High-Level Design

FR-1: Submit a processing job

Components introduced:

  1. Processing API β€” validates the request, generates presigned URLs, computes the idempotency hash, and enqueues jobs.
  2. S3 (source bucket) β€” holds uploaded source images. Clients write directly via presigned URL.
  3. Redis (idempotency cache) β€” stores hash(source + operations) β†’ job_id so duplicate submissions return the existing result instantly.
  4. Postgres (jobs table) β€” durable record of every job and its state.
flowchart LR
    CLIENT["Client"]:::client
    API["Processing API"]:::service
    S3[("S3<br/>source bucket")]:::data
    REDIS[("Redis<br/>idem cache")]:::data
    PG[("Postgres<br/>jobs")]:::data

    CLIENT -->|"1. GET presigned URL"| API
    CLIENT -->|"2. PUT image directly"| S3
    CLIENT -->|"3. POST job"| API
    API -->|"4. Check idem hash"| REDIS
    API -->|"5. Insert job row"| PG

    classDef client fill:#4c3a5e,stroke:#818cf8,color:#e2e8f0
    classDef service fill:#1a3a2a,stroke:#4ade80,color:#e2e8f0
    classDef data fill:#3b3520,stroke:#fbbf24,color:#e2e8f0

Step-by-step flow:

  1. Client requests a presigned upload URL from the API.
  2. Client uploads the raw image directly to S3 β€” the API never touches the bytes.
  3. Client submits POST /v1/jobs with the source_key and operations array.
  4. API computes idempotency_key = SHA256(source_key + canonicalized operations JSON).
  5. API checks Redis: if this hash exists, return the existing job_id immediately (cache hit = free dedup).
  6. On cache miss: insert a new job row in Postgres with status PENDING, write the hash to Redis.
  7. Return 201 Created with the job_id.

Why presigned URLs? At 1000 images/sec Γ— 2 MB average, that’s 2 GB/s of bandwidth. Proxying through the API would require enormous network capacity. Presigned URLs let S3 absorb the upload traffic directly.


FR-2: Execute the pipeline

New components:

  1. SQS Job Queue β€” decouples submission from processing. Visibility timeout ensures at-least-once delivery.
    πŸ’‘ Visibility timeout = after a worker receives a message, it becomes invisible to other workers for N seconds. If the worker doesn’t delete it (crashes), the message reappears for another worker to retry.
  2. Worker Pool β€” stateless processes that pull from SQS, execute the image pipeline using libvips, and write outputs to S3.
  3. S3 (output bucket) β€” stores processed images. Separate from source bucket for lifecycle policies.
flowchart LR
    API["Processing API"]:::service
    SQS["SQS<br/>job queue"]:::async
    WORKER["Worker Pool<br/>libvips"]:::service
    S3_SRC[("S3<br/>source")]:::data
    S3_OUT[("S3<br/>outputs")]:::data
    PG[("Postgres<br/>jobs")]:::data

    API -->|"1. SendMessage"| SQS
    SQS -->|"2. ReceiveMessage"| WORKER
    WORKER -->|"3. GetObject source"| S3_SRC
    WORKER -->|"4. PutObject result"| S3_OUT
    WORKER -->|"5. UPDATE status"| PG

    classDef service fill:#1a3a2a,stroke:#4ade80,color:#e2e8f0
    classDef async fill:#AB47BC,stroke:#4A148C,color:#fff
    classDef data fill:#3b3520,stroke:#fbbf24,color:#e2e8f0

Step-by-step flow:

  1. API sends a message to SQS containing { job_id, source_key, operations }.
  2. Worker long-polls SQS with ReceiveMessage. Visibility timeout set to 60s.
  3. Worker downloads the source image from S3 into a streaming buffer.
  4. Worker chains operations using libvips: source β†’ resize β†’ compress β†’ convert. Single streaming pass, no full-image buffers.
  5. Worker uploads output to S3 (output bucket, keyed by job_id).
  6. Worker updates Postgres: status = COMPLETED, output_url, completed_at.
  7. Worker deletes the SQS message (acknowledges success).
  8. If worker crashes, visibility timeout expires after 60s, message reappears for another worker.

Why SQS over Kafka here? SQS’s per-message visibility timeout is a natural fit for image processing where each job takes 1-5s. No partition management needed. For 10K+ images/sec, Kafka with consumer groups is the better choice.


FR-3: Retrieve results

New components:

  1. Webhook Dispatcher β€” async service that fires HTTP callbacks to the client’s registered URL on job completion.
  2. CDN β€” fronts the output S3 bucket for low-latency delivery of processed images.
flowchart LR
    WORKER["Worker"]:::service
    PG[("Postgres<br/>jobs")]:::data
    WHOOK["Webhook<br/>Dispatcher"]:::service
    CLIENT["Client"]:::client
    CDN["CDN<br/>CloudFront"]:::service
    S3_OUT[("S3<br/>outputs")]:::data

    WORKER -->|"1. Job complete"| PG
    WORKER -->|"2. Fire webhook"| WHOOK
    WHOOK -->|"3. POST callback"| CLIENT
    CLIENT -->|"4. GET output"| CDN
    CDN -->|"5. Origin fetch"| S3_OUT

    classDef client fill:#4c3a5e,stroke:#818cf8,color:#e2e8f0
    classDef service fill:#1a3a2a,stroke:#4ade80,color:#e2e8f0
    classDef data fill:#3b3520,stroke:#fbbf24,color:#e2e8f0

Step-by-step flow:

  1. Worker completes processing, updates Postgres with COMPLETED status and output URL.
  2. Worker publishes a completion event to the webhook dispatcher.
  3. Webhook dispatcher POSTs to the client’s webhook_url with { job_id, status, output_url }.
  4. If webhook delivery fails, retries with exponential backoff (1s, 5s, 30s, 2min) up to 5 attempts.
  5. Client can also poll GET /v1/jobs/{id} at any time to check status.
  6. Output images are served through a CDN for low-latency global access.

Technology Choices

Tier / Purpose Primary Alternatives Why
API Gateway Node.js or Go HTTP service FastAPI, Spring Boot Lightweight; handles presigned URL generation and job submission
Object Storage AWS S3 GCS, Azure Blob, MinIO Presigned URL uploads, cheap at-rest storage for source and output images
Job Queue SQS with visibility timeout Kafka, RabbitMQ, Google Pub/Sub Managed, per-message visibility timeout for retry, built-in DLQ
Job Metadata PostgreSQL MySQL, CockroachDB ACID for job state, indexed queries for status polling
Idempotency Cache Redis Memcached, DynamoDB Sub-ms hash lookups; TTL-based expiry for cache entries
Image Processing libvips ImageMagick, Sharp (Node bindings to libvips), Pillow Streaming architecture β€” 4-8x faster than ImageMagick, 1/10th memory
Auto-scaling K8s HPA on custom metrics AWS Auto Scaling, KEDA Scale workers on queue depth metric, not CPU
Dead-letter Queue SQS DLQ Kafka DLQ topic, RabbitMQ dead-letter exchange Isolate poison messages after max retries
Webhook Delivery Async HTTP with retries SNS, EventBridge Direct callback to client with exponential backoff

Why libvips over ImageMagick

ImageMagick loads the full image into memory (200+ MB for a 50 MP image). libvips uses a streaming tile-based pipeline β€” memory stays at ~50 MB regardless of image size. At 200 workers, that’s 40 GB vs 10 GB total.


Data Modeling

Postgres (Job metadata β€” source of truth):

CREATE TABLE jobs (
    job_id UUID PRIMARY KEY,
    tenant_id UUID NOT NULL,
    status VARCHAR(20) NOT NULL DEFAULT 'PENDING',
    source_url TEXT NOT NULL,
    operations JSONB NOT NULL,
    idempotency_key VARCHAR(64) NOT NULL,
    output_url TEXT,
    webhook_url TEXT,
    error_message TEXT,
    attempts INTEGER DEFAULT 0,
    created_at TIMESTAMP NOT NULL DEFAULT NOW(),
    completed_at TIMESTAMP,
    CONSTRAINT uq_idempotency UNIQUE(tenant_id, idempotency_key)
);
CREATE INDEX idx_jobs_status ON jobs(status) WHERE status IN ('PENDING', 'PROCESSING');
CREATE INDEX idx_jobs_tenant ON jobs(tenant_id, created_at DESC);

Redis (Idempotency cache + worker state):

Key: "idem:{hash}" β†’ job_id (where hash = SHA256(source_url + canonical JSON of operations))
TTL: 24 hours

Key: "worker:{workerId}:heartbeat" β†’ { job_id, started_at }
TTL: 30s

Access Patterns:

Query Data Source How
Submit job Redis + Postgres Check idem cache β†’ if miss, INSERT job row, enqueue
Poll job status Postgres SELECT status, output_url FROM jobs WHERE job_id = ?
Worker claims job SQS ReceiveMessage with visibility timeout
Worker completes Postgres + Redis UPDATE job status, SET idem cache with output URL
Detect stuck workers Redis TTL Heartbeat expires β†’ message becomes visible again (SQS handles this)

Deep Dives

Deep Dive 1: Worker Scaling

The problem: Queue depth fluctuates wildly. E-commerce sites upload 10x more images during flash sales. Fixed worker counts either waste money (idle at baseline) or drop latency (overwhelmed at peak).

Bad β€” fixed worker count:
Provision for peak (200 workers) and eat the cost 24/7. At baseline (100 images/sec), 180 workers sit idle. Monthly cost: 200 Γ— $150 = $30K in compute for maybe $5K worth of actual work.

Good β€” CPU-based auto-scaling:
Scale workers based on CPU utilization (target 70%). Problem: image processing is bursty β€” by the time CPU spikes, the queue is already deep and latency has blown past SLA. There’s a 2-3 minute lag between queue depth spiking and CPU reflecting it.

Great β€” queue-depth-based scaling with KEDA or HPA custom metrics:


Deep Dive 2: Idempotency and Dedup

The problem: Clients retry on timeout. The same product image gets submitted with the same operations 3x. Without dedup, we waste compute processing it 3x and potentially confuse clients with multiple output URLs.

Bad β€” no dedup:
Every submission creates a new job. Three retries = three processing runs = 3x cost. Client gets three different job_ids and doesn’t know which output to use.

Good β€” database unique constraint:
Add a unique index on (tenant_id, idempotency_key). Second submission gets a constraint violation, API returns the existing job. Problem: this requires a Postgres round-trip on every submission, and under high concurrency you get lock contention on the index.

Great β€” Redis cache with hash key:

Canonical JSON: Sort keys alphabetically before hashing so {"width":800,"height":600} and {"height":600,"width":800} produce the same hash. This is critical β€” without canonicalization, semantically identical pipelines get different hashes.


Deep Dive 3: Failure Handling

The problem: Workers crash (OOM on a 200 MB TIFF), images are corrupt (truncated upload), operations are invalid (resize to 0Γ—0). Without proper failure handling, jobs get stuck forever or retry infinitely.

Bad β€” infinite retries:
Worker crashes, message reappears, next worker crashes on the same corrupt image. Repeat forever. The β€œpoison message” consumes worker capacity and never completes.

Good β€” max retries with exponential backoff:
Set maxReceiveCount = 3 on SQS. After 3 failures, message moves to DLQ. Problem: transient failures (S3 throttling) also end up in DLQ after just 3 attempts.

Great β€” tiered retry with DLQ and alerting:


Deep Dive 4: Performance β€” libvips vs ImageMagick

The problem: At 1000 images/sec, processing speed and memory usage directly determine infrastructure cost. The wrong library choice means 4x more servers.

Bad β€” ImageMagick (load-all architecture):
Decompresses the entire image into memory. A 50 MP JPEG becomes ~288 MB of raw pixels. With 200 workers, that’s 57 GB of RAM just for pixel buffers. Workers need 4 GB each.

Good β€” Pillow/PIL with partial streaming:
Can do lazy loading for some formats, but resize still loads the full image. Better than ImageMagick, not optimal.

Great β€” libvips streaming pipeline:


Core Flows

Flow 1: Submit Job

sequenceDiagram
    actor Client
    participant API as Processing API
    participant Redis
    participant PG as Postgres
    participant SQS

    Client->>API: POST /v1/upload-url
    API-->>Client: presigned_url + source_key
    Client->>Client: PUT image to S3 via presigned URL

    Client->>API: POST /v1/jobs (source_key + operations)
    API->>API: compute hash = SHA256(source_key + ops)
    API->>Redis: GET idem:{hash}

    alt Cache Hit
        Redis-->>API: existing job_id
        API-->>Client: 200 (existing job with output_url)
    else Cache Miss
        API->>PG: INSERT job (status=PENDING)
        API->>Redis: SET idem:{hash} job_id EX 86400
        API->>SQS: SendMessage (job_id + source_key + ops)
        API-->>Client: 201 Created (job_id)
    end

Walkthrough:

  1. Client gets a presigned URL and uploads directly to S3.
  2. Client submits the job. API hashes source + operations for idempotency.
  3. Redis cache hit β†’ return existing result immediately (no reprocessing).
  4. Cache miss β†’ create job, cache the hash, enqueue to SQS.

Non-obvious failure: If API crashes between Postgres INSERT and SQS SendMessage, the job is stuck as PENDING. A reconciliation job (runs every 5 min) detects jobs in PENDING for > 2 minutes and re-enqueues them.


Flow 2: Process Job

sequenceDiagram
    participant SQS
    participant Worker
    participant S3_Src as S3 Source
    participant S3_Out as S3 Output
    participant PG as Postgres
    participant Webhook as Webhook Dispatcher

    Worker->>SQS: ReceiveMessage (visibility=60s)
    SQS-->>Worker: message (job_id, source_key, ops)
    Worker->>PG: UPDATE status=PROCESSING
    Worker->>S3_Src: GetObject (stream)
    Worker->>Worker: libvips pipeline (resize, compress, convert)
    Worker->>S3_Out: PutObject (output)
    Worker->>PG: UPDATE status=COMPLETED, output_url
    Worker->>Redis: SET idem:{hash} job_id (refresh TTL)
    Worker->>SQS: DeleteMessage
    Worker->>Webhook: POST callback (job_id, output_url)

    alt Worker Crashes
        Note over SQS: visibility timeout expires after 60s
        SQS-->>Worker: message reappears for another worker
    end

    alt Permanent Failure
        Worker->>PG: UPDATE status=FAILED, error_message
        Worker->>SQS: DeleteMessage
        Worker->>Webhook: POST callback (job_id, FAILED, error)
    end

Walkthrough:

  1. Worker long-polls SQS. Receives a message with 60s visibility timeout.
  2. Marks job as PROCESSING in Postgres.
  3. Streams source image from S3 into libvips pipeline.
  4. Executes operations in sequence (resize β†’ compress β†’ convert) as one streaming pass.
  5. Uploads output to S3, updates job as COMPLETED, refreshes idempotency cache.
  6. Deletes SQS message and fires webhook to notify client.

Non-obvious failure: Worker takes > 60s (huge image). Visibility timeout expires, another worker picks up the same message. Solution: extend visibility timeout periodically (ChangeMessageVisibility every 30s while processing), and the idempotency cache prevents duplicate entries.


Final Architecture

The consolidated view after all deep dives β€” showing every component and how they connect:

flowchart LR
    CLIENT["Client"]:::client
    API["Processing API<br/>presigned URLs + job CRUD"]:::service
    REDIS[("Redis<br/>idempotency cache")]:::data
    PG[("Postgres<br/>jobs table")]:::data
    SQS["SQS<br/>job queue + DLQ"]:::async
    S3_SRC[("S3<br/>source bucket")]:::data
    WORKERS["Worker Pool<br/>libvips streaming"]:::service
    S3_OUT[("S3<br/>output bucket")]:::data
    CDN["CDN<br/>CloudFront"]:::service
    WEBHOOK["Webhook<br/>Dispatcher"]:::service
    HPA["Auto-Scaler<br/>KEDA or HPA"]:::service

    CLIENT -->|"presigned upload"| S3_SRC
    CLIENT -->|"submit job"| API
    API -->|"check dedup"| REDIS
    API -->|"store job"| PG
    API -->|"enqueue"| SQS
    SQS -->|"pull jobs"| WORKERS
    WORKERS -->|"stream source"| S3_SRC
    WORKERS -->|"write output"| S3_OUT
    WORKERS -->|"update status"| PG
    WORKERS -->|"set cache"| REDIS
    WORKERS -->|"notify"| WEBHOOK
    WEBHOOK -->|"callback"| CLIENT
    CLIENT -->|"fetch output"| CDN
    CDN -->|"origin"| S3_OUT
    HPA -->|"scale on queue depth"| WORKERS

    classDef client fill:#4c3a5e,stroke:#818cf8,color:#e2e8f0
    classDef service fill:#1a3a2a,stroke:#4ade80,color:#e2e8f0
    classDef async fill:#AB47BC,stroke:#4A148C,color:#fff
    classDef data fill:#3b3520,stroke:#fbbf24,color:#e2e8f0

Reading this diagram:


Interview Tips

Discussion

Newest first
You

Free system design + DSA prep. If it helped you crack an interview, consider supporting.

SensAI SensAI
Beta
Listening...
Tap mic to stop voice mode

Shape what we build next

Every piece of feedback is read by the team and directly influences our roadmap.

What type of feedback?

Install SystemCraft

Add to your home screen for instant access, offline reading, and a distraction-free experience.

Offline reading Faster loads No browser tabs App-like feel

Unlock AI Features

One click to activate - no payment, no credit card. Just sign in and you're in.

AI code review and hints
SensAI chat assistant
AI mock interviews
Whiteboard analysis
100% free during early access