smaple.tr
background job processing

Background Job Processing: Queue Systems, Retry Strategies, and Dead Letter Queues [2026]

Mehmet Kurtipek
November 25, 2025
14 min read
background job processing
queue systems
BullMQ
Celery
Sidekiq
dead letter queue
worker scaling
retry strategies

Processing a payment takes 500ms. Generating a PDF report takes 30 seconds. Sending 50,000 marketing emails takes 45 minutes. Doing any of these synchronously in an HTTP request handler destroys user experience and server stability. Background job processing solves this: defer work to an asynchronous queue, return an immediate response to the user, and process the work at scale.

This guide covers the complete background job processing stack: queue pattern selection, technology comparison (BullMQ, Sidekiq, Celery, RabbitMQ, Kafka), retry strategies with exponential backoff, dead letter queue design, idempotency guarantees, priority queue architecture, and worker scaling patterns. By the end, you will have a clear framework for building reliable asynchronous processing infrastructure.

Background Job Processing: Queue Pattern Selection

Three queue patterns cover the majority of background job use cases. Selecting the wrong pattern creates reliability and scalability problems that are expensive to fix later.

Work Queue (Point-to-Point)

The work queue pattern ensures each job is processed by exactly one worker. A job enters the queue; one worker picks it up and processes it; the job is removed from the queue. If the worker fails mid-processing, the job returns to the queue for another worker to attempt.

This is the default pattern for job processing: email delivery, image resizing, report generation, webhook dispatch. Use the work queue pattern when each job must be processed once and only once.

[Producer] → Queue → [Worker 1]
                   → [Worker 2]
                   → [Worker 3]
(each job goes to exactly one worker)

Scaling: add more workers to increase throughput. Queue depth (jobs waiting) is the primary scaling signal.

Publish/Subscribe (Fan-Out)

The pub/sub pattern delivers a single message to all subscribers. When an order is placed, the Order service publishes an order.created event. The Notification service, Analytics service, and Inventory service each subscribe independently and each receive a copy of the event.

Use pub/sub when a single event should trigger independent processing by multiple consumers. This is the right pattern for event-driven architecture integration points, not for simple job queues.

Critical distinction from work queues: pub/sub does not guarantee once-and-only-once delivery per consumer group — it guarantees delivery to every subscriber. If your Inventory service crashes and misses an event, it does not get redelivered unless you use an event streaming system (Kafka) rather than a simple message broker.

Request/Reply (Async RPC)

The request/reply pattern sends a job to a queue and waits for the result. This is asynchronous remote procedure call: the caller does not block the HTTP thread, but it does expect a result before proceeding.

Use cases: long-running computations where the caller needs the result (not just confirmation of receipt), third-party API calls with unpredictable latency, batch processing with progress reporting.

Implementation: the caller sends a message with a reply_to queue name and a correlation_id. The worker processes the request and publishes the result to the reply_to queue with the matching correlation_id. The caller polls or subscribes to the reply queue.

Request/reply is more complex than fire-and-forget work queues. Use it only when you genuinely need the result — not as a way to make synchronous operations appear asynchronous.

Technology Selection

Redis + BullMQ (Node.js/TypeScript)

BullMQ is the leading job queue library for Node.js, built on Redis. It provides a rich feature set on top of Redis's fast in-memory operations:

Core features: priority queues, delayed jobs (scheduled for future execution), repeatable jobs (cron-like scheduling), rate limiting, job lifecycle hooks, sandboxed workers (isolated execution context), parent/child job flows.

Developer experience: full TypeScript support, declarative job definitions, Bull Board dashboard for visual monitoring.

Architecture: BullMQ stores jobs in Redis sorted sets and lists. Waiting jobs are in a sorted set ordered by priority and delay. Active jobs are in a list. Completed and failed jobs are in separate sorted sets with configurable retention.

import { Queue, Worker } from 'bullmq';

const emailQueue = new Queue('email', {
  connection: { host: 'redis', port: 6379 },
  defaultJobOptions: {
    attempts: 3,
    backoff: { type: 'exponential', delay: 2000 },
    removeOnComplete: { count: 1000 },
    removeOnFail: { count: 5000 },
  },
});

// Enqueue a job
await emailQueue.add('send-confirmation', {
  to: '[email protected]',
  orderId: '12345',
});

// Process jobs
const worker = new Worker('email', async (job) => {
  await sendEmail(job.data.to, job.data.orderId);
}, { connection: { host: 'redis', port: 6379 } });

Limitations: Redis is an in-memory store. Large job payloads (files, large datasets) should be stored externally (S3) with the job carrying a reference. Redis persistence (RDB/AOF) provides durability, but Redis cluster management adds operational complexity at large scale.

Best for: Node.js/TypeScript applications, moderate job volumes (< 10,000 jobs/minute), teams that already run Redis for caching.

Sidekiq (Ruby)

Sidekiq is the standard background job processor for Ruby on Rails applications. Its multi-threaded architecture processes multiple jobs concurrently within a single process, unlike older Ruby job processors that used single-threaded workers.

Core features: priority queues, scheduled jobs, retry with backoff, web UI, middleware pipeline for job instrumentation, batches (Sidekiq Pro).

Performance: a single Sidekiq process with 25 threads processes 25 concurrent jobs. For I/O-bound jobs (API calls, database queries), this provides high throughput without horizontal scaling.

At Smart Maple, we use Sidekiq in Rails-based marketplace and booking applications for webhook delivery, email batches, and data sync jobs. The key configuration pattern for webhook delivery: separate queues per destination with per-queue concurrency limits to prevent one slow destination from blocking others.

Best for: Ruby on Rails applications, job queues that need tight Rails integration, teams comfortable with Redis-based persistence.

Celery (Python)

Celery is the standard distributed task queue for Python. Unlike BullMQ and Sidekiq, Celery supports multiple broker backends: Redis, RabbitMQ, and Amazon SQS.

Core features: task chaining and groups (workflow primitives), periodic tasks (via Celery Beat), canvas workflow (chain, group, chord, map, starmap), result backend support, canvas primitives for complex workflow composition.

Workflow composition:

from celery import chain, group

# Chain: execute sequentially
result = chain(
    download_file.s(url),
    process_data.s(),
    upload_result.s()
).delay()

# Group: execute in parallel
result = group(
    resize_image.s('small'),
    resize_image.s('medium'),
    resize_image.s('large'),
).delay()

Celery Beat — the periodic task scheduler — runs as a separate process and publishes scheduled tasks to the queue. Important: do not run multiple Celery Beat instances. Only one Beat process should publish periodic tasks, or tasks will be dispatched multiple times.

Best for: Python applications, complex workflow composition (ML pipelines, ETL), teams using RabbitMQ or SQS as broker.

RabbitMQ (Language-Agnostic)

For heterogeneous environments (multiple languages, multiple services), RabbitMQ provides a language-agnostic broker with client libraries for every major language.

RabbitMQ is particularly strong for background job processing when:

  • Multiple services in different languages need to produce or consume jobs
  • Complex routing logic (jobs routed to different workers based on job attributes)
  • AMQP protocol compliance is required for enterprise integration

Consumer prefetch is the critical configuration for RabbitMQ-backed job queues:

channel.basic_qos(prefetch_count=10)

Without prefetch limits, RabbitMQ dispatches all available jobs to connected workers immediately. A slow worker accumulates thousands of jobs in its "in-flight" state, preventing other workers from picking them up. Set prefetch to 1-10 for CPU-intensive jobs, higher for fast I/O jobs.

Retry Strategies and Dead Letter Queue Design

Exponential Backoff with Jitter

The standard retry strategy for transient failures:

def calculate_delay(attempt: int, base_delay: float = 1.0, cap: float = 60.0) -> float:
    """Exponential backoff with full jitter."""
    exponential = min(cap, base_delay * (2 ** attempt))
    return random.uniform(0, exponential)

# Delay sequence (approximate):
# Attempt 1: 0-2 seconds
# Attempt 2: 0-4 seconds
# Attempt 3: 0-8 seconds
# Attempt 4: 0-16 seconds
# Attempt 5: 0-32 seconds

Jitter prevents the "thundering herd" problem: without jitter, all failed jobs retry simultaneously after the backoff period, creating a spike that overwhelms the recovering service.

Configuration in BullMQ:

{
  attempts: 5,
  backoff: {
    type: 'exponential',
    delay: 1000,  // 1 second base delay
  }
}

Failure Classification

Not all failures should retry. Classify failures before deciding on retry behavior:

Transient failures (retry with backoff):

  • Network timeout / connection refused
  • HTTP 429 (rate limited) — respect Retry-After header
  • HTTP 503 (service temporarily unavailable)
  • Database connection errors
  • Resource contention (lock timeout, queue full)

Permanent failures (no retry, route to DLQ):

  • HTTP 400 (invalid request data — retrying won't fix it)
  • HTTP 401 / 403 (authentication/authorization failure)
  • HTTP 404 (resource not found — the referenced entity may not exist)
  • Validation errors in job payload
  • Business rule violations

Retrying permanent failures wastes resources, delays the identification of bugs, and creates misleading metrics. Implement failure type detection in your error handler:

def should_retry(exception) -> bool:
    if isinstance(exception, (NetworkError, TimeoutError)):
        return True
    if isinstance(exception, HTTPError):
        return exception.status_code in {429, 502, 503, 504}
    return False  # Default: do not retry

Dead Letter Queue Architecture

Every production queue must have a dead letter queue. The DLQ receives jobs that have exhausted their retry attempts. DLQ jobs require:

Monitoring and alerting — DLQ depth should alert at any non-zero value. A job in the DLQ represents a failed business process that requires investigation.

Inspection tooling — operators need to inspect DLQ jobs: view the job payload, the failure reason, and the retry history. Bull Board (BullMQ), Sidekiq's built-in web UI, and Flower (Celery) provide this visibility.

Selective replay — after fixing the underlying issue, operators need to replay specific DLQ jobs or ranges of jobs without replaying everything. Implement a replay tool that moves jobs back to the primary queue with a reset attempt counter.

DLQ retention policy — DLQ jobs should expire after a configurable period (7-30 days). Unbounded DLQ growth indicates a systematic problem, not isolated failures.

Idempotency in Job Processing

Why Idempotency Matters

At-least-once delivery guarantees mean any job can be processed more than once. This happens when:

  • A worker processes a job and crashes before acknowledging delivery
  • A network partition makes the acknowledgment invisible to the broker
  • A worker takes longer than the job's visibility timeout, and the broker requeues it

An idempotent job produces the same outcome regardless of how many times it runs. Achieving idempotency requires explicit design:

Natural idempotency — some operations are idempotent by nature. Setting a value (UPDATE users SET status = 'active' WHERE id = ?) is idempotent. Incrementing a counter (UPDATE stats SET count = count + 1) is not.

Idempotency keys — assign a unique key to each job and record processed keys:

def process_payment_job(job_data: dict):
    idempotency_key = f"payment:{job_data['order_id']}"

    # Check if already processed (using Redis SET NX)
    if not redis.set(idempotency_key, "1", nx=True, ex=3600):
        logger.info(f"Duplicate job detected: {idempotency_key}")
        return  # Already processed

    try:
        charge_card(job_data['amount'], job_data['card_token'])
        db.execute(
            "INSERT INTO payments (order_id, amount, status) VALUES (?, ?, 'completed')",
            (job_data['order_id'], job_data['amount'])
        )
    except Exception:
        redis.delete(idempotency_key)  # Allow retry on failure
        raise

Database-level idempotency — use unique constraints and upsert patterns:

INSERT INTO email_sends (job_id, recipient, sent_at)
VALUES (?, ?, NOW())
ON CONFLICT (job_id) DO NOTHING;  -- Duplicate job is a no-op

Priority Queue Architecture

Queue Priority Design

Different job types have different urgency requirements:

Priority Jobs Max latency
Critical Security alerts, payment failures, authentication < 1 second
High Order confirmations, webhook delivery, user-facing actions < 10 seconds
Normal Bulk email, report generation, data sync < 5 minutes
Low Analytics processing, data archival, cleanup jobs < 1 hour

Implement priority as separate queues (not numeric priority on a single queue). Workers check queues in priority order: exhausting critical queue before checking high, high before normal:

// BullMQ worker processing multiple queues in priority order
const workers = [
  new Worker('critical', processor, { concurrency: 10 }),
  new Worker('high', processor, { concurrency: 20 }),
  new Worker('normal', processor, { concurrency: 30 }),
  new Worker('low', processor, { concurrency: 10 }),
];

Preventing Priority Inversion

Priority inversion: low-priority jobs fill the system and prevent high-priority jobs from being processed promptly. Mitigations:

Separate worker pools — dedicate a minimum number of workers exclusively to high-priority queues. These workers never process low-priority jobs, guaranteeing capacity.

Maximum wait time — promote jobs to a higher-priority queue if they wait longer than a threshold. A "normal" job waiting 30 minutes should graduate to "high" to prevent indefinite starvation.

Rate limiting low-priority production — limit how fast jobs can be enqueued in low-priority queues to prevent runaway batch jobs from overwhelming the system.

Worker Scaling

Horizontal Scaling

Workers are stateless processes that read from queues. Scale by adding more workers.

Queue depth-based scaling — scale up when queue depth exceeds a threshold, scale down when queue depth is low. Works well for steady-state loads.

KEDA (Kubernetes Event-Driven Autoscaling) — scales Kubernetes deployments based on queue metrics. KEDA has built-in scalers for Redis (BullMQ depth), RabbitMQ queue length, and custom metrics:

apiVersion: keda.sh/v1alpha1
kind: ScaledObject
metadata:
  name: email-worker-scaler
spec:
  scaleTargetRef:
    name: email-worker
  minReplicaCount: 1
  maxReplicaCount: 50
  triggers:
    - type: redis
      metadata:
        listName: bull:email:wait
        listLength: "50"        # Scale up when > 50 jobs waiting
        targetAverageLength: "25"  # Target 25 jobs per worker

Concurrency Configuration

Concurrency (jobs per worker process) depends on job type:

I/O-bound jobs (HTTP calls, database queries, file operations): high concurrency (20-100). The worker spends most time waiting; high concurrency keeps CPU busy.

CPU-bound jobs (image processing, compression, cryptography): concurrency = number of CPU cores. Over-concurrency causes context-switching overhead and degraded throughput.

Mixed jobs: profile the bottleneck. Measure CPU utilization and I/O wait during processing. Tune concurrency to maximize throughput without overwhelming downstream dependencies.

Job Monitoring Dashboard

Production job queues require visibility:

Metric Alert threshold Meaning
Queue depth (waiting jobs) > 1000 for 5 min Processing capacity insufficient
DLQ depth > 0 Failed jobs requiring attention
Job processing time (p99) > expected max Slow worker or external dependency
Worker utilization > 80% sustained Need to scale workers
Failed job rate > 1% Bug or external dependency issue

Bull Board (BullMQ), Sidekiq's built-in web UI, Flower (Celery), and custom Grafana dashboards via Prometheus are the standard visualization options. Choose the tool that integrates with your existing observability stack.

Common Patterns in Production

Chunked Processing for Large Datasets

Large batch jobs (process 100,000 records) should not run as a single job. Split them:

  1. Enqueue a "batch coordinator" job with the batch ID
  2. The coordinator splits work into chunks (1,000 records each) and enqueues 100 chunk jobs
  3. Each chunk job processes independently and records progress
  4. The coordinator aggregates results after all chunks complete

This enables parallel processing, retry at chunk granularity (not the entire batch), and progress visibility.

Webhook Delivery Pattern

Webhook delivery is a canonical background job use case with specific reliability requirements:

  • Each delivery attempt is a separate job (not a retry of the original)
  • Track delivery attempt count per endpoint
  • Implement exponential backoff with maximum 24-hour delay for failing endpoints
  • Mark endpoints as "unhealthy" after 10 consecutive failures; stop delivering until the endpoint recovers
  • Log every delivery attempt with request headers, response body, and latency for debugging

Scheduled Job Coordination

When running multiple worker instances, prevent duplicate execution of scheduled jobs:

Leader election — one worker instance holds a distributed lock and is the only one that enqueues scheduled jobs. Other instances remain standby.

Idempotent enqueue — use unique job IDs derived from the schedule (e.g., daily-report-2026-04-01) and check for existence before enqueuing.

from redis import Redis

redis = Redis()

def enqueue_daily_report():
    lock_key = f"scheduled:daily-report:{date.today()}"

    # Only one instance successfully acquires the lock
    if redis.set(lock_key, "1", nx=True, ex=3600):
        queue.enqueue(generate_daily_report)

Related Articles

August 11, 2026

MLOps Guide: Taking Machine Learning Models to Production [2026]

87% of machine learning models built by data science teams never reach production. The models work — they pass cross-validation, they score well on holdout sets, they demonstrate genuine predictive value. The problem is not the modeling. The problem is everything that happens between a notebook experiment and a reliable, monitored, production system. MLOps is the discipline that closes that gap. This guide covers the full MLOps stack: maturity levels, tooling choices (MLflow, DVC, Kubeflow

Read More
August 10, 2026

LLM Fine-Tuning Guide: Custom Model Training with LoRA and QLoRA [2026]

General-purpose LLMs are impressive. They can write code, summarize documents, answer questions, and translate between languages with reasonable accuracy. But "reasonable" is not good enough when your application requires consistent output format, domain-specific terminology, a particular tone, or behavior that the base model was never trained to exhibit. That gap is where fine-tuning matters. Fine-tuning updates a model's weights on your specific data, changing how the model behaves — not

Read More
August 9, 2026

Computer Vision Applications: Object Detection, OCR, and Industrial AI [2026]

Computer vision has moved well past the research phase. The models are trained, the frameworks are mature, the hardware is accessible, and the use cases are generating measurable returns. What was a specialized capability requiring deep expertise in 2018 is now deployable infrastructure — if you know which component to reach for and where the real complexity lives. This guide covers computer vision applications across industrial, medical, logistics, and document processing domains. It expl

Read More