smaple.tr
webhook

Webhook Integration: Delivery, Security, and Retry Architecture [2026]

Mehmet Kurtipek
November 4, 2025
11 min read
webhook
event-driven API
HMAC signature
idempotency
message broker
async integration

Webhook integration fails in predictable ways. The endpoint receives the payload, returns HTTP 200, and the developer marks the task done — then discovers two weeks later that 3% of webhooks were silently dropped because the processing code threw an exception after the HTTP response was sent. Or the payment provider retried a failed webhook four times, and the order was fulfilled four times. Or a replay attack replayed a payment notification from six hours ago.

The mechanics of webhook delivery are simple: an HTTP POST from a provider to your endpoint. The engineering requirements for production-grade webhook integration are substantially more complex: signature verification, idempotency, retry handling, dead letter queuing, delivery ordering, schema versioning, and observability. This guide covers each.

Webhook Integration: How Delivery Works

A webhook is an HTTP callback — the provider calls your endpoint when an event occurs. The alternative, polling, requires your system to repeatedly query the provider's API for new events, consuming bandwidth and API quota regardless of event frequency.

Webhook delivery flow:

  1. Your system registers an endpoint URL with the provider
  2. An event occurs in the provider's system
  3. The provider dispatches an HTTP POST to your endpoint containing the event payload
  4. Your endpoint returns HTTP 2xx within the provider's timeout window (typically 5–30 seconds)
  5. The provider records the delivery as successful

What "successful delivery" actually means: HTTP 2xx is returned before your timeout. Your processing code ran correctly. The event was persisted. These three things are independent, and treating HTTP 2xx as a proxy for correct processing is the most common webhook integration bug.

The correct pattern: Return HTTP 200 immediately. Queue the payload for async processing. Process asynchronously. This decouples delivery acknowledgment from processing correctness.

# Correct pattern
@app.post("/webhooks/payment")
async def handle_payment_webhook(request: Request):
    payload = await request.body()
    # Verify signature first
    verify_signature(payload, request.headers.get("X-Signature-256"))
    # Queue for async processing
    await queue.enqueue("process_payment_webhook", payload=payload)
    return {"status": "accepted"}  # HTTP 200 immediately

HMAC Signature Verification

Webhook endpoints are publicly accessible URLs. Without signature verification, any attacker who knows your endpoint URL can send arbitrary payloads. Signature verification ensures the payload came from the expected provider.

How HMAC signing works:

The provider holds a secret key shared with your application (generated at webhook registration). For each delivery, the provider:

  1. Computes HMAC-SHA256 of the raw request body using the secret key
  2. Encodes the signature (typically hex or base64)
  3. Sends the signature in a request header (X-Signature-256, Stripe-Signature, etc.)

Your endpoint recomputes the HMAC and compares. Mismatches indicate either a wrong secret or a tampered payload.

Implementation (Python):

import hmac
import hashlib
import secrets

def verify_signature(payload: bytes, signature_header: str, secret: str) -> bool:
    expected_sig = hmac.new(
        secret.encode('utf-8'),
        payload,
        hashlib.sha256
    ).hexdigest()

    # Use constant-time comparison to prevent timing attacks
    return secrets.compare_digest(
        f"sha256={expected_sig}",
        signature_header
    )

Critical implementation details:

  • Use secrets.compare_digest() (constant-time comparison), not ==. Timing attacks can extract the secret key character by character using standard equality comparison.
  • Compute HMAC over the raw request body bytes, not a parsed/re-serialized version. JSON serialization is not deterministic — field ordering can change.
  • Reject invalid signatures immediately with HTTP 401 before any processing.

Stripe's enhanced format includes a timestamp in the signature to enable replay protection:

Stripe-Signature: t=1614556800,v1=abc123...

The t value is the timestamp (Unix seconds). Reject webhooks where |current_time - t| > 300 seconds (5-minute window). This prevents replay attacks where a valid past payload is re-sent.

Retry Logic and Exponential Backoff

Webhook delivery fails. Your endpoint may be temporarily unavailable, overloaded, or returning errors. Providers retry failed deliveries — but implementations vary significantly.

Provider retry behavior examples:

Provider Max Retries Backoff Pattern Retry Window
Stripe 5 Exponential, ~1hr–72hr 3 days
GitHub 3 Fixed intervals —
Shopify 19 Exponential 48 hours
Twilio 7 Exponential —

Designing your endpoint for retries:

If your endpoint returns 4xx (except 429), most providers treat it as a non-retriable failure. Return 5xx for transient errors to trigger retries. Return 429 with a Retry-After header to signal temporary rate limiting.

If your endpoint is temporarily offline, the provider queues the delivery and retries per its backoff schedule. Design systems so that a 30-minute endpoint outage is recoverable from retry behavior — which means providers need at least one retry at >30 minutes post-failure.

Building your own retry system for outbound webhooks:

When your system sends webhooks to customers' endpoints, implement the same patterns:

import asyncio

async def deliver_webhook_with_retry(endpoint: str, payload: dict, event_id: str):
    max_attempts = 8
    base_delay = 1  # seconds

    for attempt in range(max_attempts):
        try:
            async with httpx.AsyncClient(timeout=30) as client:
                response = await client.post(
                    endpoint,
                    json=payload,
                    headers={
                        "X-Event-ID": event_id,
                        "X-Signature-256": compute_signature(payload)
                    }
                )
                if response.status_code < 500:
                    return  # Success or non-retriable failure
        except httpx.RequestError:
            pass  # Network error, retry

        if attempt < max_attempts - 1:
            delay = base_delay * (2 ** attempt)  # Exponential backoff
            await asyncio.sleep(min(delay, 3600))  # Cap at 1 hour

    # All retries exhausted
    await move_to_dead_letter_queue(event_id, payload)

Idempotency: Handling Duplicate Deliveries

Distributed systems guarantee at-least-once delivery, not exactly-once. Retries, network partitions, and provider restarts all produce duplicate webhook deliveries. Your processing code must be idempotent: processing the same event multiple times produces the same result as processing it once.

Idempotency key implementation:

Every webhook event has a unique identifier (event_id, idempotency_key, or similar). Before processing:

  1. Check if the event ID has been processed (check in a database or cache)
  2. If already processed, return HTTP 200 (acknowledge without re-processing)
  3. If not processed, process the event and record the event ID atomically
-- PostgreSQL upsert pattern for idempotent processing
INSERT INTO processed_webhooks (event_id, processed_at, result)
VALUES ($1, NOW(), $2)
ON CONFLICT (event_id) DO NOTHING;

-- Only proceed with business logic if the insert succeeded (returned 1 row)

Redis implementation for high-throughput scenarios:

async def process_idempotent(event_id: str, process_fn):
    key = f"webhook:processed:{event_id}"
    # SET NX (set if not exists) with 24-hour expiry
    if await redis.set(key, "1", nx=True, ex=86400):
        await process_fn()
    # If SET returned None, event was already processed

Idempotency scope: The idempotency check must be scoped to the semantic operation, not just the HTTP request. A payment notification delivered twice should result in one payment recorded, not two — regardless of how many times the endpoint received the payload.

Dead Letter Queue Design

After all retry attempts are exhausted, the failed event must not disappear. Dead letter queues (DLQ) capture events that could not be processed and provide:

  • Persistence: Events are not lost when processing permanently fails
  • Observability: DLQ depth is a critical operational metric
  • Recovery: When the root cause is fixed, DLQ events can be replayed

DLQ event structure (store alongside the original payload):

{
  "event_id": "evt_123",
  "original_payload": {...},
  "provider": "stripe",
  "endpoint": "/webhooks/payment",
  "failure_reason": "OrderService unavailable after 8 attempts",
  "attempt_count": 8,
  "first_attempt_at": "2026-04-01T14:30:00Z",
  "last_attempt_at": "2026-04-01T17:45:00Z",
  "error_details": "ConnectionRefusedError: [Errno 111]"
}

DLQ alerting: Alert when DLQ depth exceeds a threshold. A single event in the DLQ may be a data anomaly; ten events indicate a systemic processing problem. Configure PagerDuty/OpsGenie alerts on DLQ size metrics.

DLQ replay workflow: After fixing the root cause, replay DLQ events. Replaying requires that processing is idempotent (see previous section) — otherwise replay doubles the effects of successful past deliveries.

Event Ordering and Consistency

Webhooks do not guarantee delivery order. A payment.updated event may arrive before the payment.created event it references. Design processing to handle out-of-order events:

Optimistic creation: When processing a webhook that references an entity that does not yet exist in your system, create it with the information available. When the "earlier" event arrives, update the entity.

Event sequence numbers: Some providers include sequence numbers in payloads. Use these to detect gaps and request missing events via the provider's API.

Database constraints: Write processing logic defensively. Use upsert operations rather than insert operations. Use optimistic locking for concurrent update scenarios.

Observability

Webhook processing failures are invisible without explicit instrumentation. Key metrics to track:

Delivery metrics:

  • Webhook delivery rate (successful acknowledgments / total deliveries)
  • Processing success rate (successfully processed / successfully acknowledged)
  • DLQ depth by event type
  • P95 processing latency by endpoint

Business metrics:

  • Event processing lag (time from provider delivery to business effect)
  • DLQ event age (oldest event in the DLQ indicates how stale unprocessed events are)

Logging requirements (structured log per webhook event):

{
  "event_id": "evt_123",
  "event_type": "payment.succeeded",
  "provider": "stripe",
  "received_at": "2026-04-01T14:30:00.123Z",
  "signature_valid": true,
  "processing_started_at": "2026-04-01T14:30:00.125Z",
  "processing_completed_at": "2026-04-01T14:30:00.890Z",
  "processing_duration_ms": 765,
  "result": "success",
  "idempotency_status": "new"
}

Structured logs with consistent event IDs enable correlating delivery records with processing records across log streams.

Schema Evolution

Provider webhook schemas change over time. Fields are added, removed, or renamed. Defensive parsing prevents breaking changes from propagating to processing failures:

  • Accept unknown fields: Do not fail on unexpected fields (additive changes are backward compatible)
  • Use explicit field extraction: Extract only the fields your code uses; ignore everything else
  • Version headers: Many providers include a schema version header. Log the version; alert on unexpected versions.
  • Test with schema samples: Maintain a library of sample payloads from the provider's documentation. Run them through your parsing and processing code in CI.

Conclusion

Production-grade webhook integration requires seven engineering components beyond a simple POST handler: HMAC signature verification, idempotency with event deduplication, retry with exponential backoff, dead letter queueing, delivery acknowledgment decoupled from processing, observability instrumentation, and defensive schema parsing.

Implementing each correctly is straightforward. The failure modes of not implementing them — duplicate order fulfillment, silent event loss, replay attacks, and unobservable processing failures — have direct business consequences. The investment in correct webhook integration pays back in reduced incident rates and consistent system behavior.

For the broader event-driven architecture patterns that webhooks are one component of, see Software Architecture Patterns. For the API design patterns that complement webhook integration, see the related guides on authentication and security practices.

Testing Webhook Integration

Webhook testing requires a different approach than standard API testing — you do not control when the provider sends events.

Local development with tunneling: ngrok and similar tunneling tools expose a local development server to the public internet via a temporary URL. This allows real webhook deliveries from providers to reach your local development environment. Register the tunnel URL as your webhook endpoint in the provider's dashboard.

Webhook simulation in CI: Store a library of representative event payloads from the provider's documentation. In CI, send these payloads directly to the webhook endpoint handler function (bypassing HTTP) to test parsing, validation, and processing logic. This approach is faster and more reliable than trying to trigger real provider events in tests.

Contract testing: As your webhook processing logic evolves, the schema of expected payloads may drift from what the provider actually sends. Pact or similar contract testing tools allow you to define the expected webhook payload schema and verify that both provider and consumer agree on the contract. This catches breaking changes before they reach production.

Chaos testing for retry scenarios: Simulate endpoint unavailability by returning 5xx responses for a defined number of requests, then verifying that the provider's retry behavior delivers the event and that your idempotency key prevents duplicate processing. Most providers have sandbox environments where retry behavior can be tested safely.

Webhook Security Beyond HMAC

HMAC signature verification is necessary but not sufficient for webhook security in high-security applications.

IP allowlisting: Major providers publish the IP ranges used for webhook delivery. Configure firewall rules or application-level IP filtering to reject webhook requests from unexpected sources. This is a defense-in-depth measure — IP addresses can be spoofed, so it does not replace signature verification.

Rate limiting on webhook endpoints: Webhook endpoints should have rate limiting to prevent denial-of-service through webhook flooding. Limit by source IP and by time window. A legitimate provider delivering thousands of webhooks per second is either experiencing an incident or testing improperly.

TLS certificate pinning for high-security applications: Verify the provider's TLS certificate against a pinned public key. This prevents man-in-the-middle attacks in environments where the network is not fully trusted.

Secrets rotation: Webhook signing secrets should be rotatable without downtime. Implement a dual-secret rotation strategy: accept events signed with either the old or new secret during the rotation window, then retire the old secret after confirming all active subscriptions have been updated. Most providers support this pattern directly.

Related Articles

August 11, 2026

MLOps Guide: Taking Machine Learning Models to Production [2026]

87% of machine learning models built by data science teams never reach production. The models work — they pass cross-validation, they score well on holdout sets, they demonstrate genuine predictive value. The problem is not the modeling. The problem is everything that happens between a notebook experiment and a reliable, monitored, production system. MLOps is the discipline that closes that gap. This guide covers the full MLOps stack: maturity levels, tooling choices (MLflow, DVC, Kubeflow

Read More
August 10, 2026

LLM Fine-Tuning Guide: Custom Model Training with LoRA and QLoRA [2026]

General-purpose LLMs are impressive. They can write code, summarize documents, answer questions, and translate between languages with reasonable accuracy. But "reasonable" is not good enough when your application requires consistent output format, domain-specific terminology, a particular tone, or behavior that the base model was never trained to exhibit. That gap is where fine-tuning matters. Fine-tuning updates a model's weights on your specific data, changing how the model behaves — not

Read More
August 9, 2026

Computer Vision Applications: Object Detection, OCR, and Industrial AI [2026]

Computer vision has moved well past the research phase. The models are trained, the frameworks are mature, the hardware is accessible, and the use cases are generating measurable returns. What was a specialized capability requiring deep expertise in 2018 is now deployable infrastructure — if you know which component to reach for and where the real complexity lives. This guide covers computer vision applications across industrial, medical, logistics, and document processing domains. It expl

Read More