smaple.tr
notification infrastructure

Notification Infrastructure: Building Multi-Channel Push, Email, SMS, and In-App Delivery [2026]

Mehmet Kurtipek
February 23, 2026
13 min read
notification infrastructure
push notifications
FCM APNs
email deliverability
SMS gateway
in-app notifications
fan-out architecture

Notification fatigue is killing user engagement — but poorly designed delivery systems are a larger problem than over-sending. When a user misses a critical security alert because it landed in spam, or receives a flood of marketing push notifications at 2 AM, the failure is architectural, not just strategic.

This guide covers the full notification infrastructure stack: multi-channel delivery pipelines, fan-out architectures, FCM/APNs integration patterns, email deliverability configuration, SMS gateway selection, in-app notification center design, throttling, preference management, and analytics. By the end, you will have a clear picture of what production-grade notification infrastructure looks like at each scale.

Notification Infrastructure: Core Architecture

Effective notification systems share a common architectural pattern regardless of channel: a central orchestration service routes events through channel-specific adapters, each backed by a message queue for buffering and retry.

The Fan-Out Problem

Fan-out — sending a single event to many recipients — is the hardest scaling challenge in notification infrastructure. When a price drops on a product that 50,000 users have saved, the system must enqueue 50,000 individual delivery tasks without overloading downstream channel APIs.

Three fan-out patterns handle this at scale:

Direct fan-out — the notification service writes one task per recipient directly to the delivery queue. Works well below 10,000 recipients but creates queue pressure at scale.

Segmented fan-out — the service writes a single "fan-out job" to a pre-processing queue. A worker expands it into per-recipient tasks and distributes them across channel queues. This decouples the write spike from the delivery pipeline.

Pull-based fan-out — used for real-time feeds. Recipients subscribe to a topic and pull their own notification stream. Reduces write load but increases read infrastructure complexity.

In aggregator and marketplace platforms at Smart Maple, we use segmented fan-out for batch campaigns and direct fan-out for transactional events. The boundary is roughly 1,000 recipients: above that, segmented fan-out prevents queue saturation.

Channel Routing and Fallback

Every notification type maps to a preferred channel and fallback chain:

Notification type Primary channel Fallback
Security alerts (password reset, suspicious login) Push + SMS Email
Transactional (order confirmation, shipping update) Email + Push —
Marketing promotions Email Push (if opted in)
In-app activity updates In-app Push (if backgrounded)

Fallback logic must be explicit. If a push token is invalid, the system should not silently fail — it should attempt SMS or email based on the user's preference profile, then log the fallback path for analytics.

Push Notifications: FCM and APNs

Firebase Cloud Messaging (FCM)

FCM delivers push notifications to Android and iOS devices through a single API. The production integration patterns that matter:

Token lifecycle management — device tokens change on app reinstall, device migration, and periodic rotation. An InvalidRegistration or NotRegistered error on send means the token is stale. Tokens receiving these errors must be deleted from your database immediately. Background token cleanup at scale: run a nightly job against your FCM send log to identify tokens with no successful delivery in 30 days.

Topic messaging vs direct targeting — FCM topics allow broadcasting to all subscribers without maintaining a recipient list. Topics are appropriate for global announcements but not for personalized notifications. For user-specific delivery, maintain your own token-to-user mapping and send direct.

Notification vs data messages — notification payloads are rendered by the OS automatically when the app is backgrounded. Data payloads are passed to your app code for custom handling. In foreground, both go to app code. The practical implication: if you need custom notification behavior (update a badge, decrypt content before display), use a data message and handle rendering yourself.

HTTP v1 API — Google deprecated the legacy FCM API in 2024. The HTTP v1 API requires OAuth 2.0 service account authentication. All new integrations should use v1; legacy API integrations should migrate before the June 2024 deprecation deadline.

Apple Push Notification Service (APNs)

APNs imposes stricter rules than FCM. The key patterns:

Token vs certificate authentication — APNs supports both. Token-based authentication (using a .p8 key file) does not expire and works across environments. Certificate-based authentication expires annually and is environment-specific. New integrations should use token authentication exclusively.

Permission timing strategy — iOS requires explicit opt-in for push notifications. Asking for permission on first launch produces 40-60% decline rates industry-wide. Asking after a user completes a meaningful action (first order, first booking) doubles opt-in rates. Show a pre-permission prompt explaining the value before triggering the system dialog.

Notification service extensions — NSE allows modifying a notification's content before it is displayed: decrypt end-to-end encrypted messages, download media attachments, update badge counts based on live data. Extensions run for a maximum of 30 seconds; after that, the unmodified payload displays.

Provisional authorization — iOS 12+ allows quiet test delivery to the Notification Center without an explicit opt-in prompt. Use this to deliver value before asking for permission.

Rich Push and Interactive Notifications

Rich push (images, video, carousels) increases tap-through rates by 40-50% over text-only notifications. Implementation:

  • iOS: mutable-content: 1 triggers your NSE to attach media before display. Maximum attachment size is 10 MB for images, 50 MB for video.
  • Android: BigPictureStyle and BigTextStyle in the notification payload provide expanded views.

Action buttons allow users to act without opening the app. "Confirm appointment", "Mark as read", "Add to cart" — these increase conversion on transactional notifications. iOS allows up to 4 action buttons; Android supports up to 3.

Email Notification Delivery

Transactional Email Providers

SendGrid, Amazon SES, Postmark, Mailgun, and Resend each target different use cases:

Provider Strength Ideal for
Postmark Highest deliverability focus Transactional email only
Amazon SES Lowest unit cost High-volume, cost-sensitive
SendGrid Full-featured marketing + transactional Mixed use cases
Resend Developer-first API, modern DX Greenfield projects

For transactional email specifically, Postmark's deliverability-first architecture (separate IP pools, no marketing email allowed) produces consistently high inbox placement rates.

Email Deliverability Configuration

Three DNS records are mandatory for production email:

SPF (Sender Policy Framework) — a TXT record listing which servers are authorized to send from your domain. Without SPF, receiving servers have no way to verify your send origin.

DKIM (DomainKeys Identified Mail) — adds a cryptographic signature to every email. The receiving server verifies the signature against your public key published in DNS. DKIM proves the message content was not modified in transit.

DMARC (Domain-based Message Authentication, Reporting & Conformance) — a policy that tells receiving servers what to do when SPF or DKIM fails: none (monitor only), quarantine (deliver to spam), or reject (block delivery). Start with p=none to collect reports, then advance to quarantine and reject as you verify your sending infrastructure.

IP warming — new sending IPs have no reputation. Start at 50 messages/day and double weekly. Sending 100,000 messages from a cold IP guarantees spam classification.

Bounce handling — hard bounces (invalid address) must remove the address immediately. Soft bounces (mailbox full, server busy) get up to 3 retries over 72 hours; after that, treat as hard bounce. Continuing to send to hard bounce addresses tanks your sender reputation.

Email Template Architecture

Email rendering differs fundamentally from browser rendering. CSS support is inconsistent across clients — Gmail strips <style> blocks from the <head>, Outlook uses Word's rendering engine, and Apple Mail is comparatively modern. Practical rules:

  • Use table-based layout for structure, not CSS flexbox or grid
  • Inline critical CSS; external stylesheets are unreliable
  • Test in Litmus or Email on Acid before production sends
  • MJML or Foundation for Emails generates cross-client compatible HTML from modern template syntax

SMS Notification Infrastructure

SMS Gateway Selection

Key evaluation criteria for SMS providers:

Coverage — Twilio covers 180+ countries with direct carrier connections. Vonage (Nexmo), MessageBird, and Sinch are comparable. For APAC or Africa coverage, evaluate local aggregators alongside global providers.

Sender ID — many markets (UK, Australia, India) allow alphanumeric sender IDs ("SMARTMAPLE"). The US does not support alphanumeric sender IDs for A2P messaging — you need a dedicated 10DLC number or short code.

A2P 10DLC registration (US) — since 2022, all US application-to-person SMS requires carrier registration. Unregistered traffic is filtered. Registration requires brand and campaign approval; processing takes 2-6 weeks.

Delivery receipts — critical for reliability monitoring. Configure provider webhooks to receive DLR (delivery receipt) callbacks for every message. Track delivered/failed rates per carrier.

Character Encoding and Cost Optimization

SMS pricing is per-segment. GSM-7 encoding: 160 characters per segment. Unicode (required for non-Latin characters, emoji): 70 characters per segment. A 161-character message in Unicode = 2 segments = double cost.

Cost optimization techniques:

  • Avoid emoji in SMS unless essential (forces Unicode, halves character limit)
  • URL shorteners reduce character count for link-heavy messages
  • Consolidate multiple short messages into a single longer one when context allows
  • Push-first with SMS fallback reduces SMS volume without reducing delivery coverage

In-App Notification Center

Real-Time Delivery Architecture

In-app notifications require a persistent connection or polling mechanism:

WebSocket — bidirectional persistent connection. Best for high-frequency updates (chat, live feeds). Requires stateful server infrastructure and connection management at scale.

Server-Sent Events (SSE) — unidirectional push from server to client. Lower infrastructure complexity than WebSocket. Appropriate for notification delivery where the client does not need to send data back on the same connection.

Long polling — fallback for environments that block WebSocket or SSE. Client makes a request; server holds it open until an event arrives or timeout. Higher latency and resource consumption than SSE.

For notification centers specifically, SSE is the right default. Implement WebSocket only when you need bidirectional communication (e.g., a real-time chat component on the same page).

Notification Center Data Model

Notification {
  id: uuid
  user_id: string
  type: enum (order_update | security | promotion | system)
  title: string
  body: string
  action_url: string
  read: boolean
  read_at: timestamp
  created_at: timestamp
  metadata: json
}

Unread count computation: cache per-user unread count in Redis. Increment on notification insert; decrement (or recompute) on mark-as-read. Never compute unread count from a full table scan on every page load.

Notification grouping: "You have 5 new comments on your post" is better UX than 5 individual notifications. Group by (user_id, type, reference_id) within a time window (e.g., 1 hour) and update the group's count rather than inserting new rows.

Preference Management and Compliance

User Preference Architecture

A complete preference model covers four dimensions:

Channel preferences — push enabled/disabled, email enabled/disabled, SMS enabled/disabled per notification category.

Category preferences — marketing, transactional, security, system. Security notifications are typically non-dismissable; marketing is opt-in by default.

Frequency preferences — immediate, daily digest, weekly digest per category.

Quiet hours — time range (e.g., 10 PM – 8 AM local time) during which non-critical notifications are queued rather than delivered. Requires storing user timezone.

Store preferences in a preferences table (not on the user object) for efficient partial updates. A preference change should propagate to all channels within one delivery cycle.

GDPR (Article 6 consent), CAN-SPAM, CASL, and ePrivacy Directive require explicit opt-in for marketing communications and a one-click unsubscribe mechanism.

Required implementation:

  • List-Unsubscribe header (RFC 8058) on all marketing email — enables one-click unsubscribe from email clients without requiring the user to visit your site
  • STOP opt-out for SMS — all marketing SMS must recognize and process STOP replies
  • Push notification categories in iOS — respect system-level permission revocations within 24 hours

Transactional notifications (password reset, order confirmation, security alerts) do not require marketing opt-in. But misclassifying promotional content as transactional is a compliance risk.

Throttling and Smart Scheduling

Notification Throttling Architecture

Throttling prevents notification fatigue and protects downstream channel API rate limits:

Per-user rate limits — maximum N notifications of category X per time window. Example: no more than 3 marketing push notifications per day per user. Implement with a sliding window counter in Redis.

Global channel rate limits — FCM, APNs, and SMS providers impose per-second or per-minute send limits. Respect these with a token bucket implementation in your queue consumer.

Priority queues — security notifications bypass throttling. Marketing notifications drain last. Implement with separate queues or a priority field on the notification record.

Deduplication — prevent identical notifications from sending twice (duplicate event triggers, retry logic). Use a (user_id, type, reference_id, window) uniqueness check before enqueuing.

Time-Zone-Aware Scheduling

Calculate delivery time in the user's local timezone. For daily digest: send at 9 AM local time, not 9 AM UTC. For immediate delivery: check quiet hours before enqueue; if in quiet hours, schedule for the next allowed window.

Machine learning scheduling (predicting individual open-rate peaks) is available in some providers (Braze, Iterable) but adds complexity. For most use cases, timezone-based scheduling plus quiet hours produces most of the benefit at a fraction of the cost.

Analytics and A/B Testing

Core Notification Metrics

Track per-channel, per-category, per-segment:

Metric Formula Signal
Delivery rate delivered / sent Infrastructure health
Open rate opened / delivered Content relevance
Click-through rate clicked / opened CTA effectiveness
Conversion rate converted / clicked Business value
Opt-out rate unsubscribed / delivered Fatigue signal

An opt-out rate above 0.3% per campaign is a warning signal. Above 0.5% triggers deliverability problems with email providers.

A/B Test Variables

Variables worth testing in sequence: subject line / push title (highest impact, fastest to test), send time, CTA copy, content length, channel mix (email vs push for the same event). Do not run A/B tests on multiple variables simultaneously — it makes attribution impossible.

Statistical significance requirement: p < 0.05. For a 10% open rate, you need roughly 1,500 recipients per variant to detect a 2-percentage-point difference with 80% power. Run tests to completion — stopping early when one variant leads produces false positives at high rates.

Production Architecture Summary

A production-ready notification system requires:

  1. Central orchestration service with channel routing logic
  2. Per-channel adapter queue (push queue, email queue, SMS queue, in-app queue)
  3. Fan-out worker for high-volume broadcast events
  4. Preference store with real-time propagation
  5. Throttling layer (Redis sliding window per user per category)
  6. Dead letter queue for failed deliveries with retry and alerting
  7. Delivery webhook handler for DLR and bounce callbacks
  8. Analytics pipeline tracking per-message outcomes

The threshold for needing dedicated notification infrastructure (vs a simple email library call): roughly 1,000 monthly active users. Above that, delivery observability and preference management justify the investment. At 100,000 MAU, fan-out architecture and throttling become requirements.

Related Articles

August 11, 2026

MLOps Guide: Taking Machine Learning Models to Production [2026]

87% of machine learning models built by data science teams never reach production. The models work — they pass cross-validation, they score well on holdout sets, they demonstrate genuine predictive value. The problem is not the modeling. The problem is everything that happens between a notebook experiment and a reliable, monitored, production system. MLOps is the discipline that closes that gap. This guide covers the full MLOps stack: maturity levels, tooling choices (MLflow, DVC, Kubeflow

Read More
August 10, 2026

LLM Fine-Tuning Guide: Custom Model Training with LoRA and QLoRA [2026]

General-purpose LLMs are impressive. They can write code, summarize documents, answer questions, and translate between languages with reasonable accuracy. But "reasonable" is not good enough when your application requires consistent output format, domain-specific terminology, a particular tone, or behavior that the base model was never trained to exhibit. That gap is where fine-tuning matters. Fine-tuning updates a model's weights on your specific data, changing how the model behaves — not

Read More
August 9, 2026

Computer Vision Applications: Object Detection, OCR, and Industrial AI [2026]

Computer vision has moved well past the research phase. The models are trained, the frameworks are mature, the hardware is accessible, and the use cases are generating measurable returns. What was a specialized capability requiring deep expertise in 2018 is now deployable infrastructure — if you know which component to reach for and where the real complexity lives. This guide covers computer vision applications across industrial, medical, logistics, and document processing domains. It expl

Read More