smaple.tr
big data

Big Data Analytics Platform: Spark, Kafka and Data Lake Architecture [2026]

Mehmet Kurtipek
February 21, 2026
11 min read
big data
Apache Spark
Kafka
data lake
real-time analytics

At terabyte scale, traditional data warehouse architecture collapses under write volume, query latency, and operational complexity. Netflix processes petabytes of event data daily. Uber's real-time pricing engine consumes millions of location events per minute. The infrastructure enabling this is not exotic — it is a well-understood combination of Apache Spark for distributed computing, Apache Kafka for event streaming, and a data lake architecture that separates raw storage from analytics-ready data.

This guide covers the full big data analytics platform stack: data scale categories, Spark execution model, Kafka topic design, Lambda vs Kappa architecture decisions, the medallion data lake pattern, and production cost modeling. By the end, you will have a clear architecture map and the decision criteria to choose the right approach for your data volume and latency requirements.

Big Data Analytics Platform: Scale Categories and Architecture Fit

The first architectural decision is whether you actually have a big data problem. Applying distributed computing to a dataset that fits on a single machine adds complexity without benefit.

Scale Category Volume Right Tool Processing Time
Small Data Under 100 GB PostgreSQL, SQLite Under 1 hour, single machine
Medium Data 100 GB – 1 TB Dedicated warehouse, Power BI 1–4 hours, single server
Big Data 1 TB – 100 TB Apache Spark, Hadoop HDFS 1–10 minutes, distributed cluster
Massive Data Over 100 TB Spark on Kubernetes, cloud warehouses Auto-scaling, continuous

Most enterprise analytics systems live in the medium-to-big data boundary. The inflection point where distributed computing pays off is typically when batch jobs exceed 4 hours on a single machine, or when real-time latency requirements fall below 5 seconds for aggregate queries.

Big data challenges cluster into six categories: volume (TB/PB scale storage), velocity (real-time event streams), variety (structured and unstructured mixed), processing speed (distributed computation), cost (infrastructure vs managed services), and operational complexity (cluster management, tuning).

Apache Spark: Distributed Computing Engine

Spark is the computation layer of a big data platform. It processes data in parallel across a cluster of machines, using a Directed Acyclic Graph (DAG) execution model that optimizes query plans before running them.

Spark Architecture

The Spark execution model has four components:

Driver (Master): Holds the SparkContext, creates RDDs and DataFrames, parallelizes transformations, and orchestrates execution across the cluster. The driver runs your application code and coordinates the distributed work.

Cluster Manager: Kubernetes, YARN, or Mesos — allocates resources to the Spark application. Kubernetes has become the dominant choice for cloud-native deployments.

Executors (Workers): The actual computation units. A production cluster runs 10 to 1,000+ executors depending on data volume. Each executor has assigned CPU cores (parallelism) and memory (for caching and computation).

Storage: HDFS, S3, GCS, or Azure Blob — the distributed file system that stores input data and intermediate results.

The Spark API has layers. The lowest level is RDD (Resilient Distributed Dataset) — immutable, distributed collections with fault tolerance via lineage tracking. Above this is the DataFrame API with the Catalyst query optimizer and Tungsten memory management engine. On top of DataFrames sit Spark SQL, Spark Streaming, Spark MLlib, and GraphX.

Why Spark Outperforms Hadoop MapReduce

Spark stores intermediate computation results in memory rather than writing them to disk between stages. For iterative algorithms (like machine learning training) this produces 10–100x speedups. For single-pass ETL jobs, the gain is smaller (2–5x) but the programming model is dramatically simpler.

In distributed systems work we have run at Smart Maple — including a healthcare scheduling optimization system processing thousands of route combinations daily — the difference between Spark's in-memory model and disk-based processing is not academic. Iterative optimization algorithms that would require hours of MapReduce processing complete in minutes with Spark.

Apache Kafka: Real-Time Event Streaming

Kafka decouples data producers from data consumers. Producers write events to topics without knowing who will consume them. Consumers read at their own pace without blocking producers. This decoupling is what makes a Kafka-based platform resilient to downstream slowdowns.

Kafka Architecture

DATA PRODUCERS
├─ Application logs
├─ Database CDC (Change Data Capture)
├─ API webhooks
├─ IoT sensors
└─ Real-time events
      ↓
KAFKA BROKER CLUSTER
├─ Topic (messaging channel)
│  ├─ Partition 0 (replicated across 3 brokers)
│  ├─ Partition 1 (replicated across 3 brokers)
│  └─ Partition N
│
└─ Retention Policy
   ├─ Time-based (e.g., 7 days)
   ├─ Size-based (e.g., 1 GB)
   └─ Compacted (keep latest value per key)
      ↓
KAFKA CONSUMERS
├─ Consumer Group 1: Spark Streaming
├─ Consumer Group 2: Real-time Dashboard
├─ Consumer Group 3: ML Feature Pipeline
└─ Consumer Group N

Critical Topic Configuration Parameters

Three parameters determine a topic's behavior under production load:

Partitions: Controls write parallelism and maximum consumer concurrency. A topic with 10 partitions can be consumed by at most 10 consumer instances in a group simultaneously. Under-partition a high-volume topic and you create a bottleneck. Over-partition and you add unnecessary coordination overhead. A rule of thumb: target 1 MB/s throughput per partition.

Replication Factor: 3 is the standard for production. With replication factor 3, a cluster tolerates 2 broker failures without data loss. Combined with min.insync.replicas=2 and acks=all, you get strong durability guarantees.

Retention: Time-based retention (7 days is common) enables consumer groups to replay recent events. For event sourcing patterns where you need full history, use log compaction rather than time-based deletion.

Lambda vs Kappa Architecture

The central architectural decision for a big data platform is whether to maintain separate batch and streaming pipelines (Lambda) or unify them under a single streaming pipeline (Kappa).

Lambda Architecture runs two parallel pipelines: a batch layer for historical accuracy (Spark batch jobs running nightly or hourly) and a speed layer for real-time results (Spark Streaming or Flink with sub-minute latency). A serving layer merges results from both layers for queries.

Pros: Batch layer provides high accuracy for historical analysis. Speed layer provides low latency for operational queries. Cons: You maintain two separate codebases for the same business logic. When logic changes, both must be updated consistently. Operational overhead is significant.

Kappa Architecture eliminates the batch layer. All data flows through a single streaming pipeline. Historical reprocessing is achieved by replaying events from Kafka (Kafka's configurable retention enables this).

Pros: Single codebase for all data processing logic. Simpler operations. Cons: Reprocessing large historical datasets through a streaming framework requires careful resource management.

Decision heuristic: If your analytics latency requirement is under 1 minute and you can replay from Kafka, Kappa is simpler. If you need batch backfills from external sources that were never in Kafka, or if regulatory accuracy requirements mandate periodic reconciliation, Lambda provides the safety net.

Data Lake Architecture: The Medallion Pattern

A data lake without structure becomes a data swamp — unqueryyable, unreliable, expensive. The Medallion Architecture imposes three quality layers on raw storage.

Bronze Layer (Raw Data): Data stored exactly as received, no transformations. Format: Parquet (columnar, compression-efficient). Retention: 90 days. Purpose: Historical analysis, audit trail, raw replay capability. This layer is append-only and should never be modified.

Silver Layer (Validated and Enriched): Bronze data cleaned, deduplicated, validated against schema, and enriched with joins to reference tables. Format: Parquet or Delta Lake. Transformations include null handling, type casting, deduplication, and standardization. This layer is safe for feature engineering and ML training.

Gold Layer (Business-Ready): Aggregated metrics, KPIs, pre-joined tables optimized for BI consumption. Format: Delta Lake (ACID transactions). Aggregation levels: hourly, daily, monthly. Retention: 5+ years. This is what your dashboards and reporting tools query.

Delta Lake deserves specific mention. It adds ACID transaction semantics to object storage (S3, GCS), enabling reliable concurrent writes, schema enforcement, and time travel (query any previous version of a table). For production data lakes, Delta Lake or Apache Iceberg has replaced raw Parquet for Gold and Silver layers.

PySpark Batch Processing: Bronze to Gold

A standard PySpark batch job follows the Bronze → Silver → Gold transformation pattern:

from pyspark.sql import SparkSession
from pyspark.sql.functions import col, count, when, avg, to_date

spark = SparkSession.builder \
    .appName("analytics-batch") \
    .config("spark.sql.adaptive.enabled", "true") \
    .getOrCreate()

# 1. Read Bronze (raw, validated)
raw_df = spark.read.parquet("s3://data-lake/bronze/events/")

# 2. Silver: clean, deduplicate, enrich
silver_df = raw_df \
    .filter(col("event_id").isNotNull()) \
    .dropDuplicates(["event_id"]) \
    .withColumn("event_date", to_date(col("created_at"))) \
    .cache()

# 3. Gold: aggregate for BI consumption
gold_df = silver_df.groupBy("event_date", "dimension_id").agg(
    count("event_id").alias("total_events"),
    (count(when(col("status") == "completed", 1)) /
     count("event_id") * 100).alias("completion_rate_pct"),
    avg("duration_minutes").alias("avg_duration")
).orderBy("event_date")

# 4. Write Gold with Delta Lake
gold_df.write \
    .format("delta") \
    .mode("overwrite") \
    .option("mergeSchema", "true") \
    .save("s3://data-lake/gold/daily_metrics/")

Key Spark performance settings for production:

  • spark.sql.adaptive.enabled=true — AQE dynamically optimizes query plans at runtime, handles data skew automatically
  • spark.sql.shuffle.partitions — default 200 is too high for small datasets, too low for large ones; tune to 2x your executor cores
  • spark.sql.adaptive.coalescePartitions.enabled=true — merges small partitions after shuffle to reduce overhead

Spark Streaming: Real-Time Pipeline from Kafka

from pyspark.sql import SparkSession
from pyspark.sql.functions import from_json, col, window, count
from pyspark.sql.types import StructType, StructField, StringType, IntegerType, TimestampType

spark = SparkSession.builder \
    .appName("realtime-streaming-analytics") \
    .getOrCreate()

# Read from Kafka
kafka_df = spark.readStream \
    .format("kafka") \
    .option("kafka.bootstrap.servers", "kafka-cluster:9092") \
    .option("subscribe", "events.stream") \
    .option("startingOffsets", "latest") \
    .load()

# Define schema and parse JSON
schema = StructType([
    StructField("event_id", IntegerType()),
    StructField("entity_id", IntegerType()),
    StructField("status", StringType()),
    StructField("timestamp", TimestampType())
])

parsed_df = kafka_df.select(
    from_json(col("value").cast("string"), schema).alias("data")
).select("data.*")

# 5-minute sliding window aggregation
windowed = parsed_df.groupBy(
    window(col("timestamp"), "5 minutes", "1 minute"),
    "entity_id"
).agg(count("event_id").alias("event_count"))

# Write to Redis for real-time dashboard
query = windowed.writeStream \
    .format("redis") \
    .option("host", "redis-cache") \
    .option("checkpointLocation", "/tmp/spark-checkpoint") \
    .start()

query.awaitTermination()

Production Infrastructure and Cost Modeling

A realistic cost model for a medium-scale big data platform on AWS (500 entities, ~50M events/month):

Component Specification Monthly Cost
EMR Cluster (batch) 1x r5.4xlarge master + 20x r5.2xlarge workers $6,600
Kafka Cluster 3x m5.2xlarge brokers + 3x t3.xlarge ZooKeeper $1,350
Spark Streaming Kubernetes cluster $2,000
S3 Storage ~2 TB + 10 TB egress $950
ElastiCache Redis 2x cache.r6g.xlarge $800
Aurora PostgreSQL db.r5.2xlarge (warehouse) $2,500
Monitoring CloudWatch + DataDog $500
Total ~$14,700/month

Cost reduction strategies:

  • EMR Spot Instances for batch workers: 60–80% savings on worker nodes (risk: spot interruption)
  • S3 Intelligent-Tiering: automatic cost optimization for infrequently accessed data
  • Databricks alternative: For pure batch workloads, Databricks Community Edition or managed Databricks costs 70–85% less than a self-managed EMR cluster
  • Reserved Instances for baseline Kafka and streaming infrastructure: 40% savings vs on-demand

Common Big Data Platform Failures

Failure Pattern Root Cause Resolution
Out-of-memory errors Large shuffles without partition tuning Increase shuffle partitions, enable spill to disk
Slow batch jobs Unoptimized query plan, missing partitioning Use EXPLAIN on queries, add partition predicates
Growing Kafka consumer lag Consumer throughput below producer rate Scale consumer group, optimize processing logic
Data skew hotspots Uneven partition key distribution Salt keys, use broadcast joins for small tables
Cost explosion Executor count misconfigured Implement auto-scaling policies, use Spot/Preemptible

Conclusion

A production big data analytics platform requires three coordinated layers: event streaming (Kafka), distributed computation (Spark), and structured storage (data lake with medallion architecture). The Lambda/Kappa decision drives whether you maintain one or two processing pipelines. Delta Lake or Iceberg transforms a raw data lake into a reliable, queryable analytical store.

The entry cost (~$15K/month for medium scale) is significant, but the unit economics improve rapidly with scale: at 50M events/month, the cost-per-event approaches $0.04. For organizations generating value from data at this scale, the platform investment pays for itself through better operational decisions, faster analytics cycles, and eliminated manual reporting effort.


Author: Smart Maple Data Engineering Team Updated: April 2026

Related Articles

August 11, 2026

MLOps Guide: Taking Machine Learning Models to Production [2026]

87% of machine learning models built by data science teams never reach production. The models work — they pass cross-validation, they score well on holdout sets, they demonstrate genuine predictive value. The problem is not the modeling. The problem is everything that happens between a notebook experiment and a reliable, monitored, production system. MLOps is the discipline that closes that gap. This guide covers the full MLOps stack: maturity levels, tooling choices (MLflow, DVC, Kubeflow

Read More
August 10, 2026

LLM Fine-Tuning Guide: Custom Model Training with LoRA and QLoRA [2026]

General-purpose LLMs are impressive. They can write code, summarize documents, answer questions, and translate between languages with reasonable accuracy. But "reasonable" is not good enough when your application requires consistent output format, domain-specific terminology, a particular tone, or behavior that the base model was never trained to exhibit. That gap is where fine-tuning matters. Fine-tuning updates a model's weights on your specific data, changing how the model behaves — not

Read More
August 9, 2026

Computer Vision Applications: Object Detection, OCR, and Industrial AI [2026]

Computer vision has moved well past the research phase. The models are trained, the frameworks are mature, the hardware is accessible, and the use cases are generating measurable returns. What was a specialized capability requiring deep expertise in 2018 is now deployable infrastructure — if you know which component to reach for and where the real complexity lives. This guide covers computer vision applications across industrial, medical, logistics, and document processing domains. It expl

Read More