smaple.tr
process mining

Process Mining: Celonis, Event Logs, and Process Discovery Guide [2026]

Mehmet Kurtipek
March 13, 2026
10 min read
process mining
Celonis
event logs
process discovery
conformance checking

87% of organizations believe their core processes run as designed. Process mining analysis of the actual event log data typically reveals a different picture: average cycle times 40–60% longer than assumed, conformance rates below 70%, and bottlenecks concentrated in steps that managers consider routine. Process mining bridges the gap between how processes are supposed to work and how they actually work.

This guide covers the complete process mining methodology: event log structure, discovery algorithms, conformance checking, bottleneck analysis, Celonis vs PM4Py selection, and the integration patterns that make process mining an ongoing operational capability rather than a one-time analysis. By the end, you will understand how to extract, analyze, and act on process data from enterprise systems.

What Process Mining Reveals

Process mining extracts process behavior from event logs — the timestamped records that enterprise systems (ERP, CRM, BPM, hospital information systems) generate as they execute business processes.

Traditional process analysis depends on stakeholder interviews, shadowing observations, and documented SOPs. These methods reveal the intended process. Process mining reveals the actual process — including the variants, exceptions, and timing patterns that never appear in the documentation.

Typical discoveries from process mining:

  • A purchase-to-pay process with 47 documented variants versus the 3 variants in the official SOP
  • An invoice approval process where 23% of invoices are reopened after initial approval (rework cycle)
  • Average time-in-queue for a specific approval step: 4.8 days (versus the assumed "same day")
  • 12% of cases follow a path that includes a manual override of the standard routing rule

These discoveries are not visible from system reports or stakeholder interviews. They emerge only from analyzing the complete case history across all instances.

Event Log Structure

Process mining requires an event log: a structured record of activity occurrences tied to cases.

Minimum event log schema

Every event must have three elements:

Case ID: The unique identifier for the process instance (order number, patient ID, invoice number, loan application ID). All events with the same case ID belong to the same process instance.

Activity: The action that occurred (e.g., "Invoice Received", "Approval Requested", "Payment Processed"). Activity names should be consistent — the same action should always use the same name.

Timestamp: When the activity occurred. Granularity depends on the process — transaction processing may require millisecond precision; monthly financial close can use day-level timestamps.

Extended event attributes

Optional attributes enrich the analysis:

  • Resource: Who or which system executed the activity (user ID, system name, department)
  • Cost: Transaction cost at this activity step
  • Duration: Time the activity itself took (vs the waiting time before it)
  • Outcome: Result of the activity (approved, rejected, flagged)
  • Data attributes: Any process-relevant data captured at this event

Event log sources

ERP systems are the richest process mining source — SAP, Oracle, and Dynamics generate comprehensive event logs for financial and supply chain processes. The SAP Change Document tables (CDHDR/CDPOS) capture field-level changes with timestamps and user information.

CRM systems log sales activities, opportunity stage transitions, and customer interaction history. Hospital information systems (HIS) log patient journey events. Workflow engines (Camunda, ServiceNow, Pega) log task-level events for BPM processes.

Extraction approach: Most process mining projects require custom SQL queries to extract and transform raw event data into the required schema. This extraction is typically the most time-consuming step in a process mining project.

Process Mining Types

Process discovery

Discovery algorithms construct a process model from the event log without any prior knowledge of the intended process. The output is a graphical representation of how the process actually behaves — which activities occur, in which sequences, with what frequencies.

Discovery algorithms:

Alpha algorithm: The original discovery algorithm. Simple and fast but sensitive to noise and does not handle loops well. Used for educational purposes and simple processes.

Heuristic Miner: More noise-tolerant than Alpha. Uses frequency thresholds to distinguish genuine process paths from noisy behavior. Practical for most real-world event logs.

Inductive Miner: Produces sound process models (no deadlocks, complete replay fitness) by finding the dominant behavior and handling infrequent paths as exceptions. The recommended algorithm for most process mining projects.

Fuzzy Miner: Designed for highly complex processes with many variants. Simplifies the model by aggregating infrequent behavior, producing a more readable visualization.

Conformance checking

Conformance checking measures how closely the actual process (event log) matches a reference model (the intended process). Two conformance metrics:

Fitness: What percentage of cases in the log can be fully replayed on the reference model? A fitness of 0.95 means 95% of cases follow paths that are valid according to the model.

Precision: How much behavior allowed by the reference model is actually observed in the log? Low precision indicates an overly permissive model that allows paths never taken in practice.

Token replay: The most common conformance technique. Each token represents a case. The token "plays through" the model following the event log sequence. Tokens that cannot complete replay identify non-conforming behavior. The conformance score is the ratio of successfully replayed tokens.

Alignment-based conformance: More accurate than token replay. Finds the minimum number of log moves (activities in log but not in model) and model moves (activities in model but not in log) required to align each case with the reference model.

Performance analysis

Performance analysis adds time and resource dimensions to the process model.

Waiting time vs execution time: Cases spend most of their elapsed time waiting between activities, not in active execution. Process mining separates these two time components for each transition in the process. The waiting time distribution (median, 90th percentile, maximum) reveals SLA risk and bottleneck locations.

Bottleneck identification: The transitions with highest median waiting time or highest case backlog are the process bottlenecks. Improving these transitions produces the largest cycle time reduction.

Resource analysis: Which users or teams handle the most cases? Where is the load concentrated? Do high-load resources show higher error rates or longer processing times? Resource analysis informs staffing decisions and routing optimization.

Predictive process monitoring

Machine learning applied to process event streams enables real-time prediction of case outcomes.

Remaining time prediction: Given a case in progress, how long until it completes? Models trained on historical case trajectories predict remaining time based on current state (activity reached, time elapsed, resource assigned, case attributes).

Outcome prediction: Will this case result in an SLA breach? Will this loan application be approved or rejected? Predictive models flag high-risk cases early, enabling proactive intervention.

Next activity prediction: Given the current state of a case, what is the most likely next activity? Anomaly detection uses this prediction to flag cases following unusual paths.

Celonis: Enterprise Process Mining Platform

Celonis is the market-leading enterprise process mining platform. It connects directly to enterprise system databases, builds event logs automatically from raw transaction tables, and provides a full analysis environment.

Celonis architecture

Extractor: Connects to source systems (SAP, Oracle, Salesforce) and extracts event log data. Pre-built extractors for common enterprise systems reduce extraction development time.

Event log processing: Celonis transforms raw database extracts into conforming event logs, handles data quality issues, and manages incremental updates.

Process graph analytics: Celonis creates the process graph and calculates all discovery, conformance, and performance metrics.

OCPM (Object-Centric Process Mining): Celonis's 2025 capability that mines processes involving multiple interacting objects (orders, items, invoices) simultaneously — more realistic than single-object case analysis.

Execution Management System (EMS): Action flows that trigger real-time interventions when process monitoring detects anomalies. A detected SLA risk triggers a notification, reassignment, or workflow escalation automatically.

Celonis use cases by function

Accounts Payable: Process mining reveals invoice processing variants, approval bottlenecks, duplicate payment risk, and payment timing patterns. Celonis EMS triggers early payment discount capture when liquidity allows.

Order-to-Cash: Analyzes the full cycle from order receipt to cash collection. Identifies customers with consistent payment delays, order error patterns, and credit check bottlenecks.

Purchase-to-Pay: Compliance monitoring (are procurement policy rules being followed?), vendor performance (on-time delivery by vendor), and contract utilization (what percentage of purchases follow preferred vendor contracts?).

Healthcare (patient flow): Patient journey analysis from registration through discharge. Identifies care pathway variations, discharge delay patterns, and readmission risk factors.

Celonis pricing

Celonis pricing is enterprise-tier and not publicly listed. Published analyst estimates range from $50,000 to $500,000+ annually depending on data volume, number of processes mined, and EMS action volume. For large enterprises with high-value processes (AP, O2C, P2P), this investment is justified by the identified savings. For smaller organizations or targeted analytical use, PM4Py is a more accessible starting point.

PM4Py: Open-Source Process Mining

PM4Py is the leading open-source Python library for process mining. It provides discovery, conformance, and performance analysis with a code-based interface.

Installation and basic usage

import pm4py

# Load event log from CSV
log = pm4py.read_xes('event_log.xes')

# Or from DataFrame
import pandas as pd
df = pd.read_csv('events.csv')
log = pm4py.format_dataframe(df, case_id='case_id', activity_key='activity', timestamp_key='timestamp')

# Discover process model
process_tree = pm4py.discover_process_tree_inductive(log)
petri_net, initial_marking, final_marking = pm4py.convert_to_petri_net(process_tree)

# Visualize
pm4py.view_petri_net(petri_net, initial_marking, final_marking)

# Conformance checking
fitness = pm4py.fitness_token_based_replay(log, petri_net, initial_marking, final_marking)
print(f"Fitness: {fitness['average_trace_fitness']:.2%}")

# Performance analysis
performance = pm4py.get_all_case_durations(log)
print(f"Median case duration: {sorted(performance)[len(performance)//2]/3600:.1f} hours")

PM4Py capabilities

  • Process discovery (Inductive Miner, Heuristic Miner, Alpha Miner, Fuzzy Miner)
  • Petri net, process tree, and BPMN visualization
  • Token replay and alignment-based conformance
  • Performance analysis (case duration, waiting times, service times)
  • Dotted chart visualization for case arrival patterns
  • Social network analysis (handover of work, joint activity)
  • Decision mining (which attributes predict routing at decision points?)

PM4Py vs Celonis

Dimension PM4Py Celonis
Type Open-source Python library Enterprise SaaS platform
Interface Code-based Visual, no-code
Data volume Millions of events Billions of events
ERP connectors Custom extraction required Pre-built extractors
Real-time monitoring Requires custom engineering Built-in (EMS)
Cost Free $50K–500K/year
Learning curve Python competency required 1–2 weeks for analysts
Best for Custom analysis, pilots, research Enterprise operational use

Implementation Path

Step 1: Process selection. Choose a process with high business impact, accessible event log data, and a specific performance question to answer. Accounts payable, order management, and customer onboarding are common starting points.

Step 2: Event log extraction. Extract raw event data from the source system. Transform to case-activity-timestamp schema. Quality check: case completeness, activity consistency, timestamp accuracy.

Step 3: Discovery analysis. Run Inductive Miner on the extracted log. Examine the top 10 process variants by case frequency. Identify the process's actual happy path and the most common deviation paths.

Step 4: Conformance and performance analysis. Measure fitness against the reference process. Identify the waiting time distribution by transition. Locate the top 3 bottlenecks by waiting time.

Step 5: Action and measurement. Implement process improvements targeting the identified bottlenecks. Re-run process mining after 30–60 days to measure the impact.

Process mining is most valuable when it becomes an ongoing capability — not a one-time analysis but a continuous process monitoring function that feeds the improvement cycle.

Contact Smart Maple to design your process mining implementation.

Related Articles

August 11, 2026

MLOps Guide: Taking Machine Learning Models to Production [2026]

87% of machine learning models built by data science teams never reach production. The models work — they pass cross-validation, they score well on holdout sets, they demonstrate genuine predictive value. The problem is not the modeling. The problem is everything that happens between a notebook experiment and a reliable, monitored, production system. MLOps is the discipline that closes that gap. This guide covers the full MLOps stack: maturity levels, tooling choices (MLflow, DVC, Kubeflow

Read More
August 10, 2026

LLM Fine-Tuning Guide: Custom Model Training with LoRA and QLoRA [2026]

General-purpose LLMs are impressive. They can write code, summarize documents, answer questions, and translate between languages with reasonable accuracy. But "reasonable" is not good enough when your application requires consistent output format, domain-specific terminology, a particular tone, or behavior that the base model was never trained to exhibit. That gap is where fine-tuning matters. Fine-tuning updates a model's weights on your specific data, changing how the model behaves — not

Read More
August 9, 2026

Computer Vision Applications: Object Detection, OCR, and Industrial AI [2026]

Computer vision has moved well past the research phase. The models are trained, the frameworks are mature, the hardware is accessible, and the use cases are generating measurable returns. What was a specialized capability requiring deep expertise in 2018 is now deployable infrastructure — if you know which component to reach for and where the real complexity lives. This guide covers computer vision applications across industrial, medical, logistics, and document processing domains. It expl

Read More