You can also find this library at CRAN and download it directly from R and RStudio.
LISTEN TO THIS POST AS A PODCAST:
Imagine you’re monitoring a manufacturing line, and the sensor readings suddenly look… different. The numbers aren’t obviously wrong — they’re within range — but something has shifted. Or picture an economist staring at decades of GDP data, trying to pinpoint exactly when a country’s growth model fundamentally changed. Or a doctor watching a patient’s vitals, waiting for the moment “stable” tips into “critical.”
All of these scenarios share the same underlying question: when did the system switch from one regime to another?
This is the problem of regime change detection (also called changepoint detection), and it’s one of the most practically important — and theoretically rich — problems in time series analysis. A new open-source R package called RegimeChange tackles this problem with unusual breadth, combining classical statistics, Bayesian inference, and even deep learning under a single, coherent interface. In this post, I’ll walk you through what it does, how it works, and why it matters — in plain language, without cutting corners on the technical details.
The Problem: Signal vs. Noise
Every observable system fluctuates. The central question of changepoint detection is deceptively simple: is this fluctuation just noise (random variation within the same regime), or is it signal (evidence that the system has transitioned to a qualitatively different state)?
This is harder than it sounds. There’s an irreducible tension between two types of error:
- False alarm (Type I): You detect a change where none exists. In manufacturing, this means shutting down a line for no reason. In medicine, it’s a false positive that triggers unnecessary intervention.
- Missed detection (Type II): You fail to detect a real change. In finance, this means holding a position through a regime shift. In epidemiology, it means missing the start of an outbreak.
Every detection algorithm negotiates this trade-off through its threshold: how much evidence do you demand before declaring that the world has changed? Set it too low and you cry wolf; set it too high and you react too late. There is no universal answer — the right threshold depends on the relative cost of each type of error in your specific context.
What Makes RegimeChange Different
Several R packages already address changepoint detection — the venerable changepoint package, wbs, not, ecp, and the Python ruptures library, to name a few. Each has strengths and limitations. RegimeChange distinguishes itself by integrating approaches that are usually siloed:
- Frequentist and Bayesian methods together. Most packages pick one paradigm. RegimeChange gives you both, plus deep learning, all callable through the same
detect_regimes()function. - Offline and online modes. You can analyze a complete dataset retrospectively (“when did the changes occur?”) or process streaming data in real time (“is a change happening right now?”) using the same library.
- Native uncertainty quantification. Every changepoint estimate comes with confidence intervals (via bootstrap for frequentist methods) or full posterior distributions (for Bayesian methods). This isn’t just a point estimate — it’s a statement about how confident you should be.
- Robustness to messy data. Real-world data has outliers, heavy tails, and autocorrelation. RegimeChange includes configurable robust estimation and AR(1) pre-whitening — features that, as the benchmarks show, make a dramatic difference on contaminated data.
- An optional Julia backend. For large datasets, the package can transparently dispatch to high-performance Julia implementations while keeping the R interface you’re used to.
Let’s look at each of these in more detail.
The Three Method Families
RegimeChange organizes its detection algorithms into three families. Here’s what each one does, explained simply.
Frequentist Methods
These are the classical, well-established workhorses of changepoint detection.
CUSUM (Cumulative Sum) is the granddaddy of them all, dating to E. S. Page’s 1954 paper. The idea is beautifully intuitive: you accumulate deviations from a target (or “baseline”) value over time. If the system is in the same regime, positive and negative deviations cancel out, and the cumulative sum hovers near zero. When a change occurs, deviations start piling up in one direction, and the statistic grows. When it crosses a threshold, you sound the alarm. CUSUM runs in O(n) time and is ideal for detecting a single change in mean, especially in online monitoring scenarios.
PELT (Pruned Exact Linear Time), introduced by Killick, Fearnhead, and Eckley in 2012, is the modern gold standard for detecting multiple changepoints. It uses dynamic programming to find the globally optimal segmentation of your data — the set of changepoints that minimizes a cost function (essentially: how well do the segments fit the data, penalized by the number of segments). The “pruning” part is clever: it discards candidate changepoint locations that provably can’t be optimal, which brings the average-case complexity down from O(n²) to O(n). RegimeChange supports several penalty criteria — BIC, AIC, MBIC, MDL — or you can specify a manual numeric penalty.
Binary Segmentation is a simpler, greedy alternative: find the single best changepoint, split the data there, then recurse on each segment. It’s fast (O(n log n)) but doesn’t guarantee the global optimum. Wild Binary Segmentation (WBS), from Fryzlewicz (2014), improves on this by searching over many random sub-intervals rather than the full dataset, making it much more robust when changepoints are close together.
The package also includes FPOP (Functional Pruning Optimal Partitioning), which maintains piecewise-quadratic cost functions for even better theoretical guarantees; E-Divisive, a nonparametric method using energy statistics that can detect changes in any aspect of the distribution; Kernel-based CPD, which maps data into a reproducing kernel Hilbert space to capture complex distributional shifts; and NOT (Narrowest-Over-Threshold), which excels at precise localization.
Bayesian Methods
Where frequentist methods give you a point estimate and (if you ask) a confidence interval, Bayesian methods give you something richer: a full probability distribution over when the change might have occurred.
BOCPD (Bayesian Online Changepoint Detection), from Adams and MacKay (2007), is the flagship. At each time step, it maintains a posterior distribution over the “run length” — the number of observations since the last changepoint. When a new data point arrives, it updates this distribution: either the run length grows by one (no change), or it resets to zero (change detected). The probability of a reset at each step is governed by a “hazard function,” typically a geometric distribution representing a constant probability of change per time step.
What makes BOCPD powerful is that it naturally produces, at every time point, the probability that a changepoint just occurred. This is a fundamentally different kind of output than a binary “change/no-change” flag — it’s a calibrated measure of confidence that updates with every new observation. RegimeChange implements BOCPD with several conjugate prior models: Normal-Gamma (for unknown mean and variance), Normal with known variance, Gamma-Poisson (for count data), and Normal-Wishart (for multivariate data).
Shiryaev-Roberts is another Bayesian sequential method, based on work by Shiryaev from the 1960s. It accumulates likelihood ratios from all possible past changepoint times and is asymptotically optimal for minimizing detection delay — that is, for detecting changes as quickly as possible once they occur.
Deep Learning Methods
For complex, nonlinear patterns that classical methods struggle with, RegimeChange offers optional deep learning detectors (requiring keras and tensorflow):
- Autoencoders are trained to reconstruct “normal” patterns. When the data changes regime, reconstruction error spikes — the model can’t rebuild what it hasn’t seen before.
- Temporal Convolutional Networks (TCNs) use dilated causal convolutions to model long-range temporal dependencies, operating either in supervised mode (if you have labeled changepoints for training) or unsupervised mode (predicting the next value and flagging prediction errors).
- Transformers bring self-attention to the problem, capturing long-range relationships in the time series.
- Contrastive Predictive Coding (CPC) learns representations through self-supervised contrastive learning, detecting changes where learned embeddings shift significantly.
An ensemble mode combines multiple deep learning methods, requiring a minimum number of them to agree before declaring a changepoint — a strategy that trades some sensitivity for robustness.
One Interface to Rule Them All
The genius of RegimeChange is that you don’t need to learn a different API for each method. Everything flows through a single function:
library(RegimeChange)# Generate data with a changepoint at t=200set.seed(42)data <- c(rnorm(200, 0, 1), rnorm(200, 2, 1))# Detect using PELT (the default offline method)result <- detect_regimes(data, method = "pelt")print(result)plot(result)# Switch to Bayesian online detectionresult <- detect_regimes(data, method = "bocpd", prior = normal_gamma(mu0 = 0, kappa0 = 1, alpha0 = 1, beta0 = 1))# Use CUSUM for online monitoringresult <- detect_regimes(data, method = "cusum", mode = "online", threshold = 5)
You specify what kind of change you’re looking for (type = "mean", "variance", "both", "trend", or "distribution"), how many changes you expect ("single", "multiple", or a specific integer), and the penalty criterion for model complexity. The function returns a regime_result object containing the detected changepoints, their confidence intervals, segment-by-segment statistics (mean, variance, etc.), and — for Bayesian methods — the full posterior distribution.
Online Detection
For real-time monitoring, you create a persistent detector object and feed it data one observation at a time:
detector <- regime_detector(method = "bocpd", prior = normal_gamma(), threshold = 0.5)for (x in data_stream) { detector <- update(detector, x) if (detector$last_result$alarm) { message("Changepoint detected at time ", detector$last_result$t) detector <- reset(detector) }}
This is the pattern you’d use for industrial monitoring, fraud detection, or epidemiological surveillance — anywhere data arrives sequentially and you need to react in real time.
Robustness: Where RegimeChange Really Shines
Real data is messy. It has outliers, heavy tails, and autocorrelation. Standard changepoint methods, which typically assume independent, normally distributed observations, can fail dramatically under these conditions.
RegimeChange addresses this with a configurable robustness layer in its PELT implementation. When you set robust = TRUE, the package:
- Winsorizes the data (clips extreme values beyond a specified quantile) to limit the influence of outliers.
- Uses Huber M-estimation or Tukey biweight loss functions instead of standard squared-error loss. These functions grow linearly (Huber) or flatten out entirely (Tukey) for large residuals, so a single outlier can’t dominate the cost calculation.
- Employs the Qn scale estimator (from Rousseeuw and Croux, 1993), which is far more efficient than the median absolute deviation (MAD) — 82% efficiency vs. 37% — while maintaining the same 50% breakdown point (meaning up to half the data can be contaminated before the estimator breaks).
You can also set robust = "auto", which analyzes the data’s contamination level — comparing the MAD to the standard deviation and counting outlier proportions — and automatically selects the appropriate robustness level ("none", "mild", "moderate", or "aggressive").
For autocorrelated data, the correct_ar = TRUE option estimates the AR(1) coefficient and applies a pre-whitening transformation (subtracting the lagged value scaled by the autocorrelation coefficient) before running detection. This prevents autocorrelation from inflating false positive rates.
The Benchmark Numbers
The package’s wiki reports extensive benchmarks against established R packages (changepoint, wbs, not, ecp) across 17 scenarios with 50 replications each. The results are striking:
- Overall: RegimeChange’s automatic selector achieved the highest mean F1 score (0.804) across all methods tested.
- Contaminated data: The robust mode achieved F1 = 0.998 across all contamination scenarios, versus 0.479 for the standard
changepointPELT — a 108% improvement on messy data. - Subtle variance changes (2:1 ratio): RegimeChange scored 0.667 vs. 0.400 for reference packages — a 67% improvement.
- Strong autocorrelation (AR coefficient 0.7): RegimeChange scored 0.900 vs. 0.720 — a 25% improvement.
In 16 of 17 scenarios, RegimeChange either won or tied. The single loss was marginal (0.017 in one heavy-tailed scenario).
Uncertainty Quantification: Beyond Point Estimates
A changepoint location of “observation 200” is only half the story. The other half is: how confident are we? RegimeChange takes this seriously, providing two types of uncertainty:
Location uncertainty — how precise is the estimate? For frequentist methods, the package uses block bootstrap resampling: it repeatedly resamples blocks of the data (preserving local dependence structure), reruns detection, and computes 95% confidence intervals from the distribution of bootstrap estimates. For Bayesian methods, uncertainty comes naturally from the posterior distribution over run lengths.
Existence uncertainty — is there even a change at all? BOCPD produces, at each time point, the posterior probability that a changepoint just occurred. This is arguably more informative than a p-value: it’s a direct probabilistic statement that you can use for decision-making.
The Julia Backend
R is wonderful for data analysis, but it can be slow for computationally intensive tasks. RegimeChange ships with an optional Julia backend — a complete reimplementation of the core algorithms (PELT, FPOP, BOCPD, CUSUM, Kernel CPD, WBS, and multivariate PELT) in Julia, a language designed for numerical performance.
The integration is seamless. You initialize Julia once:
init_julia()julia_available() # Returns TRUE if ready
After that, the package automatically dispatches to Julia for large datasets (n > 1000 by default) while using R for smaller ones. You can benchmark the difference with benchmark_backends(). The Julia implementation includes thoughtful numerical engineering: Welford’s algorithm for numerically stable running mean/variance, Kahan compensated summation for cumulative sums, log-domain computations to avoid underflow in BOCPD, and careful handling of ill-conditioned covariance matrices in the multivariate case.
If Julia isn’t installed, the package simply falls back to the R implementation — no errors, no fuss.
Evaluation and Comparison
RegimeChange includes a comprehensive evaluation framework. If you know the true changepoints (e.g., in simulated data or a well-studied real dataset), you can compute a full battery of metrics:
metrics <- evaluate(result, true_changepoints = c(100, 250), tolerance = 10)print(metrics)
This gives you precision, recall, and F1 score (with a tolerance window for matching), Hausdorff distance (the worst-case localization error), Rand Index and Adjusted Rand Index (segmentation agreement), and the covering metric (a weighted IoU measure). You can also run compare_methods() to test multiple algorithms on the same data and see the results side by side.
The package ships with several built-in datasets: well_log (geophysical measurements), industrial_sensor data, economic_cycles, and simulated_changepoints — giving you realistic examples to experiment with.
When to Use What
Here’s a practical guide based on the library’s design and benchmark results:
| Your Situation | Recommendation |
|---|---|
| Clean data, clear mean shifts | PELT with default settings |
| Outliers or heavy tails | PELT with robust = TRUE (or "auto") |
| Autocorrelated time series | PELT with correct_ar = TRUE |
| Need real-time detection | BOCPD or CUSUM via regime_detector() |
| Need probability of change (not just yes/no) | BOCPD |
| Nonparametric, no distributional assumptions | E-Divisive or Kernel CPD |
| Complex nonlinear patterns (and you have data to train on) | Deep learning methods |
| High-dimensional multivariate data | Sparse Projection CPD |
| Not sure which method to use | Ensemble mode, or let method = "auto" choose |
The Bigger Picture
What I find compelling about RegimeChange is its philosophical stance. The wiki frames changepoint detection not just as a technical problem but as an epistemological one: how do we distinguish real change from noise? The likelihood ratio — the ratio of how probable an observation is under the “changed” hypothesis vs. the “unchanged” hypothesis — is literally a quantification of surprise. Each observation casts a vote, evidence accumulates, and the decision threshold represents how much surprise we demand before acting.
This framing connects the technical machinery to the real-world stakes. Whether you’re monitoring a factory, tracking a disease outbreak, or studying climate data, the fundamental structure is the same: the world was in one state, it transitioned to another, and you need to detect that transition as reliably and quickly as possible.
RegimeChange gives you a unified toolkit for that task — one that spans the classical-to-modern spectrum, handles messy data gracefully, tells you not just where the change is but how confident you should be, and scales from your laptop to a Julia-accelerated backend when you need speed.
Getting Started
RegimeChange is available on GitHub and installs like any R package from source:
# install.packages("devtools")devtools::install_github("IsadoreNabi/RegimeChange")
It requires R ≥ 4.0.0 and depends on ggplot2, rlang, cli, and magrittr. The Julia backend is optional (requires Julia ≥ 1.6 and the JuliaCall R package), as are the deep learning methods (require keras and tensorflow). The package ships with three vignettes — an introduction, an offline detection guide, and a Bayesian methods tutorial — and the wiki is a comprehensive reference covering everything from mathematical foundations to detailed benchmarks.
For researchers and practitioners who need rigorous, flexible, and robust changepoint detection in R, RegimeChange is well worth a close look.


Leave a Comment/Deja un Comentario