Anomaly Detection in Python with PyOD: A Practical Guide
Most real-world datasets contain a small fraction of records that do not behave like the rest: a fraudulent payment, a faulty sensor reading, an intrusion attempt, or a defective unit on a production line. PyOD is a mature, scikit-learn-style library that bundles more than 40 outlier detection algorithms behind one consistent interface, which makes it straightforward to try several methods and compare them. This tutorial walks through the core concepts, the unified API, the main algorithm families, and a realistic end-to-end workflow.
What Is Outlier / Anomaly Detection?
Outlier detection (often used interchangeably with anomaly detection) is the task of identifying observations that deviate so much from the majority that they are likely generated by a different process. Unlike standard classification, you usually have very few labeled anomalies, or none at all, so most of the work is unsupervised: the model learns what "normal" looks like and scores each point by how far it sits from that normal region.
A few important framing points:
- Anomalies are rare by definition. If 40% of your data is "anomalous," it is probably not an anomaly problem but an imbalanced classification problem.
- The boundary between normal and abnormal is rarely crisp. Detectors produce a continuous score; you choose a cutoff.
- "Anomalous" is context-dependent. A transaction of USD 5,000 is normal for one customer and suspicious for another.
Real-World Use Cases
- Fraud detection. Credit card transactions, insurance claims, account takeovers. Fraud patterns shift over time, so unsupervised or semi-supervised detectors complement rule engines.
- Fault and defect detection. Vibration, temperature, and pressure sensors on industrial equipment; defective items in manufacturing.
- Network intrusion detection. Unusual traffic volume, port-scanning behaviour, or connection patterns that differ from a learned baseline.
- Quality control. Out-of-spec measurements in a production batch, or readings that drift outside expected tolerances.
- Data cleaning. Catching corrupted rows, unit errors, and sensor glitches before they pollute a downstream model.
Installation
PyOD itself is lightweight and depends mainly on NumPy, SciPy, and scikit-learn.
pip install pyod
Some detectors have optional dependencies. The neural models (for example AutoEncoder) require PyTorch, and a few utilities use additional packages:
# Required only for the neural-network based detectors
pip install torch
Optional, used by some combination/visualization helpers
pip install combo matplotlib
You can confirm the installation from a Python shell:
import pyod
print(pyod.version)
The Unified, scikit-learn-Style API
Every detector in PyOD follows the same lifecycle, which is the main reason the library is pleasant to use. Once you learn one model, you know all of them.
The common pieces are:
fit(X)— learn the model from training data. Detection is typically unsupervised, so you do not pass labels.decisionfunction(X)— return a raw outlier score for new data. Higher means more anomalous.predict(X)— return binary labels (0 = inlier, 1 = outlier).predictproba(X)— return calibrated-ish probabilities of being an outlier.decisionscores— the raw outlier scores computed on the training data duringfit.labels— the binary labels assigned to the training data.threshold— the score cutoff used to separate inliers from outliers.contamination— the constructor parameter that tells the model the expected proportion of outliers (default0.1). It is used to setthreshold.
The contamination value is central. PyOD ranks all training points by score and labels the top contamination fraction as outliers. It does not change how scores are computed, only where the cutoff falls. If you set it too high you get false positives; too low and you miss real anomalies.
A First Example: KNN on a 2D Dataset
PyOD ships a data generator that creates a mixture of inliers and injected outliers, which is handy for learning and for sanity checks.
from pyod.models.knn import KNN
from pyod.utils.data import generatedata
0.1 of the points are outliers
Xtrain, Xtest, ytrain, ytest = generatedata(
ntrain=500,
ntest=200,
nfeatures=2,
contamination=0.1,
randomstate=42,
)
clf = KNN(contamination=0.1)
clf.fit(Xtrain)
Results on the training data are stored on the fitted object
trainscores = clf.decisionscores # raw scores
trainlabels = clf.labels # 0 / 1
print("Threshold:", clf.threshold)
Apply to unseen data
testscores = clf.decisionfunction(Xtest)
testlabels = clf.predict(Xtest)
testproba = clf.predictproba(Xtest)
print("Predicted outliers in test set:", testlabels.sum())
The KNN detector scores each point by the distance to its k-th nearest neighbour: points far from their neighbours sit in low-density regions and score high. You can switch the aggregation with the method argument (largest, mean, or median).
Key Algorithm Families
There is no single best detector. Different methods assume different things about what makes a point anomalous, so it pays to understand the families and try a few.
Proximity-Based Methods
These rely on distances or local density. They work well when anomalies are spatially separated from dense clusters, but they are sensitive to feature scaling and can struggle in very high dimensions.
from pyod.models.knn import KNN
from pyod.models.lof import LOF
from pyod.models.cblof import CBLOF
knn = KNN(nneighbors=20, contamination=0.05)
lof = LOF(nneighbors=20, contamination=0.05) # Local Outlier Factor
cblof = CBLOF(nclusters=8, contamination=0.05) # Cluster-Based LOF
for name, model in [("KNN", knn), ("LOF", lof), ("CBLOF", cblof)]:
model.fit(Xtrain)
print(name, "flagged", model.labels.sum(), "training outliers")
LOF compares the local density of a point with that of its neighbours, so it catches anomalies that are only abnormal relative to a local cluster. CBLOF first clusters the data and scores points by their distance to large clusters, which scales better than pure LOF on larger datasets.
Linear and Statistical Methods
These model the structure of the normal data with linear projections or probability distributions.
from pyod.models.pca import PCA
from pyod.models.mcd import MCD
from pyod.models.ecod import ECOD
from pyod.models.copod import COPOD
pca = PCA(contamination=0.05) # reconstruction error in principal subspace
mcd = MCD(contamination=0.05) # robust Mahalanobis distance
ecod = ECOD(contamination=0.05) # empirical cumulative distribution
copod = COPOD(contamination=0.05) # copula-based
for name, model in [("PCA", pca), ("MCD", mcd), ("ECOD", ecod), ("COPOD", copod)]:
model.fit(Xtrain)
print(name, "threshold:", round(model.threshold, 3))
ECOD and COPOD deserve special mention. Both are parameter-free and fast: they estimate the tail probability of each feature and aggregate, with no neighbours, clusters, or distance computations to tune. They handle high-dimensional data gracefully and make excellent default baselines. When you do not know where to start, run ECOD first.
Ensemble and Tree-Based Methods
These combine many weak detectors or use tree structures to isolate points.
from pyod.models.iforest import IForest
from pyod.models.featurebagging import FeatureBagging
from pyod.models.lof import LOF
iforest = IForest(nestimators=200, contamination=0.05, randomstate=42)
fb = FeatureBagging(baseestimator=LOF(), nestimators=20, contamination=0.05)
iforest.fit(Xtrain)
fb.fit(Xtrain)
print("IForest outliers:", iforest.labels.sum())
print("FeatureBagging outliers:", fb.labels.sum())
IForest (Isolation Forest) isolates points with random splits; anomalies require fewer splits to isolate, so they get higher scores. It is fast, handles many features, and is a strong general-purpose choice. FeatureBagging trains base detectors on random feature subsets and aggregates them, which improves robustness in high dimensions.
Neural Methods
For larger, more complex datasets, an autoencoder learns to reconstruct normal data; points that reconstruct poorly are flagged. This requires PyTorch.
from pyod.models.autoencoder import AutoEncoder
ae = AutoEncoder(
hiddenneuronlist=[32, 16, 16, 32],
epochnum=30,
batchsize=64,
contamination=0.05,
)
ae.fit(Xtrain)
print("AutoEncoder outliers:", ae.labels.sum())
Neural detectors need more data and more tuning than the statistical methods, so reach for them when simpler approaches plateau, not as a first step.
Combining Multiple Detectors
Because no single method dominates, combining several can give more robust results. PyOD provides combination utilities, but first you must normalize scores so that different detectors are comparable.
import numpy as np
from pyod.models.knn import KNN
from pyod.models.iforest import IForest
from pyod.models.copod import COPOD
from pyod.utils.utility import standardizer
from pyod.models.combination import average, maximization, aom
detectors = [KNN(), IForest(randomstate=42), COPOD()]
Collect a score column per detector
trainscores = np.zeros([Xtrain.shape[0], len(detectors)])
testscores = np.zeros([Xtest.shape[0], len(detectors)])
for i, clf in enumerate(detectors):
clf.fit(Xtrain)
trainscores[:, i] = clf.decisionscores
testscores[:, i] = clf.decisionfunction(Xtest)
Normalize so every detector contributes on the same scale
trainscoresnorm, testscoresnorm = standardizer(trainscores, testscores)
Three common combination strategies
combavg = average(testscoresnorm)
combmax = maximization(testscoresnorm)
combaom = aom(testscoresnorm, nbuckets=3)
print("Average combination shape:", combavg.shape)
The strategies differ in temperament:
averageis stable and reduces variance; a good default.maximizationtakes the highest score across detectors, so it is sensitive — useful when any single detector flagging a point should count.aom(Average of Maximum) andmoa(Maximum of Average) split detectors into buckets and blend the two behaviours, trading off stability and sensitivity.
Skipping standardizer is a common mistake: an Isolation Forest score and a KNN distance live on completely different scales, and averaging them raw lets one detector dominate.
Thresholding: From Scores to Labels
The raw score is the honest output of a detector; the binary label is a decision you impose on top of it. PyOD derives threshold from contamination during fit, but you are free to choose your own cutoff on the scores.
import numpy as np
from pyod.models.iforest import IForest
clf = IForest(contamination=0.05, randomstate=42)
clf.fit(Xtrain)
scores = clf.decisionfunction(Xtest)
Option 1: use the model's learned threshold
labelsdefault = clf.predict(Xtest)
Option 2: pick a custom percentile cutoff
cutoff = np.percentile(clf.decisionscores, 97) # top 3% as outliers
labelscustom = (scores > cutoff).astype(int)
print("Default flagged:", labelsdefault.sum())
print("Custom flagged:", labelscustom.sum())
Choosing the cutoff is a business decision as much as a statistical one. If each flagged case triggers an expensive manual review, you want a conservative (high) threshold. If a missed anomaly is catastrophic, you accept more false positives.
Evaluation
When You Have Labels
If you have even a small labeled holdout, use it. PyOD's evaluateprint reports ROC-AUC and precision@n in one line.
from pyod.models.knn import KNN
from pyod.utils.data import generatedata
from pyod.utils.data import evaluateprint
Xtrain, Xtest, ytrain, ytest = generatedata(
ntrain=500, ntest=200, contamination=0.1, randomstate=42
)
clf = KNN()
clf.fit(Xtrain)
Scores on the test set
ytestscores = clf.decisionfunction(Xtest)
evaluateprint("KNN", ytest, ytestscores)
- ROC-AUC measures ranking quality across all thresholds: can the detector put anomalies above normal points? It is threshold-independent.
- Precision@n asks: of the n points the detector ranks most anomalous (where n is the number of true anomalies), how many are correct? This matches the practical reality of a limited review budget.
When You Don't Have Labels
This is the common case, and it is genuinely hard. Without ground truth you cannot compute ROC-AUC honestly. Practical tactics:
- Have a domain expert manually review the top-ranked points and judge precision@n informally.
- Check stability: do several different algorithms agree on the same flagged points? High agreement is reassuring.
- Inject synthetic anomalies with known properties and confirm the detector catches them.
- Track the flag rate over time; sudden spikes usually mean drift or a data pipeline problem, not a surge of real anomalies.
Be honest about uncertainty. An unsupervised detector that "found 200 anomalies" has found 200 unusual points; whether they are the anomalies you care about is a separate question that data alone cannot answer.
Model Persistence
Fitted detectors are ordinary Python objects, so joblib works well for saving and loading.
import joblib
from pyod.models.iforest import IForest
clf = IForest(contamination=0.05, randomstate=42)
clf.fit(Xtrain)
joblib.dump(clf, "iforestdetector.joblib")
Later, in another process
loaded = joblib.load("iforestdetector.joblib")
newlabels = loaded.predict(Xtest)
Persist the fitted scaler alongside the model. Scoring new data with a different scaling than training is a frequent and silent source of bad predictions.
End-to-End Example: Flagging Abnormal Transactions
The following pipeline reflects how you would approach a tabular fraud / abnormal-transaction problem in practice: scale the features, choose contamination deliberately, fit a robust detector, threshold, and inspect the top cases.
import numpy as np
import pandas as pd
from sklearn.preprocessing import StandardScaler
from pyod.models.ecod import ECOD
from pyod.models.iforest import IForest
Synthetic stand-in for a transactions table.
In production this comes from your warehouse.
rng = np.random.defaultrng(42)
n = 5000
normal = pd.DataFrame({
"amount": rng.lognormal(mean=3.0, sigma=0.5, size=n),
"ntx24h": rng.poisson(5, size=n),
"merchantrisk": rng.normal(0.2, 0.1, size=n),
"distancekm": rng.exponential(10, size=n),
})
A handful of injected abnormal transactions
abnormal = pd.DataFrame({
"amount": rng.lognormal(mean=6.0, sigma=0.6, size=50),
"ntx24h": rng.poisson(40, size=50),
"merchantrisk": rng.normal(0.85, 0.1, size=50),
"distancekm": rng.exponential(400, size=50),
})
df = pd.concat([normal, abnormal], ignoreindex=True)
features = ["amount", "ntx24h", "merchantrisk", "distancekm"]
1. Scale: distance- and distribution-based detectors need comparable features
scaler = StandardScaler()
X = scaler.fittransform(df[features])
2. Choose contamination from a prior estimate of the fraud rate.
Roughly 50 / 5050 here; in real data, start from historical fraud rates.
contamination = 0.01
3. Fit a parameter-free baseline and a tree ensemble, then compare
ecod = ECOD(contamination=contamination)
iforest = IForest(contamination=contamination, nestimators=200, randomstate=42)
ecod.fit(X)
iforest.fit(X)
df["ecodscore"] = ecod.decisionscores
df["iforestscore"] = iforest.decisionscores
df["ecodflag"] = ecod.labels
df["iforestflag"] = iforest.labels
4. Agreement between detectors is a useful confidence signal
df["bothflag"] = ((df["ecodflag"] == 1) & (df["iforestflag"] == 1)).astype(int)
print("ECOD flagged:", int(df["ecodflag"].sum()))
print("IForest flagged:", int(df["iforestflag"].sum()))
print("Both agree:", int(df["bothflag"].sum()))
5. Inspect the highest-risk transactions for manual review
top = df.sortvalues("iforestscore", ascending=False).head(10)
print(top[features + ["iforestscore", "bothflag"]])
In a real system you would replace the synthetic frame with your warehouse query, calibrate contamination against historical fraud rates, route the agreed-upon flags to a review queue, and feed any confirmed labels back to improve thresholds over time.
Scope Notes: GPU and Time Series
- GPU acceleration. For very large datasets, the companion library
pytodreimplements several PyOD detectors on top of PyTorch tensors so they can run on a GPU. The API mirrors PyOD, so migration is mostly an import change. Reach for it only when CPU runtime is an actual bottleneck. - Time series. PyOD treats each row as an independent observation; it does not model temporal order on its own. For time-series anomaly detection you typically engineer windowed or lagged features first, or use a library built for sequences (the related
todsproject targets this). Do not feed raw timestamps and expect temporal context.
Best Practices
- Always scale your features. Distance- and density-based detectors (
KNN,LOF,CBLOF,PCA,MCD) are dominated by large-magnitude features otherwise. Tree-based methods are less sensitive, but scaling rarely hurts. - Choose
contaminationdeliberately. Base it on a prior estimate of the anomaly rate, not the default0.1. When unsure, lean conservative and review the score distribution. - Start with parameter-free baselines.
ECODandCOPODgive a strong, fast, tuning-free reference before you invest in heavier models. - Ensemble for robustness. Combine a few diverse detectors with
standardizerplusaverage; agreement across methods is a meaningful confidence signal. - Be skeptical of evaluation without labels. Unsupervised scores rank unusualness, not your specific definition of "bad." Validate with domain experts and cross-detector agreement.
- Persist the scaler with the model. Mismatched preprocessing at scoring time silently degrades predictions.
- Separate the score from the decision. Keep raw scores around; set the threshold according to the cost of false positives versus missed anomalies.
Conclusion and Key Takeaways
PyOD turns anomaly detection from a scattered collection of papers into a single, consistent toolkit. The unified fit / decision_function / predict API means you can swap algorithms with a one-line change and compare them fairly. Start simple with parameter-free detectors like ECOD, add an Isolation Forest for a tree-based view, and combine them when you need robustness.
Key takeaways:
- Anomaly detection is mostly unsupervised: the model learns "normal" and scores deviation.
contaminationsets the threshold, not the scores; choose it from domain knowledge.- Scale features, especially for proximity-based methods.
ECODandCOPODare fast, parameter-free, high-dimensional-friendly defaults.- Combine diverse detectors with score normalization for robustness.
- Evaluate with ROC-AUC and precision@n when labels exist; rely on expert review and cross-method agreement when they do not.
- Persist the fitted scaler together with the detector.
With these building blocks you can move from a quick KNN sanity check to a production-grade detection pipeline using the same library and the same mental model throughout.