MLflow vs Neptune.ai: Complete Guide to Experiment Tracking for MLOps
Experiment tracking is a crucial component in MLOps that enables data science teams to track, compare, and reproduce machine learning experiments. In this tutorial, we'll compare two popular platforms: MLflow (open-source) and Neptune.ai (managed service), and learn how to use both.
Why is Experiment Tracking Important?
Without proper experiment tracking, ML teams often face:
- Reproducibility crisis: Unable to reproduce previous experiment results
- Lost experiments: Losing configurations that produced the best model
- Collaboration issues: Difficult to share results across teams
- Technical debt: Spreadsheets and manual notes that don't scale
Overview: MLflow vs Neptune.ai
| Aspect | MLflow | Neptune.ai |
|--------|--------|------------|
| Type | Open-source | Managed SaaS |
| Hosting | Self-hosted / Managed | Cloud-hosted |
| Pricing | Free (infra cost) | Free tier + paid plans |
| Setup | Manual setup | Instant |
| UI | Basic | Advanced |
| Collaboration | Limited | Built-in |
| Integrations | 15+ frameworks | 25+ frameworks |
| Model Registry | Yes | Yes |
| Best For | Full control, on-prem | Quick start, teams |
Part 1: MLflow
1.1 Installing MLflow
# Install MLflow
pip install mlflow
For tracking server with database backend
pip install mlflow[extras]
Start tracking server (local)
mlflow ui --port 5000
Or with backend store
mlflow server \
--backend-store-uri sqlite:///mlflow.db \
--default-artifact-root ./mlruns \
--host 0.0.0.0 \
--port 5000
1.2 Basic Experiment Tracking
import mlflow
import mlflow.sklearn
from sklearn.ensemble import RandomForestClassifier
from sklearn.datasets import loadiris
from sklearn.modelselection import traintestsplit
from sklearn.metrics import accuracyscore, f1score
Set tracking URI (optional, default: ./mlruns)
mlflow.settrackinguri("http://localhost:5000")
Set experiment name
mlflow.setexperiment("iris-classification")
Load data
X, y = loadiris(returnXy=True)
Xtrain, Xtest, ytrain, ytest = traintestsplit(X, y, testsize=0.2)
Start run
with mlflow.startrun(runname="random-forest-v1"):
# Log parameters
params = {
"nestimators": 100,
"maxdepth": 5,
"randomstate": 42
}
mlflow.logparams(params)
# Train model
model = RandomForestClassifier(*params)
model.fit(Xtrain, ytrain)
# Predict and evaluate
ypred = model.predict(Xtest)
accuracy = accuracyscore(ytest, ypred)
f1 = f1score(ytest, ypred, average='weighted')
# Log metrics
mlflow.logmetrics({
"accuracy": accuracy,
"f1score": f1
})
# Log model
mlflow.sklearn.logmodel(model, "model")
# Log artifacts (additional files)
with open("featureimportance.txt", "w") as f:
for name, importance in zip(loadiris().featurenames, model.featureimportances):
f.write(f"{name}: {importance:.4f}\n")
mlflow.logartifact("featureimportance.txt")
print(f"Run ID: {mlflow.activerun().info.runid}")
print(f"Accuracy: {accuracy:.4f}")
1.3 Hyperparameter Tuning with MLflow
import mlflow
from sklearn.ensemble import RandomForestClassifier
from sklearn.modelselection import crossvalscore
from sklearn.datasets import loadiris
import itertools
mlflow.setexperiment("iris-hyperparameter-tuning")
X, y = loadiris(returnXy=True)
Hyperparameter grid
paramgrid = {
"nestimators": [50, 100, 200],
"maxdepth": [3, 5, 10, None],
"minsamplessplit": [2, 5, 10]
}
Generate all combinations
keys = paramgrid.keys()
combinations = list(itertools.product(paramgrid.values()))
bestscore = 0
bestrunid = None
for combo in combinations:
params = dict(zip(keys, combo))
with mlflow.startrun():
# Log parameters
mlflow.logparams(params)
# Train and evaluate
model = RandomForestClassifier(params, randomstate=42)
scores = crossvalscore(model, X, y, cv=5, scoring='accuracy')
meanscore = scores.mean()
stdscore = scores.std()
# Log metrics
mlflow.logmetrics({
"cvaccuracymean": meanscore,
"cvaccuracystd": stdscore
})
# Track best
if meanscore > bestscore:
bestscore = meanscore
bestrunid = mlflow.activerun().info.runid
# Add tags
mlflow.settag("modeltype", "RandomForest")
print(f"Best score: {bestscore:.4f}")
print(f"Best run ID: {bestrunid}")
1.4 MLflow Model Registry
import mlflow
from mlflow.tracking import MlflowClient
client = MlflowClient()
Register model from run
runid = "your-run-id"
modeluri = f"runs:/{runid}/model"
Register model
result = mlflow.registermodel(modeluri, "IrisClassifier")
print(f"Model version: {result.version}")
Transition model stage
client.transitionmodelversionstage(
name="IrisClassifier",
version=1,
stage="Staging" # None, Staging, Production, Archived
)
Load model from registry
model = mlflow.pyfunc.loadmodel("models:/IrisClassifier/Staging")
Or specific version
model = mlflow.pyfunc.loadmodel("models:/IrisClassifier/1")
List all versions
for mv in client.searchmodelversions("name='IrisClassifier'"):
print(f"Version: {mv.version}, Stage: {mv.currentstage}")
1.5 MLflow Projects
Create file MLproject:
name: iris-training
condaenv: conda.yaml
entrypoints:
main:
parameters:
nestimators: {type: int, default: 100}
maxdepth: {type: int, default: 5}
command: "python train.py --nestimators {nestimators} --maxdepth {maxdepth}"
validate:
command: "python validate.py"
File conda.yaml:
name: iris-env
channels:
- defaults
dependencies:
- python=3.10
- scikit-learn
- pandas
- pip
- pip:
- mlflow
Run project:
# Run locally
mlflow run . -P nestimators=200 -P maxdepth=10
Run from GitHub
mlflow run git@github.com:user/repo.git -P nestimators=200
1.6 MLflow with Deep Learning (PyTorch)
import mlflow
import mlflow.pytorch
import torch
import torch.nn as nn
import torch.optim as optim
from torch.utils.data import DataLoader, TensorDataset
from sklearn.datasets import loadiris
from sklearn.modelselection import traintestsplit
from sklearn.preprocessing import StandardScaler
import numpy as np
mlflow.setexperiment("pytorch-iris")
Prepare data
X, y = loadiris(returnXy=True)
scaler = StandardScaler()
X = scaler.fittransform(X)
Xtrain, Xtest, ytrain, ytest = traintestsplit(X, y, testsize=0.2)
Convert to tensors
Xtraint = torch.FloatTensor(Xtrain)
ytraint = torch.LongTensor(ytrain)
Xtestt = torch.FloatTensor(Xtest)
ytestt = torch.LongTensor(ytest)
traindataset = TensorDataset(Xtraint, ytraint)
trainloader = DataLoader(traindataset, batchsize=16, shuffle=True)
Define model
class IrisNet(nn.Module):
def init(self, hiddensize=64):
super().init()
self.fc1 = nn.Linear(4, hiddensize)
self.fc2 = nn.Linear(hiddensize, 32)
self.fc3 = nn.Linear(32, 3)
self.relu = nn.ReLU()
def forward(self, x):
x = self.relu(self.fc1(x))
x = self.relu(self.fc2(x))
return self.fc3(x)
Training with MLflow
with mlflow.startrun():
# Hyperparameters
params = {
"hiddensize": 64,
"learningrate": 0.01,
"epochs": 100,
"batchsize": 16
}
mlflow.logparams(params)
model = IrisNet(hiddensize=params["hiddensize"])
criterion = nn.CrossEntropyLoss()
optimizer = optim.Adam(model.parameters(), lr=params["learningrate"])
# Training loop
for epoch in range(params["epochs"]):
model.train()
totalloss = 0
for batchX, batchy in trainloader:
optimizer.zerograd()
outputs = model(batchX)
loss = criterion(outputs, batchy)
loss.backward()
optimizer.step()
totalloss += loss.item()
# Log metrics per epoch
avgloss = totalloss / len(trainloader)
mlflow.logmetric("trainloss", avgloss, step=epoch)
# Evaluation
if epoch % 10 == 0:
model.eval()
with torch.nograd():
outputs = model(Xtestt)
, predicted = torch.max(outputs, 1)
accuracy = (predicted == ytestt).sum().item() / len(ytestt)
mlflow.logmetric("testaccuracy", accuracy, step=epoch)
# Log final model
mlflow.pytorch.logmodel(model, "model")
# Log model summary
mlflow.settag("modelarchitecture", str(model))
Part 2: Neptune.ai
2.1 Setting Up Neptune.ai
# Install Neptune
pip install neptune
Set API token (from neptune.ai dashboard)
export NEPTUNEAPITOKEN="your-api-token"
2.2 Basic Experiment Tracking
import neptune
from sklearn.ensemble import RandomForestClassifier
from sklearn.datasets import loadiris
from sklearn.modelselection import traintestsplit
from sklearn.metrics import accuracyscore, f1score
Initialize Neptune run
run = neptune.initrun(
project="your-workspace/your-project",
apitoken="your-api-token", # Or from env variable
name="random-forest-v1",
tags=["classification", "sklearn", "iris"]
)
Load data
X, y = loadiris(returnXy=True)
Xtrain, Xtest, ytrain, ytest = traintestsplit(X, y, testsize=0.2)
Log parameters
params = {
"nestimators": 100,
"maxdepth": 5,
"randomstate": 42
}
run["parameters"] = params
Train model
model = RandomForestClassifier(params)
model.fit(Xtrain, ytrain)
Evaluate
ypred = model.predict(Xtest)
accuracy = accuracyscore(ytest, ypred)
f1 = f1score(ytest, ypred, average='weighted')
Log metrics
run["metrics/accuracy"] = accuracy
run["metrics/f1score"] = f1
Log feature importance
for name, importance in zip(loadiris().featurenames, model.featureimportances):
run[f"featureimportance/{name}"] = importance
Log model file
import joblib
joblib.dump(model, "model.pkl")
run["model"].upload("model.pkl")
Stop run
run.stop()
print(f"Accuracy: {accuracy:.4f}")
2.3 Tracking Training Progress (Series)
import neptune
import numpy as np
run = neptune.initrun(
project="your-workspace/your-project"
)
Simulate training loop
for epoch in range(100):
# Simulate metrics
trainloss = 1.0 / (epoch + 1) + np.random.random() 0.1
valloss = 1.2 / (epoch + 1) + np.random.random() 0.1
accuracy = 1 - valloss + np.random.random() * 0.05
# Log series data (auto-increments step)
run["train/loss"].append(trainloss)
run["val/loss"].append(valloss)
run["val/accuracy"].append(accuracy)
# Log with explicit step
run["train/epoch"].append(epoch)
run.stop()
2.4 Logging Artifacts and Visualizations
import neptune
import matplotlib.pyplot as plt
import pandas as pd
from sklearn.metrics import confusionmatrix, ConfusionMatrixDisplay
from sklearn.datasets import loadiris
from sklearn.ensemble import RandomForestClassifier
from sklearn.modelselection import traintestsplit
run = neptune.initrun(project="your-workspace/your-project")
Load and train
X, y = loadiris(returnXy=True)
Xtrain, Xtest, ytrain, ytest = traintestsplit(X, y, testsize=0.2)
model = RandomForestClassifier(nestimators=100)
model.fit(Xtrain, ytrain)
ypred = model.predict(Xtest)
1. Log confusion matrix as image
cm = confusionmatrix(ytest, ypred)
disp = ConfusionMatrixDisplay(cm, displaylabels=loadiris().targetnames)
disp.plot()
plt.savefig("confusionmatrix.png")
run["visualizations/confusionmatrix"].upload("confusionmatrix.png")
2. Log DataFrame
df = pd.DataFrame({
"actual": ytest,
"predicted": ypred,
"correct": ytest == ypred
})
run["predictions"].upload(neptune.types.File.ashtml(df))
3. Log interactive chart with Neptune
from neptune.types import File
Feature importance bar chart
fig, ax = plt.subplots()
ax.barh(loadiris().featurenames, model.featureimportances)
ax.setxlabel("Importance")
ax.settitle("Feature Importance")
run["visualizations/featureimportance"].upload(fig)
4. Log source code
run["sourcecode"].upload("train.py")
5. Log dataset info
run["dataset/info"] = {
"name": "Iris",
"nsamples": len(X),
"nfeatures": X.shape[1],
"nclasses": len(set(y))
}
run.stop()
2.5 Neptune with PyTorch
import neptune
import torch
import torch.nn as nn
import torch.optim as optim
from torch.utils.data import DataLoader, TensorDataset
from sklearn.datasets import loadiris
from sklearn.modelselection import traintestsplit
from sklearn.preprocessing import StandardScaler
Initialize Neptune
run = neptune.initrun(
project="your-workspace/your-project",
name="pytorch-iris",
tags=["pytorch", "neural-network"]
)
Data preparation
X, y = loadiris(returnXy=True)
scaler = StandardScaler()
X = scaler.fittransform(X)
Xtrain, Xtest, ytrain, ytest = traintestsplit(X, y, testsize=0.2)
Xtraint = torch.FloatTensor(Xtrain)
ytraint = torch.LongTensor(ytrain)
Xtestt = torch.FloatTensor(Xtest)
ytestt = torch.LongTensor(ytest)
trainloader = DataLoader(TensorDataset(Xtraint, ytraint), batchsize=16)
Model
class Net(nn.Module):
def init(self):
super().init()
self.fc1 = nn.Linear(4, 64)
self.fc2 = nn.Linear(64, 32)
self.fc3 = nn.Linear(32, 3)
def forward(self, x):
x = torch.relu(self.fc1(x))
x = torch.relu(self.fc2(x))
return self.fc3(x)
model = Net()
criterion = nn.CrossEntropyLoss()
optimizer = optim.Adam(model.parameters(), lr=0.01)
Log model architecture
run["model/architecture"] = str(model)
run["parameters"] = {
"learningrate": 0.01,
"epochs": 100,
"batchsize": 16,
"optimizer": "Adam"
}
Training
for epoch in range(100):
model.train()
epochloss = 0
for batchX, batchy in trainloader:
optimizer.zerograd()
outputs = model(batchX)
loss = criterion(outputs, batchy)
loss.backward()
optimizer.step()
epochloss += loss.item()
avgloss = epochloss / len(trainloader)
run["train/loss"].append(avgloss)
# Validation
model.eval()
with torch.nograd():
outputs = model(Xtestt)
valloss = criterion(outputs, ytestt).item()
, predicted = torch.max(outputs, 1)
accuracy = (predicted == ytestt).sum().item() / len(ytestt)
run["val/loss"].append(valloss)
run["val/accuracy"].append(accuracy)
Save and log model
torch.save(model.statedict(), "model.pt")
run["model/weights"].upload("model.pt")
run.stop()
2.6 Neptune Model Registry
import neptune
from neptune.types import File
Initialize model registry
modelversion = neptune.initmodelversion(
model="YOUR-PROJECT-KEY-MOD", # Model ID from Neptune
project="your-workspace/your-project"
)
Log model metadata
modelversion["model/framework"] = "scikit-learn"
modelversion["model/algorithm"] = "RandomForest"
Log model file
modelversion["model/binary"].upload("model.pkl")
Log performance metrics
modelversion["validation/accuracy"] = 0.95
modelversion["validation/f1score"] = 0.94
Log training run reference
modelversion["run/id"] = "YOUR-RUN-ID"
Change stage
modelversion.changestage("staging") # none, staging, production, archived
modelversion.stop()
Part 3: Practical Comparison
3.1 Side-by-Side Code Comparison
Logging Parameters:# MLflow
mlflow.logparams({
"learningrate": 0.01,
"epochs": 100
})
Neptune
run["parameters"] = {
"learningrate": 0.01,
"epochs": 100
}
Logging Metrics:
# MLflow
mlflow.logmetric("accuracy", 0.95)
mlflow.logmetrics({"loss": 0.1, "f1": 0.94})
Neptune
run["metrics/accuracy"] = 0.95
run["metrics/loss"] = 0.1
Logging Series (Training Loop):
# MLflow
for epoch in range(100):
mlflow.logmetric("loss", lossvalue, step=epoch)
Neptune
for epoch in range(100):
run["train/loss"].append(lossvalue) # Auto step
Logging Artifacts:
# MLflow
mlflow.logartifact("model.pkl")
mlflow.logartifacts("./outputs")
Neptune
run["model"].upload("model.pkl")
run["outputs"].uploadfiles("./outputs")
3.2 Unified Wrapper
from abc import ABC, abstractmethod
class ExperimentTracker(ABC):
@abstractmethod
def logparams(self, params: dict): pass
@abstractmethod
def logmetric(self, name: str, value: float, step: int = None): pass
@abstractmethod
def logartifact(self, path: str): pass
@abstractmethod
def endrun(self): pass
class MLflowTracker(ExperimentTracker):
def init(self, experimentname: str):
import mlflow
mlflow.setexperiment(experimentname)
mlflow.startrun()
self.mlflow = mlflow
def logparams(self, params: dict):
self.mlflow.logparams(params)
def logmetric(self, name: str, value: float, step: int = None):
self.mlflow.logmetric(name, value, step=step)
def logartifact(self, path: str):
self.mlflow.logartifact(path)
def endrun(self):
self.mlflow.endrun()
class NeptuneTracker(ExperimentTracker):
def init(self, project: str, apitoken: str = None):
import neptune
self.run = neptune.initrun(project=project, apitoken=apitoken)
def logparams(self, params: dict):
self.run["parameters"] = params
def logmetric(self, name: str, value: float, step: int = None):
self.run[f"metrics/{name}"].append(value)
def logartifact(self, path: str):
self.run["artifacts"].upload(path)
def endrun(self):
self.run.stop()
Usage - easily switch between trackers
def trainmodel(tracker: ExperimentTracker):
tracker.logparams({"lr": 0.01, "epochs": 100})
for epoch in range(100):
loss = 1.0 / (epoch + 1)
tracker.logmetric("loss", loss, step=epoch)
tracker.logartifact("model.pkl")
tracker.endrun()
Use MLflow
tracker = MLflowTracker("my-experiment")
trainmodel(tracker)
Or use Neptune
tracker = NeptuneTracker("workspace/project")
train_model(tracker)
3.3 When to Use Each?
Choose MLflow if:- Need self-hosted solution (data sensitivity)
- Limited budget (infrastructure cost only)
- Already have infrastructure (Kubernetes, cloud)
- Need Databricks integration
- Small team with ML engineers who can maintain
- Need quick start without setup
- Distributed team needing collaboration
- Budget available for managed service
- Need advanced visualization
- Focus on experiment tracking (not full MLOps)
Conclusion
| Feature | MLflow | Neptune.ai |
|---------|--------|------------|
| Setup | Requires effort | Instant |
| Cost | Infra only | Subscription |
| UI | Functional | Modern |
| Collaboration | Basic | Excellent |
| Flexibility | High | Medium |
| Learning Curve | Medium | Low |
Recommendations:- Startup/Small team: Neptune.ai (free tier is enough to start)
- Enterprise/On-prem: MLflow (full control)
- Hybrid: Use both with unified wrapper
Both platforms are excellent for experiment tracking. The choice depends on your team's and organization's specific needs.