MLflow vs Neptune.ai: Complete Guide to Experiment Tracking for MLOps

# MLflow vs Neptune.ai: Panduan Lengkap Experiment Tracking untuk MLOps Experiment tracking adalah komponen krusial dalam MLOps yang memungkinkan tim data science untuk melacak, membandingkan, dan me...

By Ruby Abdullah · · tutorial
MLOpsMLflowNeptune.aiExperiment TrackingMachine LearningPython

MLflow vs Neptune.ai: Complete Guide to Experiment Tracking for MLOps

Experiment tracking is a crucial component in MLOps that enables data science teams to track, compare, and reproduce machine learning experiments. In this tutorial, we'll compare two popular platforms: MLflow (open-source) and Neptune.ai (managed service), and learn how to use both.

Why is Experiment Tracking Important?

Without proper experiment tracking, ML teams often face:

  • Reproducibility crisis: Unable to reproduce previous experiment results
  • Lost experiments: Losing configurations that produced the best model
  • Collaboration issues: Difficult to share results across teams
  • Technical debt: Spreadsheets and manual notes that don't scale

Overview: MLflow vs Neptune.ai

| Aspect | MLflow | Neptune.ai |

|--------|--------|------------|

| Type | Open-source | Managed SaaS |

| Hosting | Self-hosted / Managed | Cloud-hosted |

| Pricing | Free (infra cost) | Free tier + paid plans |

| Setup | Manual setup | Instant |

| UI | Basic | Advanced |

| Collaboration | Limited | Built-in |

| Integrations | 15+ frameworks | 25+ frameworks |

| Model Registry | Yes | Yes |

| Best For | Full control, on-prem | Quick start, teams |

Part 1: MLflow

1.1 Installing MLflow

# Install MLflow

pip install mlflow

For tracking server with database backend

pip install mlflow[extras]

Start tracking server (local)

mlflow ui --port 5000

Or with backend store

mlflow server \

--backend-store-uri sqlite:///mlflow.db \

--default-artifact-root ./mlruns \

--host 0.0.0.0 \

--port 5000

1.2 Basic Experiment Tracking

import mlflow

import mlflow.sklearn

from sklearn.ensemble import RandomForestClassifier

from sklearn.datasets import loadiris

from sklearn.modelselection import traintestsplit

from sklearn.metrics import accuracyscore, f1score

Set tracking URI (optional, default: ./mlruns)

mlflow.settrackinguri("http://localhost:5000")

Set experiment name

mlflow.setexperiment("iris-classification")

Load data

X, y = loadiris(returnXy=True)

Xtrain, Xtest, ytrain, ytest = traintestsplit(X, y, testsize=0.2)

Start run

with mlflow.startrun(runname="random-forest-v1"):

# Log parameters

params = {

"nestimators": 100,

"maxdepth": 5,

"randomstate": 42

}

mlflow.logparams(params)

# Train model

model = RandomForestClassifier(*params)

model.fit(Xtrain, ytrain)

# Predict and evaluate

ypred = model.predict(Xtest)

accuracy = accuracyscore(ytest, ypred)

f1 = f1score(ytest, ypred, average='weighted')

# Log metrics

mlflow.logmetrics({

"accuracy": accuracy,

"f1score": f1

})

# Log model

mlflow.sklearn.logmodel(model, "model")

# Log artifacts (additional files)

with open("featureimportance.txt", "w") as f:

for name, importance in zip(loadiris().featurenames, model.featureimportances):

f.write(f"{name}: {importance:.4f}\n")

mlflow.logartifact("featureimportance.txt")

print(f"Run ID: {mlflow.activerun().info.runid}")

print(f"Accuracy: {accuracy:.4f}")

1.3 Hyperparameter Tuning with MLflow

import mlflow

from sklearn.ensemble import RandomForestClassifier

from sklearn.modelselection import crossvalscore

from sklearn.datasets import loadiris

import itertools

mlflow.setexperiment("iris-hyperparameter-tuning")

X, y = loadiris(returnXy=True)

Hyperparameter grid

paramgrid = {

"nestimators": [50, 100, 200],

"maxdepth": [3, 5, 10, None],

"minsamplessplit": [2, 5, 10]

}

Generate all combinations

keys = paramgrid.keys()

combinations = list(itertools.product(paramgrid.values()))

bestscore = 0

bestrunid = None

for combo in combinations:

params = dict(zip(keys, combo))

with mlflow.startrun():

# Log parameters

mlflow.logparams(params)

# Train and evaluate

model = RandomForestClassifier(params, randomstate=42)

scores = crossvalscore(model, X, y, cv=5, scoring='accuracy')

meanscore = scores.mean()

stdscore = scores.std()

# Log metrics

mlflow.logmetrics({

"cvaccuracymean": meanscore,

"cvaccuracystd": stdscore

})

# Track best

if meanscore > bestscore:

bestscore = meanscore

bestrunid = mlflow.activerun().info.runid

# Add tags

mlflow.settag("modeltype", "RandomForest")

print(f"Best score: {bestscore:.4f}")

print(f"Best run ID: {bestrunid}")

1.4 MLflow Model Registry

import mlflow

from mlflow.tracking import MlflowClient

client = MlflowClient()

Register model from run

runid = "your-run-id"

modeluri = f"runs:/{runid}/model"

Register model

result = mlflow.registermodel(modeluri, "IrisClassifier")

print(f"Model version: {result.version}")

Transition model stage

client.transitionmodelversionstage(

name="IrisClassifier",

version=1,

stage="Staging" # None, Staging, Production, Archived

)

Load model from registry

model = mlflow.pyfunc.loadmodel("models:/IrisClassifier/Staging")

Or specific version

model = mlflow.pyfunc.loadmodel("models:/IrisClassifier/1")

List all versions

for mv in client.searchmodelversions("name='IrisClassifier'"):

print(f"Version: {mv.version}, Stage: {mv.currentstage}")

1.5 MLflow Projects

Create file MLproject:

name: iris-training

condaenv: conda.yaml

entrypoints:

main:

parameters:

nestimators: {type: int, default: 100}

maxdepth: {type: int, default: 5}

command: "python train.py --nestimators {nestimators} --maxdepth {maxdepth}"

validate:

command: "python validate.py"

File conda.yaml:

name: iris-env

channels:

  • defaults
dependencies:

  • python=3.10
  • scikit-learn
  • pandas
  • pip
  • pip:
  • mlflow

Run project:

# Run locally

mlflow run . -P nestimators=200 -P maxdepth=10

Run from GitHub

mlflow run git@github.com:user/repo.git -P nestimators=200

1.6 MLflow with Deep Learning (PyTorch)

import mlflow

import mlflow.pytorch

import torch

import torch.nn as nn

import torch.optim as optim

from torch.utils.data import DataLoader, TensorDataset

from sklearn.datasets import loadiris

from sklearn.modelselection import traintestsplit

from sklearn.preprocessing import StandardScaler

import numpy as np

mlflow.setexperiment("pytorch-iris")

Prepare data

X, y = loadiris(returnXy=True)

scaler = StandardScaler()

X = scaler.fittransform(X)

Xtrain, Xtest, ytrain, ytest = traintestsplit(X, y, testsize=0.2)

Convert to tensors

Xtraint = torch.FloatTensor(Xtrain)

ytraint = torch.LongTensor(ytrain)

Xtestt = torch.FloatTensor(Xtest)

ytestt = torch.LongTensor(ytest)

traindataset = TensorDataset(Xtraint, ytraint)

trainloader = DataLoader(traindataset, batchsize=16, shuffle=True)

Define model

class IrisNet(nn.Module):

def init(self, hiddensize=64):

super().init()

self.fc1 = nn.Linear(4, hiddensize)

self.fc2 = nn.Linear(hiddensize, 32)

self.fc3 = nn.Linear(32, 3)

self.relu = nn.ReLU()

def forward(self, x):

x = self.relu(self.fc1(x))

x = self.relu(self.fc2(x))

return self.fc3(x)

Training with MLflow

with mlflow.startrun():

# Hyperparameters

params = {

"hiddensize": 64,

"learningrate": 0.01,

"epochs": 100,

"batchsize": 16

}

mlflow.logparams(params)

model = IrisNet(hiddensize=params["hiddensize"])

criterion = nn.CrossEntropyLoss()

optimizer = optim.Adam(model.parameters(), lr=params["learningrate"])

# Training loop

for epoch in range(params["epochs"]):

model.train()

totalloss = 0

for batchX, batchy in trainloader:

optimizer.zerograd()

outputs = model(batchX)

loss = criterion(outputs, batchy)

loss.backward()

optimizer.step()

totalloss += loss.item()

# Log metrics per epoch

avgloss = totalloss / len(trainloader)

mlflow.logmetric("trainloss", avgloss, step=epoch)

# Evaluation

if epoch % 10 == 0:

model.eval()

with torch.nograd():

outputs = model(Xtestt)

, predicted = torch.max(outputs, 1)

accuracy = (predicted == ytestt).sum().item() / len(ytestt)

mlflow.logmetric("testaccuracy", accuracy, step=epoch)

# Log final model

mlflow.pytorch.logmodel(model, "model")

# Log model summary

mlflow.settag("modelarchitecture", str(model))

Part 2: Neptune.ai

2.1 Setting Up Neptune.ai

# Install Neptune

pip install neptune

Set API token (from neptune.ai dashboard)

export NEPTUNEAPITOKEN="your-api-token"

2.2 Basic Experiment Tracking

import neptune

from sklearn.ensemble import RandomForestClassifier

from sklearn.datasets import loadiris

from sklearn.modelselection import traintestsplit

from sklearn.metrics import accuracyscore, f1score

Initialize Neptune run

run = neptune.initrun(

project="your-workspace/your-project",

apitoken="your-api-token", # Or from env variable

name="random-forest-v1",

tags=["classification", "sklearn", "iris"]

)

Load data

X, y = loadiris(returnXy=True)

Xtrain, Xtest, ytrain, ytest = traintestsplit(X, y, testsize=0.2)

Log parameters

params = {

"nestimators": 100,

"maxdepth": 5,

"randomstate": 42

}

run["parameters"] = params

Train model

model = RandomForestClassifier(params)

model.fit(Xtrain, ytrain)

Evaluate

ypred = model.predict(Xtest)

accuracy = accuracyscore(ytest, ypred)

f1 = f1score(ytest, ypred, average='weighted')

Log metrics

run["metrics/accuracy"] = accuracy

run["metrics/f1score"] = f1

Log feature importance

for name, importance in zip(loadiris().featurenames, model.featureimportances):

run[f"featureimportance/{name}"] = importance

Log model file

import joblib

joblib.dump(model, "model.pkl")

run["model"].upload("model.pkl")

Stop run

run.stop()

print(f"Accuracy: {accuracy:.4f}")

2.3 Tracking Training Progress (Series)

import neptune

import numpy as np

run = neptune.initrun(

project="your-workspace/your-project"

)

Simulate training loop

for epoch in range(100):

# Simulate metrics

trainloss = 1.0 / (epoch + 1) + np.random.random() 0.1

valloss = 1.2 / (epoch + 1) + np.random.random() 0.1

accuracy = 1 - valloss + np.random.random() * 0.05

# Log series data (auto-increments step)

run["train/loss"].append(trainloss)

run["val/loss"].append(valloss)

run["val/accuracy"].append(accuracy)

# Log with explicit step

run["train/epoch"].append(epoch)

run.stop()

2.4 Logging Artifacts and Visualizations

import neptune

import matplotlib.pyplot as plt

import pandas as pd

from sklearn.metrics import confusionmatrix, ConfusionMatrixDisplay

from sklearn.datasets import loadiris

from sklearn.ensemble import RandomForestClassifier

from sklearn.modelselection import traintestsplit

run = neptune.initrun(project="your-workspace/your-project")

Load and train

X, y = loadiris(returnXy=True)

Xtrain, Xtest, ytrain, ytest = traintestsplit(X, y, testsize=0.2)

model = RandomForestClassifier(nestimators=100)

model.fit(Xtrain, ytrain)

ypred = model.predict(Xtest)

1. Log confusion matrix as image

cm = confusionmatrix(ytest, ypred)

disp = ConfusionMatrixDisplay(cm, displaylabels=loadiris().targetnames)

disp.plot()

plt.savefig("confusionmatrix.png")

run["visualizations/confusionmatrix"].upload("confusionmatrix.png")

2. Log DataFrame

df = pd.DataFrame({

"actual": ytest,

"predicted": ypred,

"correct": ytest == ypred

})

run["predictions"].upload(neptune.types.File.ashtml(df))

3. Log interactive chart with Neptune

from neptune.types import File

Feature importance bar chart

fig, ax = plt.subplots()

ax.barh(loadiris().featurenames, model.featureimportances)

ax.setxlabel("Importance")

ax.settitle("Feature Importance")

run["visualizations/featureimportance"].upload(fig)

4. Log source code

run["sourcecode"].upload("train.py")

5. Log dataset info

run["dataset/info"] = {

"name": "Iris",

"nsamples": len(X),

"nfeatures": X.shape[1],

"nclasses": len(set(y))

}

run.stop()

2.5 Neptune with PyTorch

import neptune

import torch

import torch.nn as nn

import torch.optim as optim

from torch.utils.data import DataLoader, TensorDataset

from sklearn.datasets import loadiris

from sklearn.modelselection import traintestsplit

from sklearn.preprocessing import StandardScaler

Initialize Neptune

run = neptune.initrun(

project="your-workspace/your-project",

name="pytorch-iris",

tags=["pytorch", "neural-network"]

)

Data preparation

X, y = loadiris(returnXy=True)

scaler = StandardScaler()

X = scaler.fittransform(X)

Xtrain, Xtest, ytrain, ytest = traintestsplit(X, y, testsize=0.2)

Xtraint = torch.FloatTensor(Xtrain)

ytraint = torch.LongTensor(ytrain)

Xtestt = torch.FloatTensor(Xtest)

ytestt = torch.LongTensor(ytest)

trainloader = DataLoader(TensorDataset(Xtraint, ytraint), batchsize=16)

Model

class Net(nn.Module):

def init(self):

super().init()

self.fc1 = nn.Linear(4, 64)

self.fc2 = nn.Linear(64, 32)

self.fc3 = nn.Linear(32, 3)

def forward(self, x):

x = torch.relu(self.fc1(x))

x = torch.relu(self.fc2(x))

return self.fc3(x)

model = Net()

criterion = nn.CrossEntropyLoss()

optimizer = optim.Adam(model.parameters(), lr=0.01)

Log model architecture

run["model/architecture"] = str(model)

run["parameters"] = {

"learningrate": 0.01,

"epochs": 100,

"batchsize": 16,

"optimizer": "Adam"

}

Training

for epoch in range(100):

model.train()

epochloss = 0

for batchX, batchy in trainloader:

optimizer.zerograd()

outputs = model(batchX)

loss = criterion(outputs, batchy)

loss.backward()

optimizer.step()

epochloss += loss.item()

avgloss = epochloss / len(trainloader)

run["train/loss"].append(avgloss)

# Validation

model.eval()

with torch.nograd():

outputs = model(Xtestt)

valloss = criterion(outputs, ytestt).item()

, predicted = torch.max(outputs, 1)

accuracy = (predicted == ytestt).sum().item() / len(ytestt)

run["val/loss"].append(valloss)

run["val/accuracy"].append(accuracy)

Save and log model

torch.save(model.statedict(), "model.pt")

run["model/weights"].upload("model.pt")

run.stop()

2.6 Neptune Model Registry

import neptune

from neptune.types import File

Initialize model registry

modelversion = neptune.initmodelversion(

model="YOUR-PROJECT-KEY-MOD", # Model ID from Neptune

project="your-workspace/your-project"

)

Log model metadata

modelversion["model/framework"] = "scikit-learn"

modelversion["model/algorithm"] = "RandomForest"

Log model file

modelversion["model/binary"].upload("model.pkl")

Log performance metrics

modelversion["validation/accuracy"] = 0.95

modelversion["validation/f1score"] = 0.94

Log training run reference

modelversion["run/id"] = "YOUR-RUN-ID"

Change stage

modelversion.changestage("staging") # none, staging, production, archived

modelversion.stop()

Part 3: Practical Comparison

3.1 Side-by-Side Code Comparison

Logging Parameters:
# MLflow

mlflow.logparams({

"learningrate": 0.01,

"epochs": 100

})

Neptune

run["parameters"] = {

"learningrate": 0.01,

"epochs": 100

}

Logging Metrics:
# MLflow

mlflow.logmetric("accuracy", 0.95)

mlflow.logmetrics({"loss": 0.1, "f1": 0.94})

Neptune

run["metrics/accuracy"] = 0.95

run["metrics/loss"] = 0.1

Logging Series (Training Loop):
# MLflow

for epoch in range(100):

mlflow.logmetric("loss", lossvalue, step=epoch)

Neptune

for epoch in range(100):

run["train/loss"].append(lossvalue) # Auto step

Logging Artifacts:
# MLflow

mlflow.logartifact("model.pkl")

mlflow.logartifacts("./outputs")

Neptune

run["model"].upload("model.pkl")

run["outputs"].uploadfiles("./outputs")

3.2 Unified Wrapper

from abc import ABC, abstractmethod

class ExperimentTracker(ABC):

@abstractmethod

def logparams(self, params: dict): pass

@abstractmethod

def logmetric(self, name: str, value: float, step: int = None): pass

@abstractmethod

def logartifact(self, path: str): pass

@abstractmethod

def endrun(self): pass

class MLflowTracker(ExperimentTracker):

def init(self, experimentname: str):

import mlflow

mlflow.setexperiment(experimentname)

mlflow.startrun()

self.mlflow = mlflow

def logparams(self, params: dict):

self.mlflow.logparams(params)

def logmetric(self, name: str, value: float, step: int = None):

self.mlflow.logmetric(name, value, step=step)

def logartifact(self, path: str):

self.mlflow.logartifact(path)

def endrun(self):

self.mlflow.endrun()

class NeptuneTracker(ExperimentTracker):

def init(self, project: str, apitoken: str = None):

import neptune

self.run = neptune.initrun(project=project, apitoken=apitoken)

def logparams(self, params: dict):

self.run["parameters"] = params

def logmetric(self, name: str, value: float, step: int = None):

self.run[f"metrics/{name}"].append(value)

def logartifact(self, path: str):

self.run["artifacts"].upload(path)

def endrun(self):

self.run.stop()

Usage - easily switch between trackers

def trainmodel(tracker: ExperimentTracker):

tracker.logparams({"lr": 0.01, "epochs": 100})

for epoch in range(100):

loss = 1.0 / (epoch + 1)

tracker.logmetric("loss", loss, step=epoch)

tracker.logartifact("model.pkl")

tracker.endrun()

Use MLflow

tracker = MLflowTracker("my-experiment")

trainmodel(tracker)

Or use Neptune

tracker = NeptuneTracker("workspace/project")

train_model(tracker)

3.3 When to Use Each?

Choose MLflow if:
  • Need self-hosted solution (data sensitivity)
  • Limited budget (infrastructure cost only)
  • Already have infrastructure (Kubernetes, cloud)
  • Need Databricks integration
  • Small team with ML engineers who can maintain

Choose Neptune.ai if:
  • Need quick start without setup
  • Distributed team needing collaboration
  • Budget available for managed service
  • Need advanced visualization
  • Focus on experiment tracking (not full MLOps)

Conclusion

| Feature | MLflow | Neptune.ai |

|---------|--------|------------|

| Setup | Requires effort | Instant |

| Cost | Infra only | Subscription |

| UI | Functional | Modern |

| Collaboration | Basic | Excellent |

| Flexibility | High | Medium |

| Learning Curve | Medium | Low |

Recommendations:
  • Startup/Small team: Neptune.ai (free tier is enough to start)
  • Enterprise/On-prem: MLflow (full control)
  • Hybrid: Use both with unified wrapper

Both platforms are excellent for experiment tracking. The choice depends on your team's and organization's specific needs.

Related Articles

Complete Comet ML Tutorial: MLOps Platform for Experiment Tracking and Model Management

Tutorial Lengkap Comet ML: Platform MLOps untuk Experiment Tracking dan Model Management Dalam dunia machine learning mo...

Azure MLflow Integration Tutorial: Experiment Tracking on Azure

Tutorial Lengkap Azure MLflow Integration: Experiment Tracking dan Model Management Azure Machine Learning menyediakan i...

Complete Weights & Biases Tutorial: Experiment Tracking for Machine Learning

Tutorial Lengkap Weights & Biases: ML Experiment Tracking dan Visualization Weights & Biases (W&B) adalah platform MLOps...

Complete MLflow Tutorial: From Setup to Production

Pendahuluan MLflow adalah platform open-source untuk mengelola end-to-end machine learning lifecycle. Dikembangkan oleh ...