MLflow vs Neptune.ai: Panduan Lengkap Experiment Tracking untuk MLOps
Experiment tracking adalah komponen krusial dalam MLOps yang memungkinkan tim data science untuk melacak, membandingkan, dan mereproduksi eksperimen machine learning. Dalam tutorial ini, kita akan membandingkan dua platform populer: MLflow (open-source) dan Neptune.ai (managed service), serta mempelajari cara menggunakan keduanya.
Mengapa Experiment Tracking Penting?
Tanpa experiment tracking yang proper, tim ML sering menghadapi:
- Reproducibility crisis: Tidak bisa mereproduksi hasil eksperimen sebelumnya
- Lost experiments: Kehilangan konfigurasi yang menghasilkan model terbaik
- Collaboration issues: Sulit berbagi hasil antar tim
- Technical debt: Spreadsheet dan catatan manual yang tidak scalable
Overview: MLflow vs Neptune.ai
| Aspek | MLflow | Neptune.ai |
|-------|--------|------------|
| Type | Open-source | Managed SaaS |
| Hosting | Self-hosted / Managed | Cloud-hosted |
| Pricing | Free (infra cost) | Free tier + paid plans |
| Setup | Manual setup | Instant |
| UI | Basic | Advanced |
| Collaboration | Limited | Built-in |
| Integrations | 15+ frameworks | 25+ frameworks |
| Model Registry | Yes | Yes |
| Best For | Full control, on-prem | Quick start, teams |
Bagian 1: MLflow
1.1 Instalasi MLflow
# Install MLflow
pip install mlflow
Untuk tracking server dengan database backend
pip install mlflow[extras]
Start tracking server (local)
mlflow ui --port 5000
Atau dengan backend store
mlflow server \
--backend-store-uri sqlite:///mlflow.db \
--default-artifact-root ./mlruns \
--host 0.0.0.0 \
--port 5000
1.2 Basic Experiment Tracking
import mlflow
import mlflow.sklearn
from sklearn.ensemble import RandomForestClassifier
from sklearn.datasets import loadiris
from sklearn.modelselection import traintestsplit
from sklearn.metrics import accuracyscore, f1score
Set tracking URI (optional, default: ./mlruns)
mlflow.settrackinguri("http://localhost:5000")
Set experiment name
mlflow.setexperiment("iris-classification")
Load data
X, y = loadiris(returnXy=True)
Xtrain, Xtest, ytrain, ytest = traintestsplit(X, y, testsize=0.2)
Start run
with mlflow.startrun(runname="random-forest-v1"):
# Log parameters
params = {
"nestimators": 100,
"maxdepth": 5,
"randomstate": 42
}
mlflow.logparams(params)
# Train model
model = RandomForestClassifier(*params)
model.fit(Xtrain, ytrain)
# Predict and evaluate
ypred = model.predict(Xtest)
accuracy = accuracyscore(ytest, ypred)
f1 = f1score(ytest, ypred, average='weighted')
# Log metrics
mlflow.logmetrics({
"accuracy": accuracy,
"f1score": f1
})
# Log model
mlflow.sklearn.logmodel(model, "model")
# Log artifacts (additional files)
with open("featureimportance.txt", "w") as f:
for name, importance in zip(loadiris().featurenames, model.featureimportances):
f.write(f"{name}: {importance:.4f}\n")
mlflow.logartifact("featureimportance.txt")
print(f"Run ID: {mlflow.activerun().info.runid}")
print(f"Accuracy: {accuracy:.4f}")
1.3 Hyperparameter Tuning dengan MLflow
import mlflow
from sklearn.ensemble import RandomForestClassifier
from sklearn.modelselection import crossvalscore
from sklearn.datasets import loadiris
import itertools
mlflow.setexperiment("iris-hyperparameter-tuning")
X, y = loadiris(returnXy=True)
Hyperparameter grid
paramgrid = {
"nestimators": [50, 100, 200],
"maxdepth": [3, 5, 10, None],
"minsamplessplit": [2, 5, 10]
}
Generate all combinations
keys = paramgrid.keys()
combinations = list(itertools.product(paramgrid.values()))
bestscore = 0
bestrunid = None
for combo in combinations:
params = dict(zip(keys, combo))
with mlflow.startrun():
# Log parameters
mlflow.logparams(params)
# Train and evaluate
model = RandomForestClassifier(params, randomstate=42)
scores = crossvalscore(model, X, y, cv=5, scoring='accuracy')
meanscore = scores.mean()
stdscore = scores.std()
# Log metrics
mlflow.logmetrics({
"cvaccuracymean": meanscore,
"cvaccuracystd": stdscore
})
# Track best
if meanscore > bestscore:
bestscore = meanscore
bestrunid = mlflow.activerun().info.runid
# Add tags
mlflow.settag("modeltype", "RandomForest")
print(f"Best score: {bestscore:.4f}")
print(f"Best run ID: {bestrunid}")
1.4 MLflow Model Registry
import mlflow
from mlflow.tracking import MlflowClient
client = MlflowClient()
Register model dari run
runid = "your-run-id"
modeluri = f"runs:/{runid}/model"
Register model
result = mlflow.registermodel(modeluri, "IrisClassifier")
print(f"Model version: {result.version}")
Transition model stage
client.transitionmodelversionstage(
name="IrisClassifier",
version=1,
stage="Staging" # None, Staging, Production, Archived
)
Load model dari registry
model = mlflow.pyfunc.loadmodel("models:/IrisClassifier/Staging")
Atau specific version
model = mlflow.pyfunc.loadmodel("models:/IrisClassifier/1")
List all versions
for mv in client.searchmodelversions("name='IrisClassifier'"):
print(f"Version: {mv.version}, Stage: {mv.currentstage}")
1.5 MLflow Projects
Buat file MLproject:
name: iris-training
condaenv: conda.yaml
entrypoints:
main:
parameters:
nestimators: {type: int, default: 100}
maxdepth: {type: int, default: 5}
command: "python train.py --nestimators {nestimators} --maxdepth {maxdepth}"
validate:
command: "python validate.py"
File conda.yaml:
name: iris-env
channels:
- defaults
dependencies:
- python=3.10
- scikit-learn
- pandas
- pip
- pip:
- mlflow
Run project:
# Run locally
mlflow run . -P nestimators=200 -P maxdepth=10
Run from GitHub
mlflow run git@github.com:user/repo.git -P nestimators=200
1.6 MLflow dengan Deep Learning (PyTorch)
import mlflow
import mlflow.pytorch
import torch
import torch.nn as nn
import torch.optim as optim
from torch.utils.data import DataLoader, TensorDataset
from sklearn.datasets import loadiris
from sklearn.modelselection import traintestsplit
from sklearn.preprocessing import StandardScaler
import numpy as np
mlflow.setexperiment("pytorch-iris")
Prepare data
X, y = loadiris(returnXy=True)
scaler = StandardScaler()
X = scaler.fittransform(X)
Xtrain, Xtest, ytrain, ytest = traintestsplit(X, y, testsize=0.2)
Convert to tensors
Xtraint = torch.FloatTensor(Xtrain)
ytraint = torch.LongTensor(ytrain)
Xtestt = torch.FloatTensor(Xtest)
ytestt = torch.LongTensor(ytest)
traindataset = TensorDataset(Xtraint, ytraint)
trainloader = DataLoader(traindataset, batchsize=16, shuffle=True)
Define model
class IrisNet(nn.Module):
def init(self, hiddensize=64):
super().init()
self.fc1 = nn.Linear(4, hiddensize)
self.fc2 = nn.Linear(hiddensize, 32)
self.fc3 = nn.Linear(32, 3)
self.relu = nn.ReLU()
def forward(self, x):
x = self.relu(self.fc1(x))
x = self.relu(self.fc2(x))
return self.fc3(x)
Training dengan MLflow
with mlflow.startrun():
# Hyperparameters
params = {
"hiddensize": 64,
"learningrate": 0.01,
"epochs": 100,
"batchsize": 16
}
mlflow.logparams(params)
model = IrisNet(hiddensize=params["hiddensize"])
criterion = nn.CrossEntropyLoss()
optimizer = optim.Adam(model.parameters(), lr=params["learningrate"])
# Training loop
for epoch in range(params["epochs"]):
model.train()
totalloss = 0
for batchX, batchy in trainloader:
optimizer.zerograd()
outputs = model(batchX)
loss = criterion(outputs, batchy)
loss.backward()
optimizer.step()
totalloss += loss.item()
# Log metrics per epoch
avgloss = totalloss / len(trainloader)
mlflow.logmetric("trainloss", avgloss, step=epoch)
# Evaluation
if epoch % 10 == 0:
model.eval()
with torch.nograd():
outputs = model(Xtestt)
, predicted = torch.max(outputs, 1)
accuracy = (predicted == ytestt).sum().item() / len(ytestt)
mlflow.logmetric("testaccuracy", accuracy, step=epoch)
# Log final model
mlflow.pytorch.logmodel(model, "model")
# Log model summary
mlflow.settag("modelarchitecture", str(model))
Bagian 2: Neptune.ai
2.1 Setup Neptune.ai
# Install Neptune
pip install neptune
Set API token (dari neptune.ai dashboard)
export NEPTUNEAPITOKEN="your-api-token"
2.2 Basic Experiment Tracking
import neptune
from sklearn.ensemble import RandomForestClassifier
from sklearn.datasets import loadiris
from sklearn.modelselection import traintestsplit
from sklearn.metrics import accuracyscore, f1score
Initialize Neptune run
run = neptune.initrun(
project="your-workspace/your-project",
apitoken="your-api-token", # Atau dari env variable
name="random-forest-v1",
tags=["classification", "sklearn", "iris"]
)
Load data
X, y = loadiris(returnXy=True)
Xtrain, Xtest, ytrain, ytest = traintestsplit(X, y, testsize=0.2)
Log parameters
params = {
"nestimators": 100,
"maxdepth": 5,
"randomstate": 42
}
run["parameters"] = params
Train model
model = RandomForestClassifier(params)
model.fit(Xtrain, ytrain)
Evaluate
ypred = model.predict(Xtest)
accuracy = accuracyscore(ytest, ypred)
f1 = f1score(ytest, ypred, average='weighted')
Log metrics
run["metrics/accuracy"] = accuracy
run["metrics/f1score"] = f1
Log feature importance
for name, importance in zip(loadiris().featurenames, model.featureimportances):
run[f"featureimportance/{name}"] = importance
Log model file
import joblib
joblib.dump(model, "model.pkl")
run["model"].upload("model.pkl")
Stop run
run.stop()
print(f"Accuracy: {accuracy:.4f}")
2.3 Tracking Training Progress (Series)
import neptune
import numpy as np
run = neptune.initrun(
project="your-workspace/your-project"
)
Simulate training loop
for epoch in range(100):
# Simulate metrics
trainloss = 1.0 / (epoch + 1) + np.random.random() 0.1
valloss = 1.2 / (epoch + 1) + np.random.random() 0.1
accuracy = 1 - valloss + np.random.random() * 0.05
# Log series data (auto-increments step)
run["train/loss"].append(trainloss)
run["val/loss"].append(valloss)
run["val/accuracy"].append(accuracy)
# Log dengan explicit step
run["train/epoch"].append(epoch)
run.stop()
2.4 Logging Artifacts dan Visualizations
import neptune
import matplotlib.pyplot as plt
import pandas as pd
from sklearn.metrics import confusionmatrix, ConfusionMatrixDisplay
from sklearn.datasets import loadiris
from sklearn.ensemble import RandomForestClassifier
from sklearn.modelselection import traintestsplit
run = neptune.initrun(project="your-workspace/your-project")
Load and train
X, y = loadiris(returnXy=True)
Xtrain, Xtest, ytrain, ytest = traintestsplit(X, y, testsize=0.2)
model = RandomForestClassifier(nestimators=100)
model.fit(Xtrain, ytrain)
ypred = model.predict(Xtest)
1. Log confusion matrix as image
cm = confusionmatrix(ytest, ypred)
disp = ConfusionMatrixDisplay(cm, displaylabels=loadiris().targetnames)
disp.plot()
plt.savefig("confusionmatrix.png")
run["visualizations/confusionmatrix"].upload("confusionmatrix.png")
2. Log DataFrame
df = pd.DataFrame({
"actual": ytest,
"predicted": ypred,
"correct": ytest == ypred
})
run["predictions"].upload(neptune.types.File.ashtml(df))
3. Log interactive chart dengan Neptune
from neptune.types import File
Feature importance bar chart
fig, ax = plt.subplots()
ax.barh(loadiris().featurenames, model.featureimportances)
ax.setxlabel("Importance")
ax.settitle("Feature Importance")
run["visualizations/featureimportance"].upload(fig)
4. Log source code
run["sourcecode"].upload("train.py")
5. Log dataset info
run["dataset/info"] = {
"name": "Iris",
"nsamples": len(X),
"nfeatures": X.shape[1],
"nclasses": len(set(y))
}
run.stop()
2.5 Neptune dengan PyTorch
import neptune
import torch
import torch.nn as nn
import torch.optim as optim
from torch.utils.data import DataLoader, TensorDataset
from sklearn.datasets import loadiris
from sklearn.modelselection import traintestsplit
from sklearn.preprocessing import StandardScaler
Initialize Neptune
run = neptune.initrun(
project="your-workspace/your-project",
name="pytorch-iris",
tags=["pytorch", "neural-network"]
)
Data preparation
X, y = loadiris(returnXy=True)
scaler = StandardScaler()
X = scaler.fittransform(X)
Xtrain, Xtest, ytrain, ytest = traintestsplit(X, y, testsize=0.2)
Xtraint = torch.FloatTensor(Xtrain)
ytraint = torch.LongTensor(ytrain)
Xtestt = torch.FloatTensor(Xtest)
ytestt = torch.LongTensor(ytest)
trainloader = DataLoader(TensorDataset(Xtraint, ytraint), batchsize=16)
Model
class Net(nn.Module):
def init(self):
super().init()
self.fc1 = nn.Linear(4, 64)
self.fc2 = nn.Linear(64, 32)
self.fc3 = nn.Linear(32, 3)
def forward(self, x):
x = torch.relu(self.fc1(x))
x = torch.relu(self.fc2(x))
return self.fc3(x)
model = Net()
criterion = nn.CrossEntropyLoss()
optimizer = optim.Adam(model.parameters(), lr=0.01)
Log model architecture
run["model/architecture"] = str(model)
run["parameters"] = {
"learningrate": 0.01,
"epochs": 100,
"batchsize": 16,
"optimizer": "Adam"
}
Training
for epoch in range(100):
model.train()
epochloss = 0
for batchX, batchy in trainloader:
optimizer.zerograd()
outputs = model(batchX)
loss = criterion(outputs, batchy)
loss.backward()
optimizer.step()
epochloss += loss.item()
avgloss = epochloss / len(trainloader)
run["train/loss"].append(avgloss)
# Validation
model.eval()
with torch.nograd():
outputs = model(Xtestt)
valloss = criterion(outputs, ytestt).item()
, predicted = torch.max(outputs, 1)
accuracy = (predicted == ytestt).sum().item() / len(ytestt)
run["val/loss"].append(valloss)
run["val/accuracy"].append(accuracy)
Save and log model
torch.save(model.statedict(), "model.pt")
run["model/weights"].upload("model.pt")
run.stop()
2.6 Neptune Model Registry
import neptune
from neptune.types import File
Initialize model registry
modelversion = neptune.initmodelversion(
model="YOUR-PROJECT-KEY-MOD", # Model ID dari Neptune
project="your-workspace/your-project"
)
Log model metadata
modelversion["model/framework"] = "scikit-learn"
modelversion["model/algorithm"] = "RandomForest"
Log model file
modelversion["model/binary"].upload("model.pkl")
Log performance metrics
modelversion["validation/accuracy"] = 0.95
modelversion["validation/f1score"] = 0.94
Log training run reference
modelversion["run/id"] = "YOUR-RUN-ID"
Change stage
modelversion.changestage("staging") # none, staging, production, archived
modelversion.stop()
Bagian 3: Perbandingan Praktis
3.1 Side-by-Side Code Comparison
Logging Parameters:# MLflow
mlflow.logparams({
"learningrate": 0.01,
"epochs": 100
})
Neptune
run["parameters"] = {
"learningrate": 0.01,
"epochs": 100
}
Logging Metrics:
# MLflow
mlflow.logmetric("accuracy", 0.95)
mlflow.logmetrics({"loss": 0.1, "f1": 0.94})
Neptune
run["metrics/accuracy"] = 0.95
run["metrics/loss"] = 0.1
Logging Series (Training Loop):
# MLflow
for epoch in range(100):
mlflow.logmetric("loss", lossvalue, step=epoch)
Neptune
for epoch in range(100):
run["train/loss"].append(lossvalue) # Auto step
Logging Artifacts:
# MLflow
mlflow.logartifact("model.pkl")
mlflow.logartifacts("./outputs")
Neptune
run["model"].upload("model.pkl")
run["outputs"].uploadfiles("./outputs")
3.2 Unified Wrapper
from abc import ABC, abstractmethod
class ExperimentTracker(ABC):
@abstractmethod
def logparams(self, params: dict): pass
@abstractmethod
def logmetric(self, name: str, value: float, step: int = None): pass
@abstractmethod
def logartifact(self, path: str): pass
@abstractmethod
def endrun(self): pass
class MLflowTracker(ExperimentTracker):
def init(self, experimentname: str):
import mlflow
mlflow.setexperiment(experimentname)
mlflow.startrun()
self.mlflow = mlflow
def logparams(self, params: dict):
self.mlflow.logparams(params)
def logmetric(self, name: str, value: float, step: int = None):
self.mlflow.logmetric(name, value, step=step)
def logartifact(self, path: str):
self.mlflow.logartifact(path)
def endrun(self):
self.mlflow.endrun()
class NeptuneTracker(ExperimentTracker):
def init(self, project: str, apitoken: str = None):
import neptune
self.run = neptune.initrun(project=project, apitoken=apitoken)
def logparams(self, params: dict):
self.run["parameters"] = params
def logmetric(self, name: str, value: float, step: int = None):
self.run[f"metrics/{name}"].append(value)
def logartifact(self, path: str):
self.run["artifacts"].upload(path)
def endrun(self):
self.run.stop()
Usage - easily switch between trackers
def trainmodel(tracker: ExperimentTracker):
tracker.logparams({"lr": 0.01, "epochs": 100})
for epoch in range(100):
loss = 1.0 / (epoch + 1)
tracker.logmetric("loss", loss, step=epoch)
tracker.logartifact("model.pkl")
tracker.endrun()
Use MLflow
tracker = MLflowTracker("my-experiment")
trainmodel(tracker)
Or use Neptune
tracker = NeptuneTracker("workspace/project")
train_model(tracker)
3.3 Kapan Menggunakan Masing-masing?
Pilih MLflow jika:- Butuh self-hosted solution (data sensitivity)
- Budget terbatas (infrastructure cost only)
- Sudah punya infrastructure (Kubernetes, cloud)
- Butuh integrasi dengan Databricks
- Tim kecil dengan ML engineer yang bisa maintain
- Butuh quick start tanpa setup
- Tim distributed yang butuh collaboration
- Budget tersedia untuk managed service
- Butuh advanced visualization
- Fokus pada experiment tracking (bukan full MLOps)
Kesimpulan
| Feature | MLflow | Neptune.ai |
|---------|--------|------------|
| Setup | Butuh effort | Instant |
| Cost | Infra only | Subscription |
| UI | Functional | Modern |
| Collaboration | Basic | Excellent |
| Flexibility | High | Medium |
| Learning Curve | Medium | Low |
Rekomendasi:- Startup/Small team: Neptune.ai (free tier cukup untuk mulai)
- Enterprise/On-prem: MLflow (full control)
- Hybrid: Gunakan keduanya dengan unified wrapper
Kedua platform sama-sama excellent untuk experiment tracking. Pilihan tergantung pada kebutuhan spesifik tim dan organisasi Anda.