MLflow vs Neptune.ai: Panduan Lengkap Experiment Tracking untuk MLOps

# MLflow vs Neptune.ai: Panduan Lengkap Experiment Tracking untuk MLOps Experiment tracking adalah komponen krusial dalam MLOps yang memungkinkan tim data science untuk melacak, membandingkan, dan me...

By Ruby Abdullah · · tutorial
MLOpsMLflowNeptune.aiExperiment TrackingMachine LearningPython

MLflow vs Neptune.ai: Panduan Lengkap Experiment Tracking untuk MLOps

Experiment tracking adalah komponen krusial dalam MLOps yang memungkinkan tim data science untuk melacak, membandingkan, dan mereproduksi eksperimen machine learning. Dalam tutorial ini, kita akan membandingkan dua platform populer: MLflow (open-source) dan Neptune.ai (managed service), serta mempelajari cara menggunakan keduanya.

Mengapa Experiment Tracking Penting?

Tanpa experiment tracking yang proper, tim ML sering menghadapi:

  • Reproducibility crisis: Tidak bisa mereproduksi hasil eksperimen sebelumnya
  • Lost experiments: Kehilangan konfigurasi yang menghasilkan model terbaik
  • Collaboration issues: Sulit berbagi hasil antar tim
  • Technical debt: Spreadsheet dan catatan manual yang tidak scalable

Overview: MLflow vs Neptune.ai

| Aspek | MLflow | Neptune.ai |

|-------|--------|------------|

| Type | Open-source | Managed SaaS |

| Hosting | Self-hosted / Managed | Cloud-hosted |

| Pricing | Free (infra cost) | Free tier + paid plans |

| Setup | Manual setup | Instant |

| UI | Basic | Advanced |

| Collaboration | Limited | Built-in |

| Integrations | 15+ frameworks | 25+ frameworks |

| Model Registry | Yes | Yes |

| Best For | Full control, on-prem | Quick start, teams |

Bagian 1: MLflow

1.1 Instalasi MLflow

# Install MLflow

pip install mlflow

Untuk tracking server dengan database backend

pip install mlflow[extras]

Start tracking server (local)

mlflow ui --port 5000

Atau dengan backend store

mlflow server \

--backend-store-uri sqlite:///mlflow.db \

--default-artifact-root ./mlruns \

--host 0.0.0.0 \

--port 5000

1.2 Basic Experiment Tracking

import mlflow

import mlflow.sklearn

from sklearn.ensemble import RandomForestClassifier

from sklearn.datasets import loadiris

from sklearn.modelselection import traintestsplit

from sklearn.metrics import accuracyscore, f1score

Set tracking URI (optional, default: ./mlruns)

mlflow.settrackinguri("http://localhost:5000")

Set experiment name

mlflow.setexperiment("iris-classification")

Load data

X, y = loadiris(returnXy=True)

Xtrain, Xtest, ytrain, ytest = traintestsplit(X, y, testsize=0.2)

Start run

with mlflow.startrun(runname="random-forest-v1"):

# Log parameters

params = {

"nestimators": 100,

"maxdepth": 5,

"randomstate": 42

}

mlflow.logparams(params)

# Train model

model = RandomForestClassifier(*params)

model.fit(Xtrain, ytrain)

# Predict and evaluate

ypred = model.predict(Xtest)

accuracy = accuracyscore(ytest, ypred)

f1 = f1score(ytest, ypred, average='weighted')

# Log metrics

mlflow.logmetrics({

"accuracy": accuracy,

"f1score": f1

})

# Log model

mlflow.sklearn.logmodel(model, "model")

# Log artifacts (additional files)

with open("featureimportance.txt", "w") as f:

for name, importance in zip(loadiris().featurenames, model.featureimportances):

f.write(f"{name}: {importance:.4f}\n")

mlflow.logartifact("featureimportance.txt")

print(f"Run ID: {mlflow.activerun().info.runid}")

print(f"Accuracy: {accuracy:.4f}")

1.3 Hyperparameter Tuning dengan MLflow

import mlflow

from sklearn.ensemble import RandomForestClassifier

from sklearn.modelselection import crossvalscore

from sklearn.datasets import loadiris

import itertools

mlflow.setexperiment("iris-hyperparameter-tuning")

X, y = loadiris(returnXy=True)

Hyperparameter grid

paramgrid = {

"nestimators": [50, 100, 200],

"maxdepth": [3, 5, 10, None],

"minsamplessplit": [2, 5, 10]

}

Generate all combinations

keys = paramgrid.keys()

combinations = list(itertools.product(paramgrid.values()))

bestscore = 0

bestrunid = None

for combo in combinations:

params = dict(zip(keys, combo))

with mlflow.startrun():

# Log parameters

mlflow.logparams(params)

# Train and evaluate

model = RandomForestClassifier(params, randomstate=42)

scores = crossvalscore(model, X, y, cv=5, scoring='accuracy')

meanscore = scores.mean()

stdscore = scores.std()

# Log metrics

mlflow.logmetrics({

"cvaccuracymean": meanscore,

"cvaccuracystd": stdscore

})

# Track best

if meanscore > bestscore:

bestscore = meanscore

bestrunid = mlflow.activerun().info.runid

# Add tags

mlflow.settag("modeltype", "RandomForest")

print(f"Best score: {bestscore:.4f}")

print(f"Best run ID: {bestrunid}")

1.4 MLflow Model Registry

import mlflow

from mlflow.tracking import MlflowClient

client = MlflowClient()

Register model dari run

runid = "your-run-id"

modeluri = f"runs:/{runid}/model"

Register model

result = mlflow.registermodel(modeluri, "IrisClassifier")

print(f"Model version: {result.version}")

Transition model stage

client.transitionmodelversionstage(

name="IrisClassifier",

version=1,

stage="Staging" # None, Staging, Production, Archived

)

Load model dari registry

model = mlflow.pyfunc.loadmodel("models:/IrisClassifier/Staging")

Atau specific version

model = mlflow.pyfunc.loadmodel("models:/IrisClassifier/1")

List all versions

for mv in client.searchmodelversions("name='IrisClassifier'"):

print(f"Version: {mv.version}, Stage: {mv.currentstage}")

1.5 MLflow Projects

Buat file MLproject:

name: iris-training

condaenv: conda.yaml

entrypoints:

main:

parameters:

nestimators: {type: int, default: 100}

maxdepth: {type: int, default: 5}

command: "python train.py --nestimators {nestimators} --maxdepth {maxdepth}"

validate:

command: "python validate.py"

File conda.yaml:

name: iris-env

channels:

  • defaults
dependencies:

  • python=3.10
  • scikit-learn
  • pandas
  • pip
  • pip:
  • mlflow

Run project:

# Run locally

mlflow run . -P nestimators=200 -P maxdepth=10

Run from GitHub

mlflow run git@github.com:user/repo.git -P nestimators=200

1.6 MLflow dengan Deep Learning (PyTorch)

import mlflow

import mlflow.pytorch

import torch

import torch.nn as nn

import torch.optim as optim

from torch.utils.data import DataLoader, TensorDataset

from sklearn.datasets import loadiris

from sklearn.modelselection import traintestsplit

from sklearn.preprocessing import StandardScaler

import numpy as np

mlflow.setexperiment("pytorch-iris")

Prepare data

X, y = loadiris(returnXy=True)

scaler = StandardScaler()

X = scaler.fittransform(X)

Xtrain, Xtest, ytrain, ytest = traintestsplit(X, y, testsize=0.2)

Convert to tensors

Xtraint = torch.FloatTensor(Xtrain)

ytraint = torch.LongTensor(ytrain)

Xtestt = torch.FloatTensor(Xtest)

ytestt = torch.LongTensor(ytest)

traindataset = TensorDataset(Xtraint, ytraint)

trainloader = DataLoader(traindataset, batchsize=16, shuffle=True)

Define model

class IrisNet(nn.Module):

def init(self, hiddensize=64):

super().init()

self.fc1 = nn.Linear(4, hiddensize)

self.fc2 = nn.Linear(hiddensize, 32)

self.fc3 = nn.Linear(32, 3)

self.relu = nn.ReLU()

def forward(self, x):

x = self.relu(self.fc1(x))

x = self.relu(self.fc2(x))

return self.fc3(x)

Training dengan MLflow

with mlflow.startrun():

# Hyperparameters

params = {

"hiddensize": 64,

"learningrate": 0.01,

"epochs": 100,

"batchsize": 16

}

mlflow.logparams(params)

model = IrisNet(hiddensize=params["hiddensize"])

criterion = nn.CrossEntropyLoss()

optimizer = optim.Adam(model.parameters(), lr=params["learningrate"])

# Training loop

for epoch in range(params["epochs"]):

model.train()

totalloss = 0

for batchX, batchy in trainloader:

optimizer.zerograd()

outputs = model(batchX)

loss = criterion(outputs, batchy)

loss.backward()

optimizer.step()

totalloss += loss.item()

# Log metrics per epoch

avgloss = totalloss / len(trainloader)

mlflow.logmetric("trainloss", avgloss, step=epoch)

# Evaluation

if epoch % 10 == 0:

model.eval()

with torch.nograd():

outputs = model(Xtestt)

, predicted = torch.max(outputs, 1)

accuracy = (predicted == ytestt).sum().item() / len(ytestt)

mlflow.logmetric("testaccuracy", accuracy, step=epoch)

# Log final model

mlflow.pytorch.logmodel(model, "model")

# Log model summary

mlflow.settag("modelarchitecture", str(model))

Bagian 2: Neptune.ai

2.1 Setup Neptune.ai

# Install Neptune

pip install neptune

Set API token (dari neptune.ai dashboard)

export NEPTUNEAPITOKEN="your-api-token"

2.2 Basic Experiment Tracking

import neptune

from sklearn.ensemble import RandomForestClassifier

from sklearn.datasets import loadiris

from sklearn.modelselection import traintestsplit

from sklearn.metrics import accuracyscore, f1score

Initialize Neptune run

run = neptune.initrun(

project="your-workspace/your-project",

apitoken="your-api-token", # Atau dari env variable

name="random-forest-v1",

tags=["classification", "sklearn", "iris"]

)

Load data

X, y = loadiris(returnXy=True)

Xtrain, Xtest, ytrain, ytest = traintestsplit(X, y, testsize=0.2)

Log parameters

params = {

"nestimators": 100,

"maxdepth": 5,

"randomstate": 42

}

run["parameters"] = params

Train model

model = RandomForestClassifier(params)

model.fit(Xtrain, ytrain)

Evaluate

ypred = model.predict(Xtest)

accuracy = accuracyscore(ytest, ypred)

f1 = f1score(ytest, ypred, average='weighted')

Log metrics

run["metrics/accuracy"] = accuracy

run["metrics/f1score"] = f1

Log feature importance

for name, importance in zip(loadiris().featurenames, model.featureimportances):

run[f"featureimportance/{name}"] = importance

Log model file

import joblib

joblib.dump(model, "model.pkl")

run["model"].upload("model.pkl")

Stop run

run.stop()

print(f"Accuracy: {accuracy:.4f}")

2.3 Tracking Training Progress (Series)

import neptune

import numpy as np

run = neptune.initrun(

project="your-workspace/your-project"

)

Simulate training loop

for epoch in range(100):

# Simulate metrics

trainloss = 1.0 / (epoch + 1) + np.random.random() 0.1

valloss = 1.2 / (epoch + 1) + np.random.random() 0.1

accuracy = 1 - valloss + np.random.random() * 0.05

# Log series data (auto-increments step)

run["train/loss"].append(trainloss)

run["val/loss"].append(valloss)

run["val/accuracy"].append(accuracy)

# Log dengan explicit step

run["train/epoch"].append(epoch)

run.stop()

2.4 Logging Artifacts dan Visualizations

import neptune

import matplotlib.pyplot as plt

import pandas as pd

from sklearn.metrics import confusionmatrix, ConfusionMatrixDisplay

from sklearn.datasets import loadiris

from sklearn.ensemble import RandomForestClassifier

from sklearn.modelselection import traintestsplit

run = neptune.initrun(project="your-workspace/your-project")

Load and train

X, y = loadiris(returnXy=True)

Xtrain, Xtest, ytrain, ytest = traintestsplit(X, y, testsize=0.2)

model = RandomForestClassifier(nestimators=100)

model.fit(Xtrain, ytrain)

ypred = model.predict(Xtest)

1. Log confusion matrix as image

cm = confusionmatrix(ytest, ypred)

disp = ConfusionMatrixDisplay(cm, displaylabels=loadiris().targetnames)

disp.plot()

plt.savefig("confusionmatrix.png")

run["visualizations/confusionmatrix"].upload("confusionmatrix.png")

2. Log DataFrame

df = pd.DataFrame({

"actual": ytest,

"predicted": ypred,

"correct": ytest == ypred

})

run["predictions"].upload(neptune.types.File.ashtml(df))

3. Log interactive chart dengan Neptune

from neptune.types import File

Feature importance bar chart

fig, ax = plt.subplots()

ax.barh(loadiris().featurenames, model.featureimportances)

ax.setxlabel("Importance")

ax.settitle("Feature Importance")

run["visualizations/featureimportance"].upload(fig)

4. Log source code

run["sourcecode"].upload("train.py")

5. Log dataset info

run["dataset/info"] = {

"name": "Iris",

"nsamples": len(X),

"nfeatures": X.shape[1],

"nclasses": len(set(y))

}

run.stop()

2.5 Neptune dengan PyTorch

import neptune

import torch

import torch.nn as nn

import torch.optim as optim

from torch.utils.data import DataLoader, TensorDataset

from sklearn.datasets import loadiris

from sklearn.modelselection import traintestsplit

from sklearn.preprocessing import StandardScaler

Initialize Neptune

run = neptune.initrun(

project="your-workspace/your-project",

name="pytorch-iris",

tags=["pytorch", "neural-network"]

)

Data preparation

X, y = loadiris(returnXy=True)

scaler = StandardScaler()

X = scaler.fittransform(X)

Xtrain, Xtest, ytrain, ytest = traintestsplit(X, y, testsize=0.2)

Xtraint = torch.FloatTensor(Xtrain)

ytraint = torch.LongTensor(ytrain)

Xtestt = torch.FloatTensor(Xtest)

ytestt = torch.LongTensor(ytest)

trainloader = DataLoader(TensorDataset(Xtraint, ytraint), batchsize=16)

Model

class Net(nn.Module):

def init(self):

super().init()

self.fc1 = nn.Linear(4, 64)

self.fc2 = nn.Linear(64, 32)

self.fc3 = nn.Linear(32, 3)

def forward(self, x):

x = torch.relu(self.fc1(x))

x = torch.relu(self.fc2(x))

return self.fc3(x)

model = Net()

criterion = nn.CrossEntropyLoss()

optimizer = optim.Adam(model.parameters(), lr=0.01)

Log model architecture

run["model/architecture"] = str(model)

run["parameters"] = {

"learningrate": 0.01,

"epochs": 100,

"batchsize": 16,

"optimizer": "Adam"

}

Training

for epoch in range(100):

model.train()

epochloss = 0

for batchX, batchy in trainloader:

optimizer.zerograd()

outputs = model(batchX)

loss = criterion(outputs, batchy)

loss.backward()

optimizer.step()

epochloss += loss.item()

avgloss = epochloss / len(trainloader)

run["train/loss"].append(avgloss)

# Validation

model.eval()

with torch.nograd():

outputs = model(Xtestt)

valloss = criterion(outputs, ytestt).item()

, predicted = torch.max(outputs, 1)

accuracy = (predicted == ytestt).sum().item() / len(ytestt)

run["val/loss"].append(valloss)

run["val/accuracy"].append(accuracy)

Save and log model

torch.save(model.statedict(), "model.pt")

run["model/weights"].upload("model.pt")

run.stop()

2.6 Neptune Model Registry

import neptune

from neptune.types import File

Initialize model registry

modelversion = neptune.initmodelversion(

model="YOUR-PROJECT-KEY-MOD", # Model ID dari Neptune

project="your-workspace/your-project"

)

Log model metadata

modelversion["model/framework"] = "scikit-learn"

modelversion["model/algorithm"] = "RandomForest"

Log model file

modelversion["model/binary"].upload("model.pkl")

Log performance metrics

modelversion["validation/accuracy"] = 0.95

modelversion["validation/f1score"] = 0.94

Log training run reference

modelversion["run/id"] = "YOUR-RUN-ID"

Change stage

modelversion.changestage("staging") # none, staging, production, archived

modelversion.stop()

Bagian 3: Perbandingan Praktis

3.1 Side-by-Side Code Comparison

Logging Parameters:
# MLflow

mlflow.logparams({

"learningrate": 0.01,

"epochs": 100

})

Neptune

run["parameters"] = {

"learningrate": 0.01,

"epochs": 100

}

Logging Metrics:
# MLflow

mlflow.logmetric("accuracy", 0.95)

mlflow.logmetrics({"loss": 0.1, "f1": 0.94})

Neptune

run["metrics/accuracy"] = 0.95

run["metrics/loss"] = 0.1

Logging Series (Training Loop):
# MLflow

for epoch in range(100):

mlflow.logmetric("loss", lossvalue, step=epoch)

Neptune

for epoch in range(100):

run["train/loss"].append(lossvalue) # Auto step

Logging Artifacts:
# MLflow

mlflow.logartifact("model.pkl")

mlflow.logartifacts("./outputs")

Neptune

run["model"].upload("model.pkl")

run["outputs"].uploadfiles("./outputs")

3.2 Unified Wrapper

from abc import ABC, abstractmethod

class ExperimentTracker(ABC):

@abstractmethod

def logparams(self, params: dict): pass

@abstractmethod

def logmetric(self, name: str, value: float, step: int = None): pass

@abstractmethod

def logartifact(self, path: str): pass

@abstractmethod

def endrun(self): pass

class MLflowTracker(ExperimentTracker):

def init(self, experimentname: str):

import mlflow

mlflow.setexperiment(experimentname)

mlflow.startrun()

self.mlflow = mlflow

def logparams(self, params: dict):

self.mlflow.logparams(params)

def logmetric(self, name: str, value: float, step: int = None):

self.mlflow.logmetric(name, value, step=step)

def logartifact(self, path: str):

self.mlflow.logartifact(path)

def endrun(self):

self.mlflow.endrun()

class NeptuneTracker(ExperimentTracker):

def init(self, project: str, apitoken: str = None):

import neptune

self.run = neptune.initrun(project=project, apitoken=apitoken)

def logparams(self, params: dict):

self.run["parameters"] = params

def logmetric(self, name: str, value: float, step: int = None):

self.run[f"metrics/{name}"].append(value)

def logartifact(self, path: str):

self.run["artifacts"].upload(path)

def endrun(self):

self.run.stop()

Usage - easily switch between trackers

def trainmodel(tracker: ExperimentTracker):

tracker.logparams({"lr": 0.01, "epochs": 100})

for epoch in range(100):

loss = 1.0 / (epoch + 1)

tracker.logmetric("loss", loss, step=epoch)

tracker.logartifact("model.pkl")

tracker.endrun()

Use MLflow

tracker = MLflowTracker("my-experiment")

trainmodel(tracker)

Or use Neptune

tracker = NeptuneTracker("workspace/project")

train_model(tracker)

3.3 Kapan Menggunakan Masing-masing?

Pilih MLflow jika:
  • Butuh self-hosted solution (data sensitivity)
  • Budget terbatas (infrastructure cost only)
  • Sudah punya infrastructure (Kubernetes, cloud)
  • Butuh integrasi dengan Databricks
  • Tim kecil dengan ML engineer yang bisa maintain

Pilih Neptune.ai jika:
  • Butuh quick start tanpa setup
  • Tim distributed yang butuh collaboration
  • Budget tersedia untuk managed service
  • Butuh advanced visualization
  • Fokus pada experiment tracking (bukan full MLOps)

Kesimpulan

| Feature | MLflow | Neptune.ai |

|---------|--------|------------|

| Setup | Butuh effort | Instant |

| Cost | Infra only | Subscription |

| UI | Functional | Modern |

| Collaboration | Basic | Excellent |

| Flexibility | High | Medium |

| Learning Curve | Medium | Low |

Rekomendasi:
  • Startup/Small team: Neptune.ai (free tier cukup untuk mulai)
  • Enterprise/On-prem: MLflow (full control)
  • Hybrid: Gunakan keduanya dengan unified wrapper

Kedua platform sama-sama excellent untuk experiment tracking. Pilihan tergantung pada kebutuhan spesifik tim dan organisasi Anda.

Artikel Terkait

Tutorial Lengkap Comet ML: Platform MLOps untuk Experiment Tracking dan Model Management

Tutorial Lengkap Comet ML: Platform MLOps untuk Experiment Tracking dan Model Management Dalam dunia machine learning mo...

Tutorial Integrasi Azure MLflow: Experiment Tracking di Azure

Tutorial Lengkap Azure MLflow Integration: Experiment Tracking dan Model Management Azure Machine Learning menyediakan i...

Tutorial Lengkap Weights & Biases: Experiment Tracking untuk Machine Learning

Tutorial Lengkap Weights & Biases: ML Experiment Tracking dan Visualization Weights & Biases (W&B) adalah platform MLOps...

Tutorial Lengkap MLflow: Dari Setup hingga Production

Pendahuluan MLflow adalah platform open-source untuk mengelola end-to-end machine learning lifecycle. Dikembangkan oleh ...