Tutorial Lengkap Azure Machine Learning: End-to-End ML Platform

# Tutorial Lengkap Azure Machine Learning: ML End-to-End di Azure Azure Machine Learning adalah platform berbasis cloud untuk membangun, melatih, dan mendeploy model machine learning. Platform ini me...

By Ruby Abdullah · · tutorial
AzureAzure MLMLOpsCloud MLPythonMachine Learning

Tutorial Lengkap Azure Machine Learning: ML End-to-End di Azure

Azure Machine Learning adalah platform berbasis cloud untuk membangun, melatih, dan mendeploy model machine learning. Platform ini menyediakan environment MLOps komprehensif dengan tools untuk seluruh lifecycle ML.

Mengapa Azure Machine Learning?

Manfaat Utama:
  • Platform terpadu: Manajemen lifecycle ML lengkap
  • AutoML: Kemampuan machine learning otomatis
  • Enterprise-ready: Security, compliance, dan governance
  • Compute fleksibel: Dari notebooks hingga GPU clusters
  • Integrasi: Konektivitas seamless dengan ekosistem Azure

Komponen Utama:
  • Workspaces
  • Compute instances dan clusters
  • Datastores dan datasets
  • Experiments dan runs
  • Models dan endpoints

Prerequisites

pip install azure-ai-ml azure-identity

Azure CLI

az login

az extension add -n ml

Quick Start

1. Buat Workspace

from azure.ai.ml import MLClient

from azure.identity import DefaultAzureCredential

from azure.ai.ml.entities import Workspace

Autentikasi

credential = DefaultAzureCredential()

Buat workspace

workspace = Workspace(

name="my-ml-workspace",

location="eastus",

displayname="ML Workspace",

description="Azure ML workspace untuk proyek ML"

)

Buat ML client

mlclient = MLClient(

credential=credential,

subscriptionid="your-subscription-id",

resourcegroupname="my-resource-group"

)

Buat workspace

mlclient.workspaces.begincreateorupdate(workspace).result()

print(f"Workspace dibuat: {workspace.name}")

2. Koneksi ke Workspace yang Ada

from azure.ai.ml import MLClient

from azure.identity import DefaultAzureCredential

mlclient = MLClient(

credential=DefaultAzureCredential(),

subscriptionid="your-subscription-id",

resourcegroupname="my-resource-group",

workspacename="my-ml-workspace"

)

print(f"Terhubung ke workspace: {mlclient.workspacename}")

Compute Resources

1. Buat Compute Instance

from azure.ai.ml.entities import ComputeInstance

computeinstance = ComputeInstance(

name="my-compute-instance",

size="StandardDS3v2",

idletimebeforeshutdownminutes=60

)

mlclient.compute.begincreateorupdate(computeinstance).result()

print("Compute instance dibuat")

2. Buat Compute Cluster

from azure.ai.ml.entities import AmlCompute

computecluster = AmlCompute(

name="cpu-cluster",

type="amlcompute",

size="StandardDS3v2",

mininstances=0,

maxinstances=4,

idletimebeforescaledown=120

)

mlclient.compute.begincreateorupdate(computecluster).result()

print("Compute cluster dibuat")

3. GPU Cluster

gpucluster = AmlCompute(

name="gpu-cluster",

type="amlcompute",

size="StandardNC6",

mininstances=0,

maxinstances=2,

idletimebeforescaledown=300

)

mlclient.compute.begincreateorupdate(gpucluster).result()

Manajemen Data

1. Register Datastore

from azure.ai.ml.entities import AzureBlobDatastore

datastore = AzureBlobDatastore(

name="my-blob-datastore",

accountname="mystorageaccount",

containername="ml-data",

credentials={

"accountkey": "your-account-key"

}

)

mlclient.datastores.createorupdate(datastore)

print("Datastore terdaftar")

2. Buat Dataset

from azure.ai.ml.entities import Data

from azure.ai.ml.constants import AssetTypes

File dataset

filedata = Data(

name="training-data",

path="azureml://datastores/my-blob-datastore/paths/data/train.csv",

type=AssetTypes.URIFILE,

description="Dataset training"

)

mlclient.data.createorupdate(filedata)

Folder dataset

folderdata = Data(

name="image-dataset",

path="azureml://datastores/my-blob-datastore/paths/images/",

type=AssetTypes.URIFOLDER,

description="Dataset gambar"

)

mlclient.data.createorupdate(folderdata)

3. Akses Data di Jobs

from azure.ai.ml import Input

Gunakan di command job

jobinputs = {

"trainingdata": Input(

type="urifile",

path="azureml://datastores/my-blob-datastore/paths/data/train.csv"

)

}

Training Jobs

1. Command Job

from azure.ai.ml import command, Input, Output

Definisikan training job

trainingjob = command(

code="./src",

command="python train.py --data ${{inputs.data}} --output ${{outputs.model}}",

inputs={

"data": Input(

type="urifile",

path="azureml://datastores/workspaceblobstore/paths/data/train.csv"

)

},

outputs={

"model": Output(type="urifolder", path="azureml://datastores/workspaceblobstore/paths/models/")

},

environment="AzureML-sklearn-1.0-ubuntu20.04-py38-cpu@latest",

compute="cpu-cluster",

displayname="sklearn-training",

experimentname="my-experiment"

)

Submit job

returnedjob = mlclient.jobs.createorupdate(trainingjob)

print(f"Job disubmit: {returnedjob.name}")

Tunggu selesai

mlclient.jobs.stream(returnedjob.name)

2. Training Script

# src/train.py

import argparse

import pandas as pd

from sklearn.ensemble import RandomForestClassifier

from sklearn.modelselection import traintestsplit

from sklearn.metrics import accuracyscore

import joblib

import os

import mlflow

def main():

parser = argparse.ArgumentParser()

parser.addargument("--data", type=str, required=True)

parser.addargument("--output", type=str, required=True)

args = parser.parseargs()

# Aktifkan MLflow autologging

mlflow.autolog()

# Load data

df = pd.readcsv(args.data)

X = df.drop("target", axis=1)

y = df["target"]

# Split data

Xtrain, Xtest, ytrain, ytest = traintestsplit(

X, y, testsize=0.2, randomstate=42

)

# Train model

model = RandomForestClassifier(nestimators=100, randomstate=42)

model.fit(Xtrain, ytrain)

# Evaluasi

predictions = model.predict(Xtest)

accuracy = accuracyscore(ytest, predictions)

print(f"Akurasi: {accuracy}")

# Simpan model

os.makedirs(args.output, existok=True)

joblib.dump(model, os.path.join(args.output, "model.joblib"))

if name == "main":

main()

3. Custom Environment

from azure.ai.ml.entities import Environment, BuildContext

Dari spesifikasi conda

env = Environment(

name="sklearn-env",

description="Scikit-learn environment",

condafile="./environment.yml",

image="mcr.microsoft.com/azureml/openmpi4.1.0-ubuntu20.04:latest"

)

mlclient.environments.createorupdate(env)

Dari Dockerfile

dockerenv = Environment(

name="custom-env",

build=BuildContext(path="./docker-context"),

description="Custom Docker environment"

)

mlclient.environments.createorupdate(dockerenv)

AutoML

1. Classification

from azure.ai.ml import automl, Input

Konfigurasi AutoML classification

classificationjob = automl.classification(

compute="cpu-cluster",

experimentname="automl-classification",

trainingdata=Input(type="mltable", path="./data/train"),

targetcolumnname="target",

primarymetric="accuracy",

ncrossvalidations=5,

enablemodelexplainability=True

)

Set limits

classificationjob.setlimits(

timeoutminutes=60,

trialtimeoutminutes=20,

maxtrials=20,

maxconcurrenttrials=4

)

Set training settings

classificationjob.settraining(

enablestackensemble=True,

enablevoteensemble=True

)

Submit job

returnedjob = mlclient.jobs.createorupdate(classificationjob)

2. Regression

regressionjob = automl.regression(

compute="cpu-cluster",

experimentname="automl-regression",

trainingdata=Input(type="mltable", path="./data/train"),

targetcolumnname="price",

primarymetric="r2score",

ncrossvalidations=5

)

regressionjob.setlimits(

timeoutminutes=120,

trialtimeoutminutes=30,

maxtrials=30

)

mlclient.jobs.createorupdate(regressionjob)

3. Forecasting

forecastingjob = automl.forecasting(

compute="cpu-cluster",

experimentname="automl-forecasting",

trainingdata=Input(type="mltable", path="./data/timeseries"),

targetcolumnname="sales",

primarymetric="normalizedrootmeansquarederror",

ncrossvalidations=3

)

Konfigurasi forecasting settings

forecastingjob.setforecastsettings(

timecolumnname="date",

forecasthorizon=30,

frequency="D"

)

mlclient.jobs.createorupdate(forecastingjob)

Registrasi Model

1. Register Model

from azure.ai.ml.entities import Model

from azure.ai.ml.constants import AssetTypes

Register dari job output

model = Model(

path=f"azureml://jobs/{returnedjob.name}/outputs/model/",

name="sklearn-classifier",

description="Random Forest classifier",

type=AssetTypes.CUSTOMMODEL

)

registeredmodel = mlclient.models.createorupdate(model)

print(f"Model terdaftar: {registeredmodel.name}:{registeredmodel.version}")

2. Register MLflow Model

mlflowmodel = Model(

path="runs:/runid/model",

name="mlflow-model",

type=AssetTypes.MLFLOWMODEL,

description="Model MLflow yang dilog"

)

mlclient.models.createorupdate(mlflowmodel)

3. List Models

# List semua models

models = mlclient.models.list()

for model in models:

print(f"{model.name}: {model.latestversion}")

Dapatkan versi spesifik

model = mlclient.models.get("sklearn-classifier", version="1")

print(f"Model path: {model.path}")

Model Deployment

1. Online Endpoint

from azure.ai.ml.entities import (

ManagedOnlineEndpoint,

ManagedOnlineDeployment,

Model,

CodeConfiguration

)

Buat endpoint

endpoint = ManagedOnlineEndpoint(

name="sklearn-endpoint",

description="Sklearn model endpoint",

authmode="key"

)

mlclient.onlineendpoints.begincreateorupdate(endpoint).result()

print("Endpoint dibuat")

Buat deployment

deployment = ManagedOnlineDeployment(

name="blue",

endpointname="sklearn-endpoint",

model="sklearn-classifier:1",

instancetype="StandardDS3v2",

instancecount=1,

codeconfiguration=CodeConfiguration(

code="./scoring",

scoringscript="score.py"

),

environment="AzureML-sklearn-1.0-ubuntu20.04-py38-cpu@latest"

)

mlclient.onlinedeployments.begincreateorupdate(deployment).result()

Set traffic

endpoint.traffic = {"blue": 100}

mlclient.onlineendpoints.begincreateorupdate(endpoint).result()

2. Scoring Script

# scoring/score.py

import json

import joblib

import numpy as np

import os

def init():

global model

modelpath = os.path.join(os.getenv("AZUREMLMODELDIR"), "model.joblib")

model = joblib.load(modelpath)

def run(rawdata):

try:

data = json.loads(rawdata)

features = np.array(data["features"])

predictions = model.predict(features)

return {"predictions": predictions.tolist()}

except Exception as e:

return {"error": str(e)}

3. Test Endpoint

import json

Siapkan test data

testdata = {

"features": [[5.1, 3.5, 1.4, 0.2], [6.2, 3.4, 5.4, 2.3]]

}

Panggil endpoint

response = mlclient.onlineendpoints.invoke(

endpointname="sklearn-endpoint",

requestfile=json.dumps(testdata)

)

print(f"Response: {response}")

Batch Endpoints

1. Buat Batch Endpoint

from azure.ai.ml.entities import BatchEndpoint, BatchDeployment

Buat endpoint

batchendpoint = BatchEndpoint(

name="batch-sklearn-endpoint",

description="Batch inference endpoint"

)

mlclient.batchendpoints.begincreateorupdate(batchendpoint).result()

Buat deployment

batchdeployment = BatchDeployment(

name="batch-deployment",

endpointname="batch-sklearn-endpoint",

model="sklearn-classifier:1",

compute="cpu-cluster",

instancecount=2,

maxconcurrencyperinstance=2,

minibatchsize=10,

outputaction="appendrow",

outputfilename="predictions.csv",

codeconfiguration=CodeConfiguration(

code="./batch-scoring",

scoringscript="batchscore.py"

),

environment="AzureML-sklearn-1.0-ubuntu20.04-py38-cpu@latest"

)

mlclient.batchdeployments.begincreateorupdate(batchdeployment).result()

2. Invoke Batch Endpoint

from azure.ai.ml import Input

Mulai batch job

job = mlclient.batchendpoints.invoke(

endpointname="batch-sklearn-endpoint",

input=Input(

path="azureml://datastores/workspaceblobstore/paths/batch-data/",

type="urifolder"

)

)

print(f"Batch job dimulai: {job.name}")

Monitor job

mlclient.jobs.stream(job.name)

Integrasi MLflow

1. Track Experiments

import mlflow

from azure.ai.ml import MLClient

Set tracking URI

mlclient = MLClient.fromconfig()

mlflow.settrackinguri(mlclient.workspaces.get().mlflowtrackinguri)

Mulai experiment

mlflow.setexperiment("my-experiment")

with mlflow.startrun():

# Log parameters

mlflow.logparam("learningrate", 0.01)

mlflow.logparam("epochs", 100)

# Train model

# ...

# Log metrics

mlflow.logmetric("accuracy", 0.95)

mlflow.logmetric("f1score", 0.93)

# Log model

mlflow.sklearn.logmodel(model, "model")

# Log artifacts

mlflow.logartifact("plots/confusionmatrix.png")

2. Query Runs

# Search runs

runs = mlflow.searchruns(

experimentnames=["my-experiment"],

filterstring="metrics.accuracy > 0.9",

orderby=["metrics.accuracy DESC"]

)

print(runs[["runid", "metrics.accuracy", "params.learningrate"]])

Best Practices

1. Resource Cleanup

# Hapus endpoint

mlclient.onlineendpoints.begindelete("sklearn-endpoint").result()

Hapus compute

mlclient.compute.begindelete("cpu-cluster").result()

Archive model

mlclient.models.archive("sklearn-classifier")

2. Manajemen Biaya

# Auto-scale compute cluster

cluster = AmlCompute(

name="cost-optimized-cluster",

size="StandardDS3v2",

mininstances=0, # Scale ke zero saat idle

maxinstances=4,

idletimebeforescaledown=120,

tier="LowPriority" # Gunakan spot instances

)

Kesimpulan

Azure Machine Learning menyediakan:

  • Platform terpadu: Lifecycle ML lengkap
  • AutoML: Pembangunan model otomatis
  • Compute scalable: Resource fleksibel
  • MLOps: Deployment dan monitoring
  • Integrasi: Ekosistem Azure
  • Key takeaways:

    • Gunakan compute clusters untuk skalabilitas
    • Manfaatkan AutoML untuk prototyping cepat
    • Track experiments dengan MLflow
    • Deploy dengan managed endpoints
    • Monitor biaya dengan auto-scaling

    Artikel Terkait

    Tutorial Lengkap Vertex AI: Platform ML Terpadu Google Cloud

    Tutorial Lengkap Vertex AI: Platform ML Terpadu di Google Cloud Vertex AI adalah platform machine learning terpadu Googl...

    Tutorial Lengkap AWS SageMaker: Machine Learning di Cloud

    Tutorial Lengkap AWS SageMaker: End-to-End ML Pipeline Amazon SageMaker adalah layanan machine learning terkelola penuh ...

    Tutorial Lengkap Comet ML: Platform MLOps untuk Experiment Tracking dan Model Management

    Tutorial Lengkap Comet ML: Platform MLOps untuk Experiment Tracking dan Model Management Dalam dunia machine learning mo...

    Tutorial Azure Databricks untuk ML: Unified Analytics Platform

    Tutorial Lengkap Azure Databricks untuk ML: Platform Analytics Terpadu Azure Databricks menyediakan platform analytics k...