Tutorial Lengkap Azure DevOps untuk MLOps: CI/CD untuk Machine Learning
Azure DevOps menyediakan kemampuan CI/CD komprehensif untuk proyek machine learning. Tutorial ini mencakup pembangunan pipeline ML otomatis, deployment model, dan continuous delivery menggunakan Azure DevOps.
Mengapa Azure DevOps untuk MLOps?
Manfaat Utama:- Automasi end-to-end: Dari kode hingga deployment
- Version control: Git repos untuk kode dan data
- Orkestrasi pipeline: Workflow ML multi-stage
- Integration: Integrasi native Azure ML
- Collaboration: Pengembangan berbasis tim
- Azure Repos: Git repositories
- Azure Pipelines: Automasi CI/CD
- Azure Artifacts: Package management
- Azure Boards: Work tracking
Prerequisites
pip install azure-devops azure-ai-ml
Azure CLI
az login
az extension add --name azure-devops
az devops configure --defaults organization=https://dev.azure.com/myorg
Setup Proyek
1. Buat DevOps Project
# Buat project
az devops project create --name "MLOps-Project" --org https://dev.azure.com/myorg
Buat repository
az repos create --name "ml-models" --project "MLOps-Project"
2. Struktur Repository
ml-models/
├── src/
│ ├── train.py
│ ├── evaluate.py
│ └── score.py
├── tests/
│ └── testmodel.py
├── pipelines/
│ ├── train-pipeline.yml
│ ├── deploy-pipeline.yml
│ └── cd-pipeline.yml
├── infrastructure/
│ └── arm-templates/
├── environment.yml
├── requirements.txt
└── azure-pipelines.yml
3. Service Connection
# Buat service connection ke Azure
az devops service-endpoint azurerm create \
--azure-rm-service-principal-id "your-sp-id" \
--azure-rm-subscription-id "your-subscription-id" \
--azure-rm-subscription-name "Your Subscription" \
--azure-rm-tenant-id "your-tenant-id" \
--name "azure-ml-connection"
CI Pipeline untuk ML
1. Basic CI Pipeline
# azure-pipelines.yml
trigger:
branches:
include:
- main
- develop
paths:
include:
- src/
- tests/
pool:
vmImage: 'ubuntu-latest'
variables:
pythonVersion: '3.9'
stages:
- stage: Build
displayName: 'Build dan Test'
jobs:
- job: BuildJob
steps:
- task: UsePythonVersion@0
inputs:
versionSpec: '$(pythonVersion)'
displayName: 'Gunakan Python $(pythonVersion)'
- script: |
python -m pip install --upgrade pip
pip install -r requirements.txt
pip install pytest pytest-cov
displayName: 'Install dependencies'
- script: |
python -m pytest tests/ --cov=src --cov-report=xml
displayName: 'Jalankan tests'
- task: PublishTestResults@2
inputs:
testResultsFiles: '*/test-.xml'
testRunTitle: 'Python Tests'
- task: PublishCodeCoverageResults@1
inputs:
codeCoverageTool: Cobertura
summaryFileLocation: '$(System.DefaultWorkingDirectory)/*/coverage.xml'
2. Linting dan Code Quality
# Tambahkan ke azure-pipelines.yml
- script: |
pip install flake8 black mypy
flake8 src/ --max-line-length=100
black --check src/
mypy src/
displayName: 'Pemeriksaan kualitas kode'
Training Pipeline
1. ML Training Pipeline
# pipelines/train-pipeline.yml
trigger:
branches:
include:
- main
paths:
include:
- src/
- data/
variables:
- group: ml-variables
- name: resourceGroup
value: 'ml-rg'
- name: workspaceName
value: 'ml-workspace'
- name: computeName
value: 'cpu-cluster'
stages:
- stage: Train
displayName: 'Train Model'
jobs:
- job: TrainJob
pool:
vmImage: 'ubuntu-latest'
steps:
- task: UsePythonVersion@0
inputs:
versionSpec: '3.9'
- task: AzureCLI@2
displayName: 'Install Azure ML CLI'
inputs:
azureSubscription: 'azure-ml-connection'
scriptType: 'bash'
scriptLocation: 'inlineScript'
inlineScript: |
az extension add -n ml
- task: AzureCLI@2
displayName: 'Submit Training Job'
inputs:
azureSubscription: 'azure-ml-connection'
scriptType: 'bash'
scriptLocation: 'inlineScript'
inlineScript: |
az ml job create \
--file jobs/train-job.yml \
--resource-group $(resourceGroup) \
--workspace-name $(workspaceName) \
--set compute=$(computeName)
2. Definisi Training Job
# jobs/train-job.yml
$schema: https://azuremlschemas.azureedge.net/latest/commandJob.schema.json
type: command
code: ./src
command: python train.py --data ${{inputs.data}} --output ${{outputs.model}}
inputs:
data:
type: urifolder
path: azureml:training-data:1
outputs:
model:
type: urifolder
environment: azureml:sklearn-env:1
compute: azureml:cpu-cluster
experimentname: mlops-training
displayname: model-training-$(Build.BuildId)
3. Training Script
# src/train.py
import argparse
import os
import mlflow
import pandas as pd
from sklearn.ensemble import RandomForestClassifier
from sklearn.modelselection import traintestsplit
from sklearn.metrics import accuracyscore, f1score
import joblib
def main():
parser = argparse.ArgumentParser()
parser.addargument("--data", type=str, required=True)
parser.addargument("--output", type=str, required=True)
args = parser.parseargs()
# Aktifkan autologging
mlflow.autolog()
# Load data
df = pd.readcsv(os.path.join(args.data, "train.csv"))
X = df.drop("target", axis=1)
y = df["target"]
Xtrain, Xtest, ytrain, ytest = traintestsplit(
X, y, testsize=0.2, randomstate=42
)
# Train model
with mlflow.startrun():
model = RandomForestClassifier(nestimators=100, randomstate=42)
model.fit(Xtrain, ytrain)
# Evaluasi
predictions = model.predict(Xtest)
accuracy = accuracyscore(ytest, predictions)
f1 = f1score(ytest, predictions, average="weighted")
mlflow.logmetric("accuracy", accuracy)
mlflow.logmetric("f1score", f1)
# Simpan model
os.makedirs(args.output, existok=True)
joblib.dump(model, os.path.join(args.output, "model.joblib"))
print(f"Accuracy: {accuracy:.4f}")
print(f"F1 Score: {f1:.4f}")
if name == "main":
main()
Pipeline Registrasi Model
1. Pipeline Register Model
# pipelines/register-pipeline.yml
stages:
- stage: Register
displayName: 'Register Model'
dependsOn: Train
jobs:
- job: RegisterJob
steps:
- task: AzureCLI@2
displayName: 'Register Model'
inputs:
azureSubscription: 'azure-ml-connection'
scriptType: 'bash'
scriptLocation: 'inlineScript'
inlineScript: |
# Dapatkan job terbaru
JOBNAME=$(az ml job list \
--resource-group $(resourceGroup) \
--workspace-name $(workspaceName) \
--query "[0].name" -o tsv)
# Register model
az ml model create \
--name my-model \
--version $(Build.BuildId) \
--path azureml://jobs/$JOBNAME/outputs/model \
--resource-group $(resourceGroup) \
--workspace-name $(workspaceName)
2. Model dengan Metadata
# models/model.yml
$schema: https://azuremlschemas.azureedge.net/latest/model.schema.json
name: classification-model
path: azureml://jobs/{jobname}/outputs/model
description: Random Forest classifier ditraining pada data customer
properties:
accuracy: 0.95
framework: sklearn
task: classification
tags:
team: data-science
environment: production
Pipeline Deployment
1. Deploy ke Online Endpoint
# pipelines/deploy-pipeline.yml
stages:
- stage: DeployStaging
displayName: 'Deploy ke Staging'
jobs:
- deployment: DeployStaging
environment: staging
strategy:
runOnce:
deploy:
steps:
- task: AzureCLI@2
displayName: 'Buat/Update Endpoint'
inputs:
azureSubscription: 'azure-ml-connection'
scriptType: 'bash'
scriptLocation: 'inlineScript'
inlineScript: |
# Buat endpoint jika belum ada
az ml online-endpoint create \
--name staging-endpoint \
--resource-group $(resourceGroup) \
--workspace-name $(workspaceName) \
--file endpoints/staging-endpoint.yml \
|| true
# Buat/update deployment
az ml online-deployment create \
--name blue \
--endpoint-name staging-endpoint \
--resource-group $(resourceGroup) \
--workspace-name $(workspaceName) \
--file endpoints/staging-deployment.yml
# Set traffic
az ml online-endpoint update \
--name staging-endpoint \
--resource-group $(resourceGroup) \
--workspace-name $(workspaceName) \
--traffic "blue=100"
- stage: TestStaging
displayName: 'Test Staging'
dependsOn: DeployStaging
jobs:
- job: SmokeTest
steps:
- task: AzureCLI@2
displayName: 'Jalankan Smoke Tests'
inputs:
azureSubscription: 'azure-ml-connection'
scriptType: 'bash'
scriptLocation: 'inlineScript'
inlineScript: |
# Dapatkan endpoint key
KEY=$(az ml online-endpoint get-credentials \
--name staging-endpoint \
--resource-group $(resourceGroup) \
--workspace-name $(workspaceName) \
--query "primaryKey" -o tsv)
# Dapatkan endpoint URL
URL=$(az ml online-endpoint show \
--name staging-endpoint \
--resource-group $(resourceGroup) \
--workspace-name $(workspaceName) \
--query "scoringuri" -o tsv)
# Test endpoint
curl -X POST "$URL" \
-H "Authorization: Bearer $KEY" \
-H "Content-Type: application/json" \
-d '{"features": [[5.1, 3.5, 1.4, 0.2]]}'
- stage: DeployProduction
displayName: 'Deploy ke Production'
dependsOn: TestStaging
condition: and(succeeded(), eq(variables['Build.SourceBranch'], 'refs/heads/main'))
jobs:
- deployment: DeployProduction
environment: production
strategy:
runOnce:
deploy:
steps:
- task: AzureCLI@2
displayName: 'Deploy ke Production'
inputs:
azureSubscription: 'azure-ml-connection'
scriptType: 'bash'
scriptLocation: 'inlineScript'
inlineScript: |
az ml online-deployment create \
--name blue-$(Build.BuildId) \
--endpoint-name production-endpoint \
--resource-group $(resourceGroup) \
--workspace-name $(workspaceName) \
--file endpoints/production-deployment.yml
2. Konfigurasi Endpoint
# endpoints/production-endpoint.yml
$schema: https://azuremlschemas.azureedge.net/latest/managedOnlineEndpoint.schema.json
name: production-endpoint
authmode: key
endpoints/production-deployment.yml
$schema: https://azuremlschemas.azureedge.net/latest/managedOnlineDeployment.schema.json
name: blue
endpointname: production-endpoint
model: azureml:classification-model:latest
instancetype: StandardDS3v2
instancecount: 2
codeconfiguration:
code: ./scoring
scoringscript: score.py
environment: azureml:sklearn-env:1
Blue-Green Deployment
1. Pipeline Blue-Green
# pipelines/blue-green-pipeline.yml
stages:
- stage: DeployGreen
displayName: 'Deploy Green'
jobs:
- job: DeployGreen
steps:
- task: AzureCLI@2
displayName: 'Deploy Versi Green'
inputs:
azureSubscription: 'azure-ml-connection'
scriptType: 'bash'
scriptLocation: 'inlineScript'
inlineScript: |
# Deploy versi baru sebagai green
az ml online-deployment create \
--name green \
--endpoint-name production-endpoint \
--resource-group $(resourceGroup) \
--workspace-name $(workspaceName) \
--file endpoints/green-deployment.yml
- stage: TestGreen
displayName: 'Test Green'
dependsOn: DeployGreen
jobs:
- job: TestGreen
steps:
- task: AzureCLI@2
displayName: 'Test Green Deployment'
inputs:
azureSubscription: 'azure-ml-connection'
scriptType: 'bash'
scriptLocation: 'inlineScript'
inlineScript: |
python tests/integrationtests.py --deployment green
- stage: ShiftTraffic
displayName: 'Shift Traffic'
dependsOn: TestGreen
jobs:
- job: ShiftTraffic
steps:
- task: AzureCLI@2
displayName: 'Traffic Shift Bertahap'
inputs:
azureSubscription: 'azure-ml-connection'
scriptType: 'bash'
scriptLocation: 'inlineScript'
inlineScript: |
# 10% ke green
az ml online-endpoint update \
--name production-endpoint \
--traffic "blue=90 green=10"
sleep 300 # Monitor 5 menit
# 50% ke green
az ml online-endpoint update \
--name production-endpoint \
--traffic "blue=50 green=50"
sleep 300
# 100% ke green
az ml online-endpoint update \
--name production-endpoint \
--traffic "blue=0 green=100"
- stage: Cleanup
displayName: 'Cleanup Deployment Lama'
dependsOn: ShiftTraffic
jobs:
- job: Cleanup
steps:
- task: AzureCLI@2
displayName: 'Hapus Blue Deployment'
inputs:
azureSubscription: 'azure-ml-connection'
scriptType: 'bash'
scriptLocation: 'inlineScript'
inlineScript: |
az ml online-deployment delete \
--name blue \
--endpoint-name production-endpoint \
--yes
Pipeline Monitoring
1. Model Monitoring
# pipelines/monitor-pipeline.yml
schedules:
- cron: "0 0 " # Harian tengah malam
displayName: Monitoring harian
branches:
include:
- main
stages:
- stage: Monitor
displayName: 'Model Monitoring'
jobs:
- job: MonitorJob
steps:
- task: AzureCLI@2
displayName: 'Cek Model Drift'
inputs:
azureSubscription: 'azure-ml-connection'
scriptType: 'bash'
scriptLocation: 'inlineScript'
inlineScript: |
python monitoring/checkdrift.py \
--endpoint production-endpoint \
--threshold 0.1
- task: AzureCLI@2
displayName: 'Cek Performance'
inputs:
azureSubscription: 'azure-ml-connection'
scriptType: 'bash'
scriptLocation: 'inlineScript'
inlineScript: |
python monitoring/checkperformance.py \
--endpoint production-endpoint \
--accuracy-threshold 0.85
2. Script Deteksi Drift
# monitoring/checkdrift.py
import argparse
from azure.ai.ml import MLClient
from azure.identity import DefaultAzureCredential
def checkdrift(endpointname, threshold):
mlclient = MLClient.fromconfig(credential=DefaultAzureCredential())
# Dapatkan prediksi terbaru
# Bandingkan dengan distribusi training
# Alert jika drift terdeteksi
driftscore = calculatedrift()
if driftscore > threshold:
print(f"##vso[task.logissue type=warning]Data drift terdeteksi: {driftscore}")
print("##vso[task.setvariable variable=driftDetected]true")
else:
print(f"Drift score: {driftscore} - dalam rentang yang diterima")
print("##vso[task.setvariable variable=driftDetected]false")
if name == "main":
parser = argparse.ArgumentParser()
parser.addargument("--endpoint", required=True)
parser.addargument("--threshold", type=float, default=0.1)
args = parser.parseargs()
checkdrift(args.endpoint, args.threshold)
Variable Groups
1. Buat Variable Group
# Buat variable group
az pipelines variable-group create \
--name "ml-variables" \
--variables \
AZURESUBSCRIPTIONID="your-subscription-id" \
AZURERESOURCEGROUP="ml-rg" \
AZUREMLWORKSPACE="ml-workspace" \
--authorize
2. Gunakan di Pipeline
variables:
- group: ml-variables
- name: environment
value: 'production'
steps:
- script: |
echo "Deploying ke $(AZUREML_WORKSPACE)"
Best Practices
1. Pipeline Templates
# templates/train-template.yml
parameters:
- name: computeTarget
type: string
- name: experimentName
type: string
steps:
- task: AzureCLI@2
inputs:
scriptType: 'bash'
inlineScript: |
az ml job create \
--compute ${{ parameters.computeTarget }} \
--experiment-name ${{ parameters.experimentName }}
Penggunaan di pipeline utama
stages:
- stage: Train
jobs:
- template: templates/train-template.yml
parameters:
computeTarget: cpu-cluster
experimentName: my-experiment
2. Environment Approvals
# Konfigurasi environment dengan approvals
environments:
- name: production
resourceName: production
resourceType: VirtualMachine
tags: production
Kesimpulan
Azure DevOps untuk MLOps menyediakan:
Key takeaways:
- Struktur repos untuk proyek ML
- Gunakan pipeline multi-stage
- Implementasikan testing yang proper
- Gunakan blue-green deployments
- Monitor model secara berkelanjutan