Azure ML Pipelines Tutorial: ML Pipeline Automation

# Tutorial Lengkap Azure ML Pipelines: CI/CD untuk Machine Learning Azure ML Pipelines memungkinkan Anda membangun workflow machine learning yang reproducible dan reusable. Pipeline mengotomatisasi l...

By Ruby Abdullah · · tutorial
AzureAzure MLPipelinesMLOpsAutomationCI/CD

Complete Azure ML Pipelines Tutorial: CI/CD for Machine Learning

Azure ML Pipelines enable you to build reproducible, reusable machine learning workflows. They automate the end-to-end ML lifecycle from data preparation to model deployment with version control and collaboration.

Why Azure ML Pipelines?

Key Benefits:
  • Reproducibility: Version-controlled workflows
  • Reusability: Modular pipeline components
  • Automation: Scheduled and triggered pipelines
  • Collaboration: Team-based development
  • Integration: Azure DevOps and GitHub Actions

Use Cases:
  • Automated model training
  • Data preprocessing workflows
  • Batch inference pipelines
  • MLOps CI/CD
  • Feature engineering automation

Prerequisites

pip install azure-ai-ml azure-identity

Azure CLI

az login

az extension add -n ml

Quick Start

1. Connect to Workspace

from azure.ai.ml import MLClient

from azure.identity import DefaultAzureCredential

mlclient = MLClient(

credential=DefaultAzureCredential(),

subscriptionid="your-subscription-id",

resourcegroupname="my-resource-group",

workspacename="my-ml-workspace"

)

2. Simple Pipeline

from azure.ai.ml import dsl, Input, Output

from azure.ai.ml.entities import Pipeline

Define pipeline

@dsl.pipeline(

compute="cpu-cluster",

description="Simple training pipeline"

)

def trainingpipeline(trainingdata):

# Preprocessing step

preprocessstep = preprocesscomponent(

inputdata=trainingdata

)

# Training step

trainstep = traincomponent(

trainingdata=preprocessstep.outputs.outputdata

)

return {

"modeloutput": trainstep.outputs.model

}

Create pipeline

pipeline = trainingpipeline(

trainingdata=Input(type="urifile", path="azureml:my-dataset:1")

)

Submit pipeline

pipelinejob = mlclient.jobs.createorupdate(

pipeline,

experimentname="training-pipeline"

)

print(f"Pipeline submitted: {pipelinejob.name}")

Pipeline Components

1. Create Component from Code

from azure.ai.ml import command

from azure.ai.ml.entities import Component

Data preprocessing component

preprocesscomponent = command(

name="preprocessdata",

displayname="Preprocess Data",

description="Clean and prepare data for training",

inputs={

"inputdata": Input(type="urifile")

},

outputs={

"outputdata": Output(type="urifolder")

},

code="./components/preprocess",

command="python preprocess.py --input ${{inputs.inputdata}} --output ${{outputs.outputdata}}",

environment="AzureML-sklearn-1.0-ubuntu20.04-py38-cpu@latest"

)

Register component

preprocesscomponent = mlclient.components.createorupdate(preprocesscomponent)

print(f"Component registered: {preprocesscomponent.name}")

2. Component Script

# components/preprocess/preprocess.py

import argparse

import pandas as pd

import os

def main():

parser = argparse.ArgumentParser()

parser.addargument("--input", type=str, required=True)

parser.addargument("--output", type=str, required=True)

args = parser.parseargs()

# Load data

df = pd.readcsv(args.input)

# Preprocess

df = df.dropna()

df = df.dropduplicates()

# Normalize numeric columns

numericcols = df.selectdtypes(include=['number']).columns

df[numericcols] = (df[numericcols] - df[numericcols].mean()) / df[numericcols].std()

# Save output

os.makedirs(args.output, existok=True)

df.tocsv(os.path.join(args.output, "processeddata.csv"), index=False)

print(f"Processed {len(df)} rows")

if name == "main":

main()

3. Training Component

# Training component

traincomponent = command(

name="trainmodel",

displayname="Train Model",

description="Train sklearn model",

inputs={

"trainingdata": Input(type="urifolder"),

"nestimators": Input(type="integer", default=100),

"maxdepth": Input(type="integer", default=10)

},

outputs={

"model": Output(type="urifolder")

},

code="./components/train",

command="""python train.py \

--data ${{inputs.trainingdata}} \

--nestimators ${{inputs.nestimators}} \

--maxdepth ${{inputs.maxdepth}} \

--output ${{outputs.model}}""",

environment="AzureML-sklearn-1.0-ubuntu20.04-py38-cpu@latest"

)

mlclient.components.createorupdate(traincomponent)

4. Component from YAML

# components/train/component.yaml

$schema: https://azuremlschemas.azureedge.net/latest/commandComponent.schema.json

name: trainmodel

displayname: Train Model

version: 1.0.0

type: command

inputs:

trainingdata:

type: urifolder

nestimators:

type: integer

default: 100

maxdepth:

type: integer

default: 10

outputs:

model:

type: urifolder

code: ./

command: >-

python train.py

--data ${{inputs.trainingdata}}

--nestimators ${{inputs.nestimators}}

--maxdepth ${{inputs.maxdepth}}

--output ${{outputs.model}}

environment: azureml:AzureML-sklearn-1.0-ubuntu20.04-py38-cpu@latest

# Load from YAML

from azure.ai.ml import loadcomponent

traincomponent = loadcomponent(source="./components/train/component.yaml")

mlclient.components.createorupdate(traincomponent)

Building Complex Pipelines

1. Multi-Step Pipeline

from azure.ai.ml import dsl, Input, Output

@dsl.pipeline(

compute="cpu-cluster",

description="Complete ML pipeline"

)

def completemlpipeline(rawdata, testsize=0.2):

# Step 1: Data preprocessing

preprocess = preprocesscomponent(

inputdata=rawdata

)

# Step 2: Feature engineering

features = featureengineeringcomponent(

inputdata=preprocess.outputs.outputdata

)

# Step 3: Train-test split

split = splitcomponent(

inputdata=features.outputs.outputdata,

testsize=testsize

)

# Step 4: Model training

train = traincomponent(

trainingdata=split.outputs.traindata,

nestimators=100,

maxdepth=10

)

# Step 5: Model evaluation

evaluate = evaluatecomponent(

model=train.outputs.model,

testdata=split.outputs.testdata

)

# Step 6: Register model if accuracy threshold met

register = registermodelcomponent(

model=train.outputs.model,

metrics=evaluate.outputs.metrics,

threshold=0.85

)

return {

"trainedmodel": train.outputs.model,

"evaluationmetrics": evaluate.outputs.metrics

}

2. Conditional Execution

from azure.ai.ml import dsl

from azure.ai.ml.dsl import condition

@dsl.pipeline()

def conditionalpipeline(data, accuracythreshold=0.85):

# Train model

train = traincomponent(trainingdata=data)

# Evaluate

evaluate = evaluatecomponent(

model=train.outputs.model,

testdata=data

)

# Conditional deployment

with condition(evaluate.outputs.accuracy >= accuracythreshold):

deploy = deploycomponent(

model=train.outputs.model

)

return {"model": train.outputs.model}

3. Parallel Execution

from azure.ai.ml import dsl

from azure.ai.ml.parallel import parallelrunfunction, RunFunction

Define parallel component

parallelstep = parallelrunfunction(

name="batchscoring",

displayname="Batch Scoring",

inputs={

"inputdata": Input(type="urifolder"),

"model": Input(type="urifolder")

},

outputs={

"predictions": Output(type="urifolder")

},

minibatchsize="10",

task=RunFunction(

code="./parallel",

entryscript="score.py",

environment="AzureML-sklearn-1.0-ubuntu20.04-py38-cpu@latest"

),

instancecount=4

)

@dsl.pipeline()

def parallelpipeline(data, model):

# Run parallel scoring

scores = parallelstep(

inputdata=data,

model=model

)

return {"predictions": scores.outputs.predictions}

Pipeline Parameters

1. Define Parameters

from azure.ai.ml import dsl, Input

@dsl.pipeline(

compute="cpu-cluster",

description="Parameterized pipeline"

)

def parameterizedpipeline(

trainingdata: Input,

learningrate: float = 0.01,

epochs: int = 100,

batchsize: int = 32

):

train = traincomponent(

data=trainingdata,

learningrate=learningrate,

epochs=epochs,

batchsize=batchsize

)

return {"model": train.outputs.model}

Run with different parameters

pipelinejob = mlclient.jobs.createorupdate(

parameterizedpipeline(

trainingdata=Input(type="urifile", path="azureml:dataset:1"),

learningrate=0.001,

epochs=200

)

)

2. Sweep for Hyperparameter Tuning

from azure.ai.ml import dsl

from azure.ai.ml.sweep import Choice, Uniform, BanditPolicy

@dsl.pipeline()

def sweeppipeline(data):

# Training with sweep

trainsweep = traincomponent(

data=data,

learningrate=Uniform(minvalue=0.001, maxvalue=0.1),

nestimators=Choice([50, 100, 200]),

maxdepth=Choice([5, 10, 15])

)

# Configure sweep

sweepjob = trainsweep.sweep(

primarymetric="accuracy",

goal="maximize",

samplingalgorithm="random",

compute="cpu-cluster"

)

sweepjob.setlimits(

maxtotaltrials=20,

maxconcurrenttrials=4,

timeout=7200

)

sweepjob.earlytermination = BanditPolicy(

slackfactor=0.1,

evaluationinterval=1

)

return {"bestmodel": sweepjob.outputs.model}

Pipeline Scheduling

1. Create Schedule

from azure.ai.ml.entities import (

JobSchedule,

RecurrenceTrigger,

RecurrencePattern

)

Create recurring schedule

schedule = JobSchedule(

name="daily-training-schedule",

trigger=RecurrenceTrigger(

frequency="day",

interval=1,

schedule=RecurrencePattern(hours=6, minutes=0),

timezone="UTC"

),

createjob=trainingpipeline(

trainingdata=Input(type="urifile", path="azureml:daily-data:latest")

)

)

mlclient.schedules.begincreateorupdate(schedule).result()

print("Schedule created")

2. Cron Schedule

from azure.ai.ml.entities import CronTrigger

cronschedule = JobSchedule(

name="weekly-retraining",

trigger=CronTrigger(

expression="0 0 0", # Every Sunday at midnight

timezone="America/NewYork"

),

createjob=retrainingpipeline(

data=Input(type="urifolder", path="azureml:weekly-data:latest")

)

)

mlclient.schedules.begincreateorupdate(cronschedule).result()

3. Manage Schedules

# List schedules

schedules = mlclient.schedules.list()

for s in schedules:

print(f"{s.name}: {s.trigger}")

Disable schedule

schedule = mlclient.schedules.get("daily-training-schedule")

schedule.isenabled = False

mlclient.schedules.begincreateorupdate(schedule).result()

Delete schedule

mlclient.schedules.begindelete("daily-training-schedule").result()

CI/CD Integration

1. Azure DevOps Pipeline

# azure-pipelines.yml

trigger:

branches:

include:

  • main
paths:

include:

  • src/
  • pipelines/

pool:

vmImage: 'ubuntu-latest'

variables:

  • group: azure-ml-vars

steps:

  • task: UsePythonVersion@0
inputs:

versionSpec: '3.9'

  • script: |
pip install azure-ai-ml azure-identity

displayName: 'Install dependencies'

  • task: AzureCLI@2
inputs:

azureSubscription: 'my-azure-subscription'

scriptType: 'bash'

scriptLocation: 'inlineScript'

inlineScript: |

python pipelines/submitpipeline.py

displayName: 'Submit ML Pipeline'

2. GitHub Actions

# .github/workflows/ml-pipeline.yml

name: ML Pipeline

on:

push:

branches: [main]

paths:

  • 'src/'
  • 'pipelines/'

jobs:

train:

runs-on: ubuntu-latest

steps:

  • uses: actions/checkout@v3
  • uses: actions/setup-python@v4
with:

python-version: '3.9'

  • name: Install dependencies
run: pip install azure-ai-ml azure-identity

  • name: Azure Login
uses: azure/login@v1

with:

creds: ${{ secrets.AZURECREDENTIALS }}

  • name: Submit Pipeline
run: python pipelines/submit
pipeline.py

env:

AZURESUBSCRIPTIONID: ${{ secrets.AZURESUBSCRIPTIONID }}

AZURERESOURCEGROUP: ${{ secrets.AZURERESOURCEGROUP }}

AZUREWORKSPACENAME: ${{ secrets.AZUREWORKSPACENAME }}

3. Pipeline Submission Script

# pipelines/submitpipeline.py

import os

from azure.ai.ml import MLClient, Input

from azure.identity import DefaultAzureCredential

def main():

# Connect to workspace

mlclient = MLClient(

credential=DefaultAzureCredential(),

subscriptionid=os.environ["AZURESUBSCRIPTIONID"],

resourcegroupname=os.environ["AZURERESOURCEGROUP"],

workspacename=os.environ["AZUREWORKSPACENAME"]

)

# Load pipeline

from mlpipeline import trainingpipeline

# Submit

pipelinejob = mlclient.jobs.createorupdate(

trainingpipeline(

trainingdata=Input(

type="urifile",

path="azureml:training-data:latest"

)

),

experimentname="ci-cd-training"

)

print(f"Pipeline submitted: {pipelinejob.name}")

print(f"Studio URL: {pipelinejob.studiourl}")

# Wait for completion

mlclient.jobs.stream(pipelinejob.name)

if name == "main":

main()

Monitoring and Debugging

1. Monitor Pipeline

# Get pipeline status

job = mlclient.jobs.get(pipelinejob.name)

print(f"Status: {job.status}")

Stream logs

mlclient.jobs.stream(pipelinejob.name)

Get child jobs (steps)

childjobs = mlclient.jobs.list(parentjobname=pipelinejob.name)

for child in childjobs:

print(f"{child.displayname}: {child.status}")

2. Download Outputs

# Download outputs

mlclient.jobs.download(

name=pipelinejob.name,

outputname="trainedmodel",

downloadpath="./outputs"

)

Best Practices

1. Modular Components

# Keep components focused and reusable

@dsl.pipeline()

def modularpipeline(data):

# Each component does one thing well

cleaned = cleandatacomponent(data=data)

features = extractfeaturescomponent(data=cleaned.outputs.data)

model = trainmodelcomponent(data=features.outputs.data)

return {"model": model.outputs.model}

2. Version Control

# Version your components

component = command(

name="trainmodel",

version="1.2.0", # Semantic versioning

# ...

)

Reference specific versions in pipelines

train = mlclient.components.get("train_model", version="1.2.0")

Conclusion

Azure ML Pipelines provide:

  • Reproducibility: Versioned workflows
  • Automation: Scheduled execution
  • Modularity: Reusable components
  • CI/CD: DevOps integration
  • Scalability: Distributed processing
  • Key takeaways:

    • Build modular, reusable components
    • Use parameters for flexibility
    • Integrate with CI/CD tools
    • Schedule pipelines for automation
    • Monitor and debug effectively

    Related Articles

    Azure DevOps for MLOps Tutorial: CI/CD for Machine Learning

    Tutorial Lengkap Azure DevOps untuk MLOps: CI/CD untuk Machine Learning Azure DevOps menyediakan kemampuan CI/CD kompreh...

    AWS SageMaker Pipelines Tutorial: ML Pipeline Automation

    Tutorial Lengkap AWS SageMaker Pipelines: Automasi ML Workflows SageMaker Pipelines adalah layanan CI/CD yang dibuat khu...

    Vertex AI Pipelines Tutorial: ML Pipeline Orchestration

    Tutorial Lengkap Vertex AI Pipelines: Orkestrasi Workflow ML Vertex AI Pipelines memungkinkan Anda mengorkestrasi workfl...

    Azure ML Managed Endpoints Tutorial: Production Model Deployment

    Tutorial Lengkap Azure ML Managed Endpoints: Deployment Model Production Azure ML Managed Endpoints menyediakan solusi f...