Complete GitHub Actions Tutorial for ML CI/CD
GitHub Actions enables you to automate Machine Learning workflows from testing to deployment. In this tutorial, we'll build a comprehensive CI/CD pipeline for ML projects, including data validation, model training, testing, and deployment.
Why CI/CD for ML?
Challenges in ML projects:
- Reproducibility: Ensuring consistent results across environments
- Testing: Validating data, models, and code
- Automation: Reducing manual work and human error
- Collaboration: Standardizing workflows across teams
- Monitoring: Tracking performance and detecting regressions
- Automated testing on every push/PR
- Scheduled retraining
- Model validation gates
- Automated deployment
- Integration with cloud services
GitHub Actions Basics
1. Workflow File Structure
# .github/workflows/ml-pipeline.yml
name: ML Pipeline
Triggers
on:
push:
branches: [main, develop]
pullrequest:
branches: [main]
schedule:
- cron: '0 0 0' # Weekly on Sunday
workflowdispatch: # Manual trigger
Environment variables
env:
PYTHONVERSION: '3.10'
MODELNAME: 'my-model'
Jobs
jobs:
test:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- name: Setup Python
uses: actions/setup-python@v5
with:
python-version: ${{ env.PYTHONVERSION }}
- name: Install dependencies
run: pip install -r requirements.txt
- name: Run tests
run: pytest tests/
2. Workflow Triggers
on:
# Push to specific branches
push:
branches: [main]
paths:
- 'src/'
- 'tests/'
- 'requirements.txt'
# Pull requests
pullrequest:
branches: [main]
# Scheduled runs
schedule:
- cron: '0 2 ' # Daily at 2 AM UTC
# Manual trigger with inputs
workflowdispatch:
inputs:
environment:
description: 'Deployment environment'
required: true
default: 'staging'
type: choice
options:
- staging
- production
retrain:
description: 'Force retrain model'
required: false
type: boolean
default: false
ML Testing Pipeline
1. Code Quality and Unit Tests
name: Code Quality & Tests
on:
push:
branches: [main, develop]
pullrequest:
branches: [main]
jobs:
lint:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- name: Setup Python
uses: actions/setup-python@v5
with:
python-version: '3.10'
- name: Install linters
run: |
pip install flake8 black isort mypy
- name: Run flake8
run: flake8 src/ tests/
- name: Check black formatting
run: black --check src/ tests/
- name: Check import sorting
run: isort --check-only src/ tests/
- name: Run mypy
run: mypy src/
test:
runs-on: ubuntu-latest
needs: lint
steps:
- uses: actions/checkout@v4
- name: Setup Python
uses: actions/setup-python@v5
with:
python-version: '3.10'
- name: Cache pip
uses: actions/cache@v4
with:
path: ~/.cache/pip
key: ${{ runner.os }}-pip-${{ hashFiles('requirements.txt') }}
- name: Install dependencies
run: pip install -r requirements.txt
- name: Run unit tests
run: pytest tests/unit/ -v --cov=src --cov-report=xml
- name: Upload coverage
uses: codecov/codecov-action@v4
with:
files: ./coverage.xml
2. Data Validation
name: Data Validation
on:
push:
paths:
- 'data/'
pullrequest:
paths:
- 'data/'
jobs:
validate-data:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- name: Setup Python
uses: actions/setup-python@v5
with:
python-version: '3.10'
- name: Install dependencies
run: |
pip install great
expectations pandas
- name: Run data validation
run: python scripts/validatedata.py
- name: Upload validation report
uses: actions/upload-artifact@v4
with:
name: data-validation-report
path: reports/data
validation.html
Data validation script:
# scripts/validatedata.py
import great
expectations as gx
import pandas as pd
import sys
def validatetrainingdata():
# Load data
df = pd.readcsv("data/train.csv")
# Create expectation suite
context = gx.getcontext()
# Define expectations
expectations = [
# No null values in critical columns
{"expectationtype": "expectcolumnvaluestonotbenull", "kwargs": {"column": "target"}},
{"expectationtype": "expectcolumnvaluestonotbenull", "kwargs": {"column": "feature1"}},
# Value ranges
{"expectationtype": "expectcolumnvaluestobebetween", "kwargs": {"column": "feature1", "minvalue": 0, "maxvalue": 100}},
# Unique values
{"expectationtype": "expectcolumnvaluestobeunique", "kwargs": {"column": "id"}},
# Row count
{"expectationtype": "expecttablerowcounttobebetween", "kwargs": {"minvalue": 1000, "maxvalue": 1000000}},
]
# Validate
validator = context.getvalidator(
batchrequest=context.getdatasource("pandas").getbatchrequest(df),
expectationsuitename="trainingdatasuite"
)
for exp in expectations:
getattr(validator, exp["expectationtype"])(exp["kwargs"])
# Get results
results = validator.validate()
if not results.success:
print("Data validation FAILED!")
print(results)
sys.exit(1)
print("Data validation PASSED!")
if name == "main":
validatetrainingdata()
3. Model Testing
name: Model Tests
on:
push:
paths:
- 'src/model/'
- 'tests/model/'
pullrequest:
paths:
- 'src/model/'
jobs:
model-tests:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- name: Setup Python
uses: actions/setup-python@v5
with:
python-version: '3.10'
- name: Install dependencies
run: pip install -r requirements.txt
- name: Download test data
run: |
# Download from DVC or storage
dvc pull data/test.csv.dvc
- name: Run model tests
run: |
pytest tests/model/ -v \
--junitxml=reports/modeltests.xml
- name: Check model performance
run: python scripts/checkmodelperformance.py
- name: Upload test results
uses: actions/upload-artifact@v4
if: always()
with:
name: model-test-results
path: reports/
Model performance check script:
# scripts/checkmodelperformance.py
import json
import sys
import joblib
from sklearn.metrics import accuracy
score, f1score
import pandas as pd
Thresholds
MIN
ACCURACY = 0.85
MINF1SCORE = 0.80
def checkperformance():
# Load model
model = joblib.load("models/model.pkl")
# Load test data
testdf = pd.readcsv("data/test.csv")
Xtest = testdf.drop("target", axis=1)
ytest = testdf["target"]
# Predict
ypred = model.predict(Xtest)
# Calculate metrics
accuracy = accuracyscore(ytest, ypred)
f1 = f1score(ytest, ypred, average="weighted")
print(f"Accuracy: {accuracy:.4f} (min: {MINACCURACY})")
print(f"F1 Score: {f1:.4f} (min: {MINF1SCORE})")
# Save metrics
metrics = {"accuracy": accuracy, "f1score": f1}
with open("reports/metrics.json", "w") as f:
json.dump(metrics, f)
# Check thresholds
if accuracy < MINACCURACY:
print(f"FAILED: Accuracy {accuracy:.4f} < {MINACCURACY}")
sys.exit(1)
if f1 < MINF1SCORE:
print(f"FAILED: F1 Score {f1:.4f} < {MINF1SCORE}")
sys.exit(1)
print("Model performance check PASSED!")
if name == "main":
checkperformance()
Training Pipeline
1. Automated Training
name: Model Training
on:
schedule:
- cron: '0 0 0' # Weekly
workflowdispatch:
inputs:
force
train:
description: 'Force training even if no data changes'
required: false
type: boolean
default: false
env:
AWSREGION: us-east-1
jobs:
check-data:
runs-on: ubuntu-latest
outputs:
datachanged: ${{ steps.check.outputs.changed }}
steps:
- uses: actions/checkout@v4
with:
fetch-depth: 2
- name: Check if data changed
id: check
run: |
if git diff --name-only HEAD~1 | grep -q "data/"; then
echo "changed=true" >> $GITHUBOUTPUT
else
echo "changed=false" >> $GITHUBOUTPUT
fi
train:
needs: check-data
if: needs.check-data.outputs.datachanged == 'true' || github.event.inputs.forcetrain == 'true'
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- name: Setup Python
uses: actions/setup-python@v5
with:
python-version: '3.10'
- name: Configure AWS credentials
uses: aws-actions/configure-aws-credentials@v4
with:
aws-access-key-id: ${{ secrets.AWSACCESSKEYID }}
aws-secret-access-key: ${{ secrets.AWSSECRETACCESSKEY }}
aws-region: ${{ env.AWSREGION }}
- name: Install dependencies
run: |
pip install -r requirements.txt
pip install dvc[s3]
- name: Pull data
run: dvc pull
- name: Train model
run: |
python src/train.py \
--data-path data/train.csv \
--output-dir models/
- name: Evaluate model
run: python src/evaluate.py
- name: Push model to DVC
run: |
dvc add models/model.pkl
dvc push
- name: Commit DVC files
run: |
git config user.name "GitHub Actions"
git config user.email "actions@github.com"
git add models/model.pkl.dvc
git commit -m "Update model [skip ci]" || echo "No changes to commit"
git push
- name: Upload training artifacts
uses: actions/upload-artifact@v4
with:
name: training-artifacts
path: |
models/
reports/
2. Training with GPU (Self-hosted Runner)
name: GPU Training
on:
workflowdispatch:
inputs:
epochs:
description: 'Number of epochs'
required: true
default: '10'
jobs:
train-gpu:
runs-on: [self-hosted, gpu] # Self-hosted runner with GPU
steps:
- uses: actions/checkout@v4
- name: Setup environment
run: |
conda activate ml-env
pip install -r requirements.txt
- name: Train with GPU
run: |
python src/train.py \
--epochs ${{ github.event.inputs.epochs }} \
--device cuda
- name: Upload model
uses: actions/upload-artifact@v4
with:
name: trained-model
path: models/
Deployment Pipeline
1. Model Deployment to AWS
name: Deploy Model
on:
push:
branches: [main]
paths:
- 'models/'
workflowdispatch:
inputs:
environment:
description: 'Deployment environment'
required: true
type: choice
options:
- staging
- production
jobs:
deploy:
runs-on: ubuntu-latest
environment: ${{ github.event.inputs.environment || 'staging' }}
steps:
- uses: actions/checkout@v4
- name: Configure AWS credentials
uses: aws-actions/configure-aws-credentials@v4
with:
aws-access-key-id: ${{ secrets.AWS
ACCESSKEYID }}
aws-secret-access-key: ${{ secrets.AWSSECRETACCESSKEY }}
aws-region: us-east-1
- name: Login to ECR
id: login-ecr
uses: aws-actions/amazon-ecr-login@v2
- name: Build and push Docker image
env:
ECRREGISTRY: ${{ steps.login-ecr.outputs.registry }}
IMAGETAG: ${{ github.sha }}
run: |
docker build -t $ECRREGISTRY/ml-model:$IMAGETAG .
docker push $ECRREGISTRY/ml-model:$IMAGETAG
- name: Deploy to ECS
run: |
aws ecs update-service \
--cluster ml-cluster \
--service ml-service \
--force-new-deployment
- name: Wait for deployment
run: |
aws ecs wait services-stable \
--cluster ml-cluster \
--services ml-service
- name: Run smoke tests
run: |
python scripts/smoketest.py \
--endpoint ${{ vars.APIENDPOINT }}
2. Deployment to Kubernetes
name: Deploy to Kubernetes
on:
push:
branches: [main]
jobs:
deploy:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- name: Setup kubectl
uses: azure/setup-kubectl@v4
- name: Configure kubeconfig
run: |
echo "${{ secrets.KUBECONFIG }}" | base64 -d > kubeconfig
export KUBECONFIG=kubeconfig
- name: Update deployment
run: |
kubectl set image deployment/ml-model \
ml-model=myregistry/ml-model:${{ github.sha }} \
-n ml-namespace
- name: Wait for rollout
run: |
kubectl rollout status deployment/ml-model \
-n ml-namespace \
--timeout=300s
- name: Verify deployment
run: |
kubectl get pods -n ml-namespace
kubectl logs -l app=ml-model -n ml-namespace --tail=50
Complete ML Pipeline
Full CI/CD Workflow
# .github/workflows/ml-complete.yml
name: Complete ML Pipeline
on:
push:
branches: [main]
pullrequest:
branches: [main]
schedule:
- cron: '0 0 * 0'
env:
PYTHONVERSION: '3.10'
jobs:
# Stage 1: Code Quality
lint:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: actions/setup-python@v5
with:
python-version: ${{ env.PYTHONVERSION }}
- run: pip install flake8 black
- run: flake8 src/ tests/
- run: black --check src/ tests/
# Stage 2: Unit Tests
unit-tests:
needs: lint
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: actions/setup-python@v5
with:
python-version: ${{ env.PYTHONVERSION }}
- run: pip install -r requirements.txt
- run: pytest tests/unit/ -v --cov=src
# Stage 3: Data Validation
data-validation:
needs: unit-tests
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: actions/setup-python@v5
with:
python-version: ${{ env.PYTHONVERSION }}
- run: pip install -r requirements.txt
- run: python scripts/validatedata.py
# Stage 4: Training
train:
needs: data-validation
runs-on: ubuntu-latest
outputs:
modelversion: ${{ steps.train.outputs.version }}
steps:
- uses: actions/checkout@v4
- uses: actions/setup-python@v5
with:
python-version: ${{ env.PYTHONVERSION }}
- run: pip install -r requirements.txt
- name: Train model
id: train
run: |
python src/train.py
VERSION=$(date +%Y%m%d%H%M%S)
echo "version=$VERSION" >> $GITHUBOUTPUT
- uses: actions/upload-artifact@v4
with:
name: model-${{ steps.train.outputs.version }}
path: models/
# Stage 5: Model Validation
model-validation:
needs: train
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: actions/setup-python@v5
with:
python-version: ${{ env.PYTHONVERSION }}
- uses: actions/download-artifact@v4
with:
name: model-${{ needs.train.outputs.modelversion }}
path: models/
- run: pip install -r requirements.txt
- run: python scripts/checkmodelperformance.py
# Stage 6: Integration Tests
integration-tests:
needs: model-validation
runs-on: ubuntu-latest
services:
postgres:
image: postgres:15
env:
POSTGRESPASSWORD: test
ports:
- 5432:5432
steps:
- uses: actions/checkout@v4
- uses: actions/setup-python@v5
with:
python-version: ${{ env.PYTHONVERSION }}
- uses: actions/download-artifact@v4
with:
name: model-${{ needs.train.outputs.modelversion }}
path: models/
- run: pip install -r requirements.txt
- run: pytest tests/integration/ -v
# Stage 7: Deploy (only on main)
deploy:
if: github.ref == 'refs/heads/main' && github.eventname == 'push'
needs: [integration-tests]
runs-on: ubuntu-latest
environment: production
steps:
- uses: actions/checkout@v4
- uses: actions/download-artifact@v4
with:
name: model-${{ needs.train.outputs.modelversion }}
path: models/
- name: Deploy to production
run: |
echo "Deploying model version ${{ needs.train.outputs.modelversion }}"
# Add deployment commands here
Secrets and Environment Variables
1. Setup Secrets
# In workflow, access secrets like this:
env:
APIKEY: ${{ secrets.APIKEY }}
AWSACCESSKEYID: ${{ secrets.AWSACCESSKEYID }}
Or in steps:
- name: Use secret
run: echo "Using secret"
env:
MYSECRET: ${{ secrets.MYSECRET }}
2. Environment-specific Variables
jobs:
deploy:
environment: production # Uses 'production' environment
steps:
- name: Deploy
run: |
echo "Deploying to ${{ vars.APIENDPOINT }}"
Best Practices
1. Caching Dependencies
- name: Cache pip
uses: actions/cache@v4
with:
path: ~/.cache/pip
key: ${{ runner.os }}-pip-${{ hashFiles('requirements.txt') }}
restore-keys: |
${{ runner.os }}-pip-
- name: Cache models
uses: actions/cache@v4
with:
path: ~/.cache/huggingface
key: ${{ runner.os }}-hf-${{ hashFiles('modelconfig.json') }}
2. Matrix Testing
jobs:
test:
strategy:
matrix:
python-version: ['3.9', '3.10', '3.11']
os: [ubuntu-latest, macos-latest]
runs-on: ${{ matrix.os }}
steps:
- uses: actions/setup-python@v5
with:
python-version: ${{ matrix.python-version }}
3. Conditional Execution
jobs:
deploy:
if: github.ref == 'refs/heads/main'
...
train:
if: |
contains(github.event.headcommit.message, '[train]') ||
github.event_name == 'schedule'
...
Conclusion
GitHub Actions provides a powerful platform for ML CI/CD:
Key takeaways:
- Use caching to speed up workflows
- Implement proper testing gates
- Use environments for different stages
- Automate retraining with schedules
- Monitor model performance in CI/CD