XGBoost & LightGBM Tutorial: Gradient Boosting Masterclass

# XGBoost & LightGBM - Masterclass Gradient Boosting ## Daftar Isi 1. [Pendahuluan](#pendahuluan) 2. [Prasyarat](#prasyarat) 3. [Memahami Gradient Boosting](#memahami-gradient-boosting) 4. [Pelatiha...

By Ruby Abdullah · · tutorial
XGBoostLightGBMGradient BoostingSHAPMachine LearningPython

XGBoost & LightGBM - Gradient Boosting Masterclass

Table of Contents

  • Introduction
  • Prerequisites
  • Understanding Gradient Boosting
  • XGBoost Training and Tuning
  • LightGBM Training and Tuning
  • XGBoost vs LightGBM Comparison
  • Hyperparameter Optimization
  • SHAP Explainability
  • Feature Importance Analysis
  • Early Stopping and Cross-Validation
  • Model Deployment
  • Best Practices
  • Conclusion

  • Introduction

    Gradient boosting is one of the most powerful machine learning techniques for structured/tabular data. XGBoost (eXtreme Gradient Boosting) and LightGBM (Light Gradient Boosting Machine) are two leading implementations that consistently win Kaggle competitions and power production ML systems across industries. This masterclass covers both libraries from fundamentals to production deployment.

    XGBoost, developed by Tianqi Chen, introduced regularization to gradient boosting and popularized the technique. LightGBM, developed by Microsoft, brought innovations like histogram-based splitting and leaf-wise tree growth for faster training. Understanding both allows you to choose the right tool for each problem.

    Prerequisites

    • Python 3.8 or higher
    • Basic understanding of machine learning concepts (classification, regression, overfitting)
    • Familiarity with scikit-learn API conventions

    pip install xgboost lightgbm
    

    pip install scikit-learn pandas numpy

    pip install shap matplotlib seaborn

    pip install optuna # For hyperparameter optimization

    Understanding Gradient Boosting

    Gradient boosting builds an ensemble of weak learners (typically decision trees) sequentially, where each new tree corrects the errors of the previous ensemble:

    import numpy as np
    

    import matplotlib.pyplot as plt

    Conceptual illustration of gradient boosting

    Step 1: Start with an initial prediction (e.g., mean of target)

    Step 2: Compute residuals (actual - predicted)

    Step 3: Fit a tree to the residuals

    Step 4: Update predictions: newpred = oldpred + learningrate treepred

    Step 5: Repeat steps 2-4

    Key differences between XGBoost and LightGBM:

    XGBoost: level-wise tree growth (breadth-first)

    LightGBM: leaf-wise tree growth (best-first) - faster but risk overfitting

    XGBoost splits nodes level by level

    LightGBM splits the leaf with the highest loss reduction

    XGBoost Training and Tuning

    Basic Classification

    import xgboost as xgb
    

    import numpy as np

    from sklearn.datasets import makeclassification

    from sklearn.modelselection import traintestsplit

    from sklearn.metrics import accuracyscore, classificationreport, rocaucscore

    Generate sample data

    X, y = makeclassification(

    nsamples=10000, nfeatures=20, ninformative=15,

    nredundant=3, randomstate=42

    )

    Xtrain, Xtest, ytrain, ytest = traintestsplit(X, y, testsize=0.2, randomstate=42)

    Using scikit-learn API

    clf = xgb.XGBClassifier(

    nestimators=500,

    maxdepth=6,

    learningrate=0.1,

    subsample=0.8,

    colsamplebytree=0.8,

    minchildweight=5,

    gamma=0.1,

    regalpha=0.1, # L1 regularization

    reglambda=1.0, # L2 regularization

    objective="binary:logistic",

    evalmetric="logloss",

    treemethod="hist", # Histogram-based method (faster)

    device="cuda", # Use GPU if available

    randomstate=42,

    njobs=-1

    )

    clf.fit(

    Xtrain, ytrain,

    evalset=[(Xtest, ytest)],

    verbose=50

    )

    ypred = clf.predict(Xtest)

    yproba = clf.predictproba(Xtest)[:, 1]

    print(f"Accuracy: {accuracyscore(ytest, ypred):.4f}")

    print(f"AUC-ROC: {rocaucscore(ytest, yproba):.4f}")

    print(classificationreport(ytest, ypred))

    XGBoost Regression

    from sklearn.datasets import makeregression
    

    from sklearn.metrics import meansquarederror, r2score

    X, y = makeregression(nsamples=10000, nfeatures=20, noise=10, randomstate=42)

    Xtrain, Xtest, ytrain, ytest = traintestsplit(X, y, testsize=0.2, randomstate=42)

    reg = xgb.XGBRegressor(

    nestimators=500,

    maxdepth=6,

    learningrate=0.05,

    subsample=0.8,

    colsamplebytree=0.8,

    minchildweight=5,

    regalpha=0.1,

    reglambda=1.0,

    objective="reg:squarederror",

    treemethod="hist",

    randomstate=42

    )

    reg.fit(

    Xtrain, ytrain,

    evalset=[(Xtest, ytest)],

    verbose=100

    )

    ypred = reg.predict(Xtest)

    print(f"RMSE: {np.sqrt(meansquarederror(ytest, ypred)):.4f}")

    print(f"R2: {r2score(ytest, ypred):.4f}")

    XGBoost Native API (DMatrix)

    # The native API offers more control and is often faster
    

    dtrain = xgb.DMatrix(Xtrain, label=ytrain, featurenames=[f"f{i}" for i in range(20)])

    dtest = xgb.DMatrix(Xtest, label=ytest, featurenames=[f"f{i}" for i in range(20)])

    params = {

    "maxdepth": 6,

    "eta": 0.1,

    "objective": "binary:logistic",

    "evalmetric": ["logloss", "auc"],

    "subsample": 0.8,

    "colsamplebytree": 0.8,

    "minchildweight": 5,

    "gamma": 0.1,

    "alpha": 0.1,

    "lambda": 1.0,

    "treemethod": "hist",

    "seed": 42

    }

    evals = [(dtrain, "train"), (dtest, "eval")]

    model = xgb.train(

    params, dtrain,

    numboostround=500,

    evals=evals,

    earlystoppingrounds=50,

    verboseeval=50

    )

    print(f"Best iteration: {model.bestiteration}")

    print(f"Best score: {model.bestscore:.4f}")

    LightGBM Training and Tuning

    Basic Classification

    import lightgbm as lgb
    

    from sklearn.datasets import makeclassification

    from sklearn.modelselection import traintestsplit

    from sklearn.metrics import accuracyscore, rocaucscore, classificationreport

    X, y = makeclassification(

    nsamples=10000, nfeatures=20, ninformative=15,

    nredundant=3, randomstate=42

    )

    Xtrain, Xtest, ytrain, ytest = traintestsplit(X, y, testsize=0.2, randomstate=42)

    Using scikit-learn API

    clf = lgb.LGBMClassifier(

    nestimators=500,

    maxdepth=-1, # No limit (controlled by numleaves)

    numleaves=31, # Key parameter for LightGBM

    learningrate=0.1,

    subsample=0.8,

    colsamplebytree=0.8,

    minchildsamples=20,

    minchildweight=1e-3,

    regalpha=0.1,

    reglambda=1.0,

    objective="binary",

    metric="binarylogloss",

    boostingtype="gbdt", # Options: gbdt, dart, rf

    njobs=-1,

    randomstate=42,

    verbose=-1

    )

    clf.fit(

    Xtrain, ytrain,

    evalset=[(Xtest, ytest)],

    callbacks=[

    lgb.earlystopping(50),

    lgb.logevaluation(50)

    ]

    )

    ypred = clf.predict(Xtest)

    yproba = clf.predictproba(Xtest)[:, 1]

    print(f"Accuracy: {accuracyscore(ytest, ypred):.4f}")

    print(f"AUC-ROC: {rocaucscore(ytest, yproba):.4f}")

    print(classificationreport(ytest, ypred))

    LightGBM Native API

    # Native API for more control
    

    traindata = lgb.Dataset(Xtrain, label=ytrain)

    testdata = lgb.Dataset(Xtest, label=ytest, reference=traindata)

    params = {

    "objective": "binary",

    "metric": ["binarylogloss", "auc"],

    "boostingtype": "gbdt",

    "numleaves": 31,

    "learningrate": 0.1,

    "featurefraction": 0.8,

    "baggingfraction": 0.8,

    "baggingfreq": 5,

    "minchildsamples": 20,

    "lambdal1": 0.1,

    "lambdal2": 1.0,

    "verbose": -1,

    "seed": 42,

    "numthreads": -1

    }

    callbacks = [

    lgb.earlystopping(50),

    lgb.logevaluation(50)

    ]

    model = lgb.train(

    params, traindata,

    numboostround=500,

    validsets=[traindata, testdata],

    validnames=["train", "eval"],

    callbacks=callbacks

    )

    print(f"Best iteration: {model.bestiteration}")

    print(f"Best score: {model.bestscore}")

    LightGBM with Categorical Features

    import pandas as pd
    

    import lightgbm as lgb

    LightGBM handles categorical features natively

    df = pd.DataFrame({

    "color": pd.Categorical(["red", "blue", "green", "red", "blue"] 2000),

    "size": pd.Categorical(["S", "M", "L", "XL", "S"] 2000),

    "weight": np.random.randn(10000),

    "price": np.random.uniform(10, 100, 10000),

    "target": np.random.randint(0, 2, 10000)

    })

    X = df.drop("target", axis=1)

    y = df["target"]

    Xtrain, Xtest, ytrain, ytest = traintestsplit(X, y, testsize=0.2, randomstate=42)

    traindata = lgb.Dataset(Xtrain, label=ytrain, categoricalfeature=["color", "size"])

    testdata = lgb.Dataset(Xtest, label=ytest, reference=traindata)

    params = {

    "objective": "binary",

    "metric": "binarylogloss",

    "numleaves": 31,

    "learningrate": 0.1,

    "verbose": -1

    }

    model = lgb.train(

    params, traindata,

    numboostround=200,

    validsets=[testdata],

    callbacks=[lgb.earlystopping(20), lgb.logevaluation(50)]

    )

    XGBoost vs LightGBM Comparison

    import time
    

    import numpy as np

    from sklearn.datasets import makeclassification

    from sklearn.modelselection import traintestsplit

    from sklearn.metrics import rocaucscore

    import xgboost as xgb

    import lightgbm as lgb

    def comparemodels(nsamples=100000, nfeatures=50):

    X, y = makeclassification(

    nsamples=nsamples, nfeatures=nfeatures,

    ninformative=30, randomstate=42

    )

    Xtrain, Xtest, ytrain, ytest = traintestsplit(X, y, testsize=0.2, randomstate=42)

    # XGBoost

    start = time.time()

    xgbmodel = xgb.XGBClassifier(

    nestimators=500, maxdepth=6, learningrate=0.1,

    treemethod="hist", randomstate=42, verbosity=0,

    earlystoppingrounds=50, evalmetric="logloss"

    )

    xgbmodel.fit(Xtrain, ytrain, evalset=[(Xtest, ytest)], verbose=False)

    xgbtime = time.time() - start

    xgbauc = rocaucscore(ytest, xgbmodel.predictproba(Xtest)[:, 1])

    # LightGBM

    start = time.time()

    lgbmodel = lgb.LGBMClassifier(

    nestimators=500, numleaves=31, learningrate=0.1,

    randomstate=42, verbose=-1

    )

    lgbmodel.fit(

    Xtrain, ytrain, evalset=[(Xtest, ytest)],

    callbacks=[lgb.earlystopping(50), lgb.logevaluation(0)]

    )

    lgbtime = time.time() - start

    lgbauc = rocaucscore(ytest, lgbmodel.predictproba(Xtest)[:, 1])

    print(f"{'Metric':<20} {'XGBoost':<15} {'LightGBM':<15}")

    print("-" 50)

    print(f"{'Training Time':<20} {xgbtime:.2f}s{'':<10} {lgbtime:.2f}s")

    print(f"{'AUC-ROC':<20} {xgbauc:.4f}{'':<10} {lgbauc:.4f}")

    print(f"{'Num Trees Used':<20} {xgbmodel.bestiteration}{'':<10} {lgbmodel.bestiteration}")

    print(f"{'Speed Ratio':<20} 1.0x{'':<11} {xgbtime/lgbtime:.1f}x faster")

    comparemodels(100000, 50)

    Key differences summarized:

    | Aspect | XGBoost | LightGBM |

    |--------|---------|----------|

    | Tree Growth | Level-wise | Leaf-wise |

    | Speed | Fast | Faster (typically 2-5x) |

    | Memory | Higher | Lower (histogram binning) |

    | Categorical | Requires encoding | Native support |

    | Overfitting Risk | Lower (level-wise) | Higher (leaf-wise, mitigate with numleaves) |

    | GPU Support | Yes | Yes |

    | Missing Values | Built-in handling | Built-in handling |

    Hyperparameter Optimization

    Using Optuna for systematic hyperparameter search:

    import optuna
    

    from sklearn.modelselection import crossvalscore

    import xgboost as xgb

    import lightgbm as lgb

    def xgboostobjective(trial):

    params = {

    "nestimators": trial.suggestint("nestimators", 100, 1000),

    "maxdepth": trial.suggestint("maxdepth", 3, 10),

    "learningrate": trial.suggestfloat("learningrate", 0.01, 0.3, log=True),

    "subsample": trial.suggestfloat("subsample", 0.5, 1.0),

    "colsamplebytree": trial.suggestfloat("colsamplebytree", 0.5, 1.0),

    "minchildweight": trial.suggestint("minchildweight", 1, 20),

    "gamma": trial.suggestfloat("gamma", 0.0, 5.0),

    "regalpha": trial.suggestfloat("regalpha", 1e-8, 10.0, log=True),

    "reglambda": trial.suggestfloat("reglambda", 1e-8, 10.0, log=True),

    "treemethod": "hist",

    "randomstate": 42,

    "verbosity": 0

    }

    model = xgb.XGBClassifier(params)

    scores = crossvalscore(model, Xtrain, ytrain, cv=5, scoring="rocauc", njobs=-1)

    return scores.mean()

    def lightgbmobjective(trial):

    params = {

    "nestimators": trial.suggestint("nestimators", 100, 1000),

    "numleaves": trial.suggestint("numleaves", 15, 127),

    "maxdepth": trial.suggestint("maxdepth", 3, 12),

    "learningrate": trial.suggestfloat("learningrate", 0.01, 0.3, log=True),

    "subsample": trial.suggestfloat("subsample", 0.5, 1.0),

    "colsamplebytree": trial.suggestfloat("colsamplebytree", 0.5, 1.0),

    "minchildsamples": trial.suggestint("minchildsamples", 5, 100),

    "regalpha": trial.suggestfloat("regalpha", 1e-8, 10.0, log=True),

    "reglambda": trial.suggestfloat("reglambda", 1e-8, 10.0, log=True),

    "randomstate": 42,

    "verbose": -1

    }

    model = lgb.LGBMClassifier(params)

    scores = crossvalscore(model, Xtrain, ytrain, cv=5, scoring="rocauc", njobs=-1)

    return scores.mean()

    Run optimization

    xgbstudy = optuna.createstudy(direction="maximize", studyname="xgboost")

    xgbstudy.optimize(xgboostobjective, ntrials=100, showprogressbar=True)

    lgbstudy = optuna.createstudy(direction="maximize", studyname="lightgbm")

    lgbstudy.optimize(lightgbmobjective, ntrials=100, showprogressbar=True)

    print(f"XGBoost best AUC: {xgbstudy.bestvalue:.4f}")

    print(f"XGBoost best params: {xgbstudy.bestparams}")

    print(f"LightGBM best AUC: {lgbstudy.bestvalue:.4f}")

    print(f"LightGBM best params: {lgbstudy.bestparams}")

    SHAP Explainability

    SHAP (SHapley Additive exPlanations) provides interpretable explanations for model predictions:

    import shap
    

    import xgboost as xgb

    import matplotlib.pyplot as plt

    Train a model

    model = xgb.XGBClassifier(nestimators=200, maxdepth=6, learningrate=0.1, randomstate=42)

    model.fit(Xtrain, ytrain)

    Create SHAP explainer

    explainer = shap.TreeExplainer(model)

    shapvalues = explainer.shapvalues(Xtest)

    Summary plot - global feature importance with direction

    shap.summaryplot(shapvalues, Xtest, featurenames=[f"feature{i}" for i in range(Xtest.shape[1])])

    plt.tightlayout()

    plt.savefig("shapsummary.png", dpi=150, bboxinches="tight")

    plt.close()

    Bar plot - mean absolute SHAP values

    shap.summaryplot(

    shapvalues, Xtest,

    featurenames=[f"feature{i}" for i in range(Xtest.shape[1])],

    plottype="bar"

    )

    plt.tightlayout()

    plt.savefig("shapbar.png", dpi=150, bboxinches="tight")

    plt.close()

    Waterfall plot for a single prediction

    shap.waterfallplot(shap.Explanation(

    values=shapvalues[0],

    basevalues=explainer.expectedvalue,

    data=Xtest[0],

    featurenames=[f"feature{i}" for i in range(Xtest.shape[1])]

    ))

    plt.tightlayout()

    plt.savefig("shapwaterfall.png", dpi=150, bboxinches="tight")

    plt.close()

    Dependence plot - interaction effects

    shap.dependenceplot(

    "feature0", shapvalues, Xtest,

    featurenames=[f"feature{i}" for i in range(Xtest.shape[1])],

    interactionindex="feature1"

    )

    plt.tightlayout()

    plt.savefig("shapdependence.png", dpi=150, bboxinches="tight")

    plt.close()

    Force plot for individual predictions

    shap.forceplot(

    explainer.expectedvalue,

    shapvalues[0:5],

    Xtest[0:5],

    featurenames=[f"feature{i}" for i in range(Xtest.shape[1])],

    matplotlib=True

    )

    plt.tightlayout()

    plt.savefig("shapforce.png", dpi=150, bboxinches="tight")

    plt.close()

    Feature Importance Analysis

    import xgboost as xgb
    

    import lightgbm as lgb

    import pandas as pd

    import matplotlib.pyplot as plt

    featurenames = [f"feature{i}" for i in range(Xtrain.shape[1])]

    XGBoost feature importance (multiple methods)

    xgbmodel = xgb.XGBClassifier(nestimators=200, randomstate=42, verbosity=0)

    xgbmodel.fit(Xtrain, ytrain)

    Built-in importance types

    importancetypes = ["weight", "gain", "cover", "totalgain", "totalcover"]

    importancedf = pd.DataFrame({"feature": featurenames})

    for imptype in importancetypes:

    scores = xgbmodel.getbooster().getscore(importancetype=imptype)

    importancedf[imptype] = [scores.get(f, 0) for f in xgbmodel.getbooster().featurenames]

    print("XGBoost Feature Importance (by gain):")

    print(importancedf.sortvalues("gain", ascending=False).head(10))

    LightGBM feature importance

    lgbmodel = lgb.LGBMClassifier(nestimators=200, randomstate=42, verbose=-1)

    lgbmodel.fit(Xtrain, ytrain)

    lgbimportance = pd.DataFrame({

    "feature": featurenames,

    "splitimportance": lgbmodel.featureimportances,

    "gainimportance": lgbmodel.booster.featureimportance(importancetype="gain")

    }).sortvalues("gainimportance", ascending=False)

    print("\nLightGBM Feature Importance (by gain):")

    print(lgbimportance.head(10))

    Permutation importance (model-agnostic)

    from sklearn.inspection import permutationimportance

    permimp = permutationimportance(

    xgbmodel, Xtest, ytest,

    nrepeats=10, randomstate=42, njobs=-1, scoring="rocauc"

    )

    permdf = pd.DataFrame({

    "feature": featurenames,

    "importancemean": permimp.importancesmean,

    "importancestd": permimp.importancesstd

    }).sortvalues("importancemean", ascending=False)

    print("\nPermutation Importance:")

    print(permdf.head(10))

    Early Stopping and Cross-Validation

    import xgboost as xgb
    

    import lightgbm as lgb

    from sklearn.modelselection import StratifiedKFold

    import numpy as np

    XGBoost built-in cross-validation

    dtrain = xgb.DMatrix(Xtrain, label=ytrain)

    params = {

    "maxdepth": 6,

    "eta": 0.1,

    "objective": "binary:logistic",

    "evalmetric": "auc",

    "treemethod": "hist",

    "seed": 42

    }

    cvresults = xgb.cv(

    params, dtrain,

    numboostround=1000,

    nfold=5,

    stratified=True,

    earlystoppingrounds=50,

    verboseeval=100,

    seed=42

    )

    print(f"Best numboostround: {len(cvresults)}")

    print(f"Best AUC: {cvresults['test-auc-mean'].iloc[-1]:.4f} +/- {cvresults['test-auc-std'].iloc[-1]:.4f}")

    LightGBM built-in cross-validation

    traindata = lgb.Dataset(Xtrain, label=ytrain)

    paramslgb = {

    "objective": "binary",

    "metric": "auc",

    "numleaves": 31,

    "learningrate": 0.1,

    "verbose": -1,

    "seed": 42

    }

    cvresultslgb = lgb.cv(

    paramslgb, traindata,

    numboostround=1000,

    nfold=5,

    stratified=True,

    callbacks=[lgb.earlystopping(50), lgb.logevaluation(100)],

    seed=42,

    returncvbooster=True

    )

    print(f"Best iteration: {cvresultslgb['cvbooster'].bestiteration}")

    print(f"Best AUC: {max(cvresultslgb['valid auc-mean']):.4f}")

    Manual cross-validation with custom logic

    skf = StratifiedKFold(nsplits=5, shuffle=True, randomstate=42)

    foldscores = []

    foldmodels = []

    for fold, (trainidx, validx) in enumerate(skf.split(Xtrain, ytrain)):

    Xfoldtrain, Xfoldval = Xtrain[trainidx], Xtrain[validx]

    yfoldtrain, yfoldval = ytrain[trainidx], ytrain[validx]

    model = xgb.XGBClassifier(

    nestimators=1000, maxdepth=6, learningrate=0.1,

    earlystoppingrounds=50, evalmetric="auc",

    treemethod="hist", randomstate=42, verbosity=0

    )

    model.fit(

    Xfoldtrain, yfoldtrain,

    evalset=[(Xfoldval, yfoldval)],

    verbose=False

    )

    valpred = model.predictproba(Xfoldval)[:, 1]

    auc = rocaucscore(yfoldval, valpred)

    foldscores.append(auc)

    foldmodels.append(model)

    print(f"Fold {fold + 1}: AUC = {auc:.4f} (best iter: {model.bestiteration})")

    print(f"\nMean AUC: {np.mean(foldscores):.4f} +/- {np.std(foldscores):.4f}")

    Model Deployment

    Saving and Loading Models

    import xgboost as xgb
    

    import lightgbm as lgb

    import joblib

    import json

    XGBoost - multiple save formats

    xgbmodel.savemodel("xgbmodel.json") # JSON format (recommended)

    xgbmodel.savemodel("xgbmodel.ubj") # Universal Binary JSON

    xgbmodel.getbooster().savemodel("xgbmodel.bin") # Binary format

    Load XGBoost model

    loadedxgb = xgb.XGBClassifier()

    loadedxgb.loadmodel("xgbmodel.json")

    LightGBM

    lgbmodel.booster.savemodel("lgbmodel.txt") # Text format

    joblib.dump(lgbmodel, "lgbmodel.joblib") # Joblib (preserves sklearn wrapper)

    Load LightGBM model

    loadedlgb = lgb.Booster(modelfile="lgbmodel.txt")

    loadedlgbsklearn = joblib.load("lgbmodel.joblib")

    FastAPI Deployment

    from fastapi import FastAPI, HTTPException
    

    from pydantic import BaseModel

    import xgboost as xgb

    import numpy as np

    from typing import List

    app = FastAPI(title="Gradient Boosting Prediction API")

    Load models at startup

    xgbmodel = xgb.XGBClassifier()

    xgbmodel.loadmodel("xgbmodel.json")

    class PredictionRequest(BaseModel):

    features: List[List[float]]

    class PredictionResponse(BaseModel):

    predictions: List[int]

    probabilities: List[List[float]]

    @app.post("/predict", responsemodel=PredictionResponse)

    async def predict(request: PredictionRequest):

    try:

    X = np.array(request.features)

    predictions = xgbmodel.predict(X).tolist()

    probabilities = xgbmodel.predictproba(X).tolist()

    return PredictionResponse(

    predictions=predictions,

    probabilities=probabilities

    )

    except Exception as e:

    raise HTTPException(statuscode=400, detail=str(e))

    @app.get("/model/info")

    async def modelinfo():

    booster = xgbmodel.getbooster()

    return {

    "numtrees": booster.numboostedrounds(),

    "numfeatures": booster.numfeatures(),

    "featurenames": booster.featurenames

    }

    Run: uvicorn deployment:app --host 0.0.0.0 --port 8000

    ONNX Export for Production

    from skl2onnx import convertsklearn
    

    from skl2onnx.common.datatypes import FloatTensorType

    import onnxruntime as ort

    import numpy as np

    Convert XGBoost model to ONNX

    initialtype = [("floatinput", FloatTensorType([None, Xtrain.shape[1]]))]

    onnxmodel = convertsklearn(xgbmodel, initialtypes=initialtype)

    with open("model.onnx", "wb") as f:

    f.write(onnxmodel.SerializeToString())

    Inference with ONNX Runtime

    session = ort.InferenceSession("model.onnx")

    inputname = session.getinputs()[0].name

    Predict

    Xsample = Xtest[:5].astype(np.float32)

    results = session.run(None, {inputname: Xsample})

    print(f"Predictions: {results[0]}")

    print(f"Probabilities: {results[1]}")

    Best Practices

  • Start Simple: Begin with default parameters and a small learning rate (0.05-0.1). Only tune after establishing a baseline.
  • Use Early Stopping: Always use early stopping to prevent overfitting and find the optimal number of trees automatically.
  • Feature Engineering First: Spend more time on feature engineering than hyperparameter tuning. Good features matter more than perfect parameters.
  • Handle Class Imbalance: Use scaleposweight (XGBoost) or isunbalance/scaleposweight (LightGBM) for imbalanced datasets.
  • Regularization: Increase regalpha, reglambda, and minchildweight to combat overfitting. Reduce maxdepth and numleaves.
  • Use Histogram-Based Methods: Always use treemethod="hist" in XGBoost for faster training without sacrificing accuracy.
  • Cross-Validate: Never rely on a single train/test split. Use stratified k-fold cross-validation for reliable estimates.
  • SHAP over Built-in Importance: Built-in feature importance can be misleading. Use SHAP values for reliable interpretability.
  • Monitor Overfitting: Watch the gap between training and validation metrics. A growing gap indicates overfitting.
  • Save Model Metadata: Always save preprocessing steps, feature names, and training parameters alongside the model.
  • Conclusion

    XGBoost and LightGBM remain the gold standard for structured data machine learning. XGBoost offers robust, well-tested performance with strong regularization options, while LightGBM provides faster training and native categorical feature support. For most practical applications, either library will deliver excellent results. The key to success lies in proper feature engineering, systematic hyperparameter optimization with tools like Optuna, rigorous cross-validation, and SHAP-based interpretability for building trust in model predictions. Choose XGBoost when you need maximum robustness, and LightGBM when training speed and memory efficiency are priorities.

    Related Articles

    SHAP Tutorial: Explainable AI and Model Interpretability

    SHAP - Panduan Praktis Explainable AI dan Interpretabilitas Model Model machine learning makin sering dipakai untuk meng...

    Complete Comet ML Tutorial: MLOps Platform for Experiment Tracking and Model Management

    Tutorial Lengkap Comet ML: Platform MLOps untuk Experiment Tracking dan Model Management Dalam dunia machine learning mo...

    MLX Tutorial: Apple's Machine Learning Framework for Apple Silicon

    Tutorial MLX: Framework Machine Learning Apple untuk Apple Silicon MLX adalah framework machine learning open-source dar...

    PyOD Tutorial: Anomaly and Outlier Detection in Python

    Deteksi Anomali di Python dengan PyOD: Panduan Praktis Sebagian besar dataset di dunia nyata mengandung sebagian kecil d...