Feature Engineering Masterclass Tutorial: Feature Techniques for ML

# Tutorial 14: Masterclass Rekayasa Fitur (Feature Engineering) ## Daftar Isi 1. [Pendahuluan](#pendahuluan) 2. [Prasyarat](#prasyarat) 3. [Mengapa Rekayasa Fitur Penting](#mengapa-rekayasa-fitur-pe...

By Ruby Abdullah · · tutorial
Feature EngineeringMachine LearningData ScienceFeaturetoolsPythonPreprocessing

Tutorial 14: Feature Engineering Masterclass

Table of Contents

  • Introduction
  • Prerequisites
  • Why Feature Engineering Matters
  • Numerical Feature Transformations
  • Categorical Feature Encoding
  • Datetime Feature Engineering
  • Text Feature Engineering
  • Feature Selection Techniques
  • Automated Feature Engineering with Featuretools
  • Building a Complete Feature Pipeline
  • Best Practices
  • Conclusion

  • Introduction

    Feature engineering is widely regarded as the most impactful activity in applied machine learning. While model architectures and hyperparameter tuning receive significant attention, the quality and expressiveness of your input features are what ultimately determine model performance. A well-engineered feature can make a simple linear model outperform a poorly-featured deep neural network.

    This masterclass covers the full spectrum of feature engineering techniques: from fundamental numerical transformations and categorical encoding strategies to advanced datetime extraction, text vectorization, and automated feature generation. Each technique is demonstrated with production-ready Python code and practical guidance on when and how to apply it.


    Prerequisites

    • Python 3.9 or higher
    • Solid understanding of pandas and NumPy
    • Basic machine learning knowledge (classification, regression)
    • Install required packages:

    pip install pandas numpy scikit-learn categoryencoders featuretools scipy nltk sentence-transformers
    


    Why Feature Engineering Matters

    Consider a dataset with a dateofbirth column. Feeding the raw date string to a model is meaningless. But engineering features like age, birthmonth, isweekendbirth, or generationcohort from that single column creates highly informative signals. This is the essence of feature engineering: transforming raw data into representations that capture the underlying patterns your model needs to learn.

    Key principles:

    • Domain knowledge is your greatest asset. Understanding the business context helps you create features that no automated tool can discover.
    • Simplicity wins. A few well-chosen features often outperform hundreds of noisy ones.
    • Avoid data leakage. Never use information from the future or the target variable during feature creation.


    Numerical Feature Transformations

    Scaling and Normalization

    Many algorithms (SVM, KNN, neural networks, regularized regression) are sensitive to feature scales. Scaling ensures all features contribute equally.

    import pandas as pd
    

    import numpy as np

    from sklearn.preprocessing import StandardScaler, MinMaxScaler, RobustScaler

    Sample data

    df = pd.DataFrame({

    'income': [30000, 45000, 120000, 55000, 250000, 80000],

    'age': [22, 35, 45, 28, 60, 40],

    'transactionscount': [5, 120, 450, 30, 800, 200],

    })

    StandardScaler: zero mean, unit variance (best for normally distributed data)

    standardscaler = StandardScaler()

    df['incomestandard'] = standardscaler.fittransform(df[['income']])

    MinMaxScaler: scale to [0, 1] range (best when you need bounded values)

    minmaxscaler = MinMaxScaler()

    df['incomeminmax'] = minmaxscaler.fittransform(df[['income']])

    RobustScaler: uses median and IQR (best for data with outliers)

    robustscaler = RobustScaler()

    df['incomerobust'] = robustscaler.fittransform(df[['income']])

    print(df[['income', 'incomestandard', 'incomeminmax', 'incomerobust']])

    Log and Power Transforms

    Skewed distributions are common in real-world data (income, prices, counts). Log and power transforms reduce skewness, making data more normally distributed.

    from sklearn.preprocessing import PowerTransformer
    
    

    Log transform (for positive values with right skew)

    df['incomelog'] = np.log1p(df['income']) # log1p handles zero values

    Square root transform (milder than log)

    df['transactionssqrt'] = np.sqrt(df['transactionscount'])

    Box-Cox transform (finds optimal power parameter automatically)

    ptboxcox = PowerTransformer(method='box-cox') # Requires strictly positive data

    df['incomeboxcox'] = ptboxcox.fittransform(df[['income']])

    Yeo-Johnson transform (handles zero and negative values)

    ptyeojohnson = PowerTransformer(method='yeo-johnson')

    df['incomeyeojohnson'] = ptyeojohnson.fittransform(df[['income']])

    Binning and Discretization

    Converting continuous variables into discrete bins can capture non-linear relationships and reduce the impact of outliers.

    from sklearn.preprocessing import KBinsDiscretizer
    
    

    Equal-width binning

    df['agebinequal'] = pd.cut(df['age'], bins=4, labels=['young', 'mid', 'senior', 'elderly'])

    Quantile-based binning (equal frequency in each bin)

    df['incomequantile'] = pd.qcut(df['income'], q=4, labels=['Q1', 'Q2', 'Q3', 'Q4'])

    KBinsDiscretizer with different strategies

    kbd = KBinsDiscretizer(nbins=5, encode='ordinal', strategy='kmeans')

    df['incomekmeansbin'] = kbd.fittransform(df[['income']])

    Custom domain-driven bins

    def categorizeincome(income):

    if income < 40000:

    return 'low'

    elif income < 80000:

    return 'medium'

    elif income < 150000:

    return 'high'

    else:

    return 'veryhigh'

    df['incomecategory'] = df['income'].apply(categorizeincome)

    Interaction and Polynomial Features

    Combining features can reveal relationships that individual features cannot express.

    from sklearn.preprocessing import PolynomialFeatures
    
    

    Manual interaction features

    df['incomepertransaction'] = df['income'] / (df['transactionscount'] + 1)

    df['ageincomeinteraction'] = df['age'] df['income']

    Polynomial features (degree 2)

    poly = PolynomialFeatures(degree=2, includebias=False, interactiononly=False)

    numericcols = ['income', 'age', 'transactionscount']

    polyfeatures = poly.fittransform(df[numericcols])

    polynames = poly.getfeaturenamesout(numericcols)

    dfpoly = pd.DataFrame(polyfeatures, columns=polynames)

    print(f"Original features: {len(numericcols)}, Polynomial features: {len(polynames)}")


    Categorical Feature Encoding

    One-Hot Encoding

    Suitable for nominal categories with low cardinality.

    from sklearn.preprocessing import OneHotEncoder
    
    

    dfcat = pd.DataFrame({

    'color': ['red', 'blue', 'green', 'red', 'blue', 'green'],

    'size': ['S', 'M', 'L', 'M', 'L', 'S'],

    'target': [1, 0, 1, 0, 1, 0],

    })

    pandas getdummies (quick and simple)

    dfencoded = pd.getdummies(dfcat, columns=['color', 'size'], dropfirst=True)

    scikit-learn OneHotEncoder (better for pipelines)

    ohe = OneHotEncoder(sparseoutput=False, drop='first', handleunknown='ignore')

    encoded = ohe.fittransform(dfcat[['color', 'size']])

    encodeddf = pd.DataFrame(encoded, columns=ohe.getfeaturenamesout())

    Target Encoding

    Maps categories to the mean of the target variable. Powerful for high-cardinality features but requires careful handling to avoid overfitting.

    import categoryencoders as ce
    
    

    dftarget = pd.DataFrame({

    'city': ['NYC', 'LA', 'Chicago', 'NYC', 'LA', 'Chicago', 'NYC', 'LA'],

    'target': [1, 0, 1, 1, 0, 0, 1, 1],

    })

    Target encoding with smoothing to prevent overfitting

    targetencoder = ce.TargetEncoder(cols=['city'], smoothing=1.0)

    dftarget['citytargetencoded'] = targetencoder.fittransform(

    dftarget['city'], dftarget['target']

    )

    Manual implementation with k-fold to prevent leakage

    from sklearn.modelselection import KFold

    def targetencodekfold(df, column, target, nsplits=5, smoothing=10):

    """Target encoding with k-fold cross-validation to prevent leakage."""

    globalmean = df[target].mean()

    encoded = pd.Series(index=df.index, dtype=float)

    kf = KFold(nsplits=nsplits, shuffle=True, randomstate=42)

    for trainidx, validx in kf.split(df):

    train = df.iloc[trainidx]

    stats = train.groupby(column)[target].agg(['mean', 'count'])

    smoothed = (stats['count'] stats['mean'] + smoothing globalmean) / (

    stats['count'] + smoothing

    )

    encoded.iloc[validx] = df.iloc[validx][column].map(smoothed).fillna(globalmean)

    return encoded

    Frequency and Count Encoding

    Replace categories with their occurrence frequency. Simple yet effective.

    def frequencyencode(df, column):
    

    """Replace category values with their frequency in the dataset."""

    freqmap = df[column].valuecounts(normalize=True).todict()

    return df[column].map(freqmap)

    def countencode(df, column):

    """Replace category values with their count in the dataset."""

    countmap = df[column].valuecounts().todict()

    return df[column].map(countmap)

    dfcat['colorfreq'] = frequencyencode(dfcat, 'color')

    dfcat['colorcount'] = countencode(dfcat, 'color')

    Ordinal and Binary Encoding

    import categoryencoders as ce
    
    

    Ordinal encoding (for naturally ordered categories)

    ordinalmap = {'S': 1, 'M': 2, 'L': 3, 'XL': 4}

    dfcat['sizeordinal'] = dfcat['size'].map(ordinalmap)

    Binary encoding (efficient for high-cardinality)

    binaryencoder = ce.BinaryEncoder(cols=['color'])

    dfbinary = binaryencoder.fittransform(dfcat[['color']])

    print(dfbinary)


    Datetime Feature Engineering

    Datetime columns contain rich temporal signals. Extracting them properly is critical for time-series and event-based models.

    import pandas as pd
    

    import numpy as np

    Create sample datetime data

    dftime = pd.DataFrame({

    'eventtime': pd.daterange('2025-01-01', periods=1000, freq='H'),

    'userid': np.random.randint(1, 50, 1000),

    'amount': np.random.uniform(10, 500, 1000),

    })

    Basic temporal components

    dftime['hour'] = dftime['eventtime'].dt.hour

    dftime['dayofweek'] = dftime['eventtime'].dt.dayofweek # 0=Monday

    dftime['dayofmonth'] = dftime['eventtime'].dt.day

    dftime['month'] = dftime['eventtime'].dt.month

    dftime['quarter'] = dftime['eventtime'].dt.quarter

    dftime['year'] = dftime['eventtime'].dt.year

    dftime['weekofyear'] = dftime['eventtime'].dt.isocalendar().week.astype(int)

    Boolean flags

    dftime['isweekend'] = dftime['dayofweek'].isin([5, 6]).astype(int)

    dftime['ismonthstart'] = dftime['eventtime'].dt.ismonthstart.astype(int)

    dftime['ismonthend'] = dftime['eventtime'].dt.ismonthend.astype(int)

    Cyclical encoding (important for models to understand that hour 23 is close to hour 0)

    dftime['hoursin'] = np.sin(2 np.pi dftime['hour'] / 24)

    dftime['hourcos'] = np.cos(2 np.pi dftime['hour'] / 24)

    dftime['dowsin'] = np.sin(2 np.pi dftime['dayofweek'] / 7)

    dftime['dowcos'] = np.cos(2 np.pi dftime['dayofweek'] / 7)

    dftime['monthsin'] = np.sin(2 np.pi dftime['month'] / 12)

    dftime['monthcos'] = np.cos(2 np.pi * dftime['month'] / 12)

    Time-based aggregations per user

    dftime = dftime.sortvalues(['userid', 'eventtime'])

    dftime['timesincelastevent'] = dftime.groupby('userid')['eventtime'].diff().dt.totalseconds()

    dftime['rollingavgamount24h'] = (

    dftime.groupby('userid')['amount']

    .transform(lambda x: x.rolling(window=24, minperiods=1).mean())

    )

    Part of day

    def getpartofday(hour):

    if 5 <= hour < 12:

    return 'morning'

    elif 12 <= hour < 17:

    return 'afternoon'

    elif 17 <= hour < 21:

    return 'evening'

    else:

    return 'night'

    dftime['partofday'] = dftime['hour'].apply(getpartofday)


    Text Feature Engineering

    TF-IDF Vectorization

    from sklearn.featureextraction.text import TfidfVectorizer
    
    

    documents = [

    "machine learning is a subset of artificial intelligence",

    "deep learning uses neural networks with many layers",

    "natural language processing deals with text data",

    "computer vision analyzes images and video",

    "reinforcement learning trains agents through rewards",

    ]

    Basic TF-IDF

    tfidf = TfidfVectorizer(

    maxfeatures=100,

    mindf=1,

    maxdf=0.95,

    ngramrange=(1, 2), # Unigrams and bigrams

    stopwords='english',

    sublineartf=True, # Apply log normalization

    )

    tfidfmatrix = tfidf.fittransform(documents)

    featurenames = tfidf.getfeaturenamesout()

    print(f"TF-IDF features: {tfidfmatrix.shape[1]}")

    Convert to DataFrame for inspection

    tfidfdf = pd.DataFrame(tfidfmatrix.toarray(), columns=featurenames)

    Statistical Text Features

    Simple but surprisingly effective features extracted from raw text:

    import re
    
    

    def extracttextfeatures(text):

    """Extract statistical features from text."""

    words = text.split()

    sentences = text.split('.')

    return {

    'charcount': len(text),

    'wordcount': len(words),

    'sentencecount': len([s for s in sentences if s.strip()]),

    'avgwordlength': np.mean([len(w) for w in words]) if words else 0,

    'maxwordlength': max([len(w) for w in words]) if words else 0,

    'uniquewordratio': len(set(words)) / len(words) if words else 0,

    'uppercaseratio': sum(1 for c in text if c.isupper()) / len(text) if text else 0,

    'digitratio': sum(1 for c in text if c.isdigit()) / len(text) if text else 0,

    'punctuationcount': len(re.findall(r'[^\w\s]', text)),

    'exclamationcount': text.count('!'),

    'questioncount': text.count('?'),

    }

    Apply to DataFrame

    dftext = pd.DataFrame({'text': documents})

    textfeatures = dftext['text'].apply(lambda x: pd.Series(extracttextfeatures(x)))

    dftext = pd.concat([dftext, textfeatures], axis=1)

    Sentence Embeddings

    For tasks requiring semantic understanding, dense embeddings outperform sparse TF-IDF:

    from sentencetransformers import SentenceTransformer
    
    

    Load a pre-trained sentence transformer model

    model = SentenceTransformer('all-MiniLM-L6-v2')

    Generate embeddings (384-dimensional vectors)

    embeddings = model.encode(documents, showprogressbar=True)

    print(f"Embedding shape: {embeddings.shape}") # (5, 384)

    Create a DataFrame with embedding features

    embeddingcols = [f'emb{i}' for i in range(embeddings.shape[1])]

    dfembeddings = pd.DataFrame(embeddings, columns=embeddingcols)

    Optionally reduce dimensions with PCA

    from sklearn.decomposition import PCA

    pca = PCA(ncomponents=50, randomstate=42)

    reducedembeddings = pca.fittransform(embeddings)

    print(f"Explained variance: {pca.explainedvarianceratio.sum():.2%}")


    Feature Selection Techniques

    After engineering many features, selecting the most relevant ones reduces overfitting, improves interpretability, and speeds up training.

    Filter Methods

    from sklearn.featureselection import mutualinfoclassif, fclassif, SelectKBest
    

    from sklearn.datasets import makeclassification

    Generate synthetic dataset

    X, y = makeclassification(nsamples=1000, nfeatures=50, ninformative=10,

    nredundant=10, randomstate=42)

    featurenames = [f'feature{i}' for i in range(50)]

    dfselect = pd.DataFrame(X, columns=featurenames)

    Mutual Information (works for non-linear relationships)

    miscores = mutualinfoclassif(X, y, randomstate=42)

    miranking = pd.Series(miscores, index=featurenames).sortvalues(ascending=False)

    print("Top 10 features by Mutual Information:")

    print(miranking.head(10))

    ANOVA F-test (best for linear relationships)

    selector = SelectKBest(scorefunc=fclassif, k=10)

    Xselected = selector.fittransform(X, y)

    selectedmask = selector.getsupport()

    selectedfeatures = [f for f, s in zip(featurenames, selectedmask) if s]

    Correlation-Based Removal

    def removecorrelatedfeatures(df, threshold=0.90):
    

    """Remove features that are highly correlated with each other."""

    corrmatrix = df.corr().abs()

    uppertriangle = corrmatrix.where(

    np.triu(np.ones(corrmatrix.shape), k=1).astype(bool)

    )

    todrop = [col for col in uppertriangle.columns

    if any(uppertriangle[col] > threshold)]

    print(f"Removing {len(todrop)} correlated features: {todrop}")

    return df.drop(columns=todrop)

    Model-Based Selection (Feature Importance)

    from sklearn.ensemble import RandomForestClassifier
    

    from sklearn.featureselection import SelectFromModel

    Train a Random Forest and use feature importances

    rf = RandomForestClassifier(nestimators=100, randomstate=42)

    rf.fit(X, y)

    importances = pd.Series(rf.featureimportances, index=featurenames)

    importances = importances.sortvalues(ascending=False)

    print("Top 10 features by Random Forest importance:")

    print(importances.head(10))

    Automatic selection based on importance threshold

    selector = SelectFromModel(rf, threshold='median')

    Xselected = selector.fittransform(X, y)

    print(f"Selected {Xselected.shape[1]} features out of {X.shape[1]}")


    Automated Feature Engineering with Featuretools

    Featuretools automates the generation of features from relational datasets using Deep Feature Synthesis (DFS).

    import featuretools as ft
    
    

    Create sample entities (tables)

    customers = pd.DataFrame({

    'customerid': range(1, 6),

    'signupdate': pd.daterange('2024-01-01', periods=5, freq='30D'),

    'country': ['US', 'UK', 'DE', 'US', 'FR'],

    })

    transactions = pd.DataFrame({

    'transactionid': range(1, 21),

    'customerid': np.random.choice(range(1, 6), 20),

    'amount': np.random.uniform(10, 500, 20).round(2),

    'category': np.random.choice(['food', 'electronics', 'clothing', 'travel'], 20),

    'transactiondate': pd.daterange('2024-06-01', periods=20, freq='3D'),

    })

    Create an EntitySet

    es = ft.EntitySet(id='ecommerce')

    es = es.adddataframe(

    dataframename='customers',

    dataframe=customers,

    index='customerid',

    timeindex='signupdate',

    )

    es = es.adddataframe(

    dataframename='transactions',

    dataframe=transactions,

    index='transactionid',

    timeindex='transactiondate',

    )

    es = es.addrelationship('customers', 'customerid', 'transactions', 'customerid')

    Run Deep Feature Synthesis

    featurematrix, featuredefs = ft.dfs(

    entityset=es,

    targetdataframename='customers',

    aggprimitives=['mean', 'sum', 'max', 'min', 'count', 'std', 'numunique'],

    transprimitives=['month', 'weekday', 'year'],

    maxdepth=2,

    )

    print(f"Generated {len(featuredefs)} features automatically")

    print(featurematrix.columns.tolist())


    Building a Complete Feature Pipeline

    Combine all techniques into a reusable sklearn pipeline:

    from sklearn.pipeline import Pipeline
    

    from sklearn.compose import ColumnTransformer

    from sklearn.impute import SimpleImputer

    from sklearn.preprocessing import StandardScaler, OneHotEncoder, FunctionTransformer

    Define column groups

    numericfeatures = ['income', 'age', 'transactionscount']

    categoricalfeatures = ['city', 'devicetype']

    Numeric pipeline

    numericpipeline = Pipeline([

    ('imputer', SimpleImputer(strategy='median')),

    ('logtransform', FunctionTransformer(np.log1p, validate=True)),

    ('scaler', StandardScaler()),

    ])

    Categorical pipeline

    categoricalpipeline = Pipeline([

    ('imputer', SimpleImputer(strategy='constant', fillvalue='unknown')),

    ('encoder', OneHotEncoder(handleunknown='ignore', sparseoutput=False)),

    ])

    Combined preprocessor

    preprocessor = ColumnTransformer([

    ('numeric', numericpipeline, numericfeatures),

    ('categorical', categoricalpipeline, categoricalfeatures),

    ])

    Full pipeline with model

    from sklearn.ensemble import GradientBoostingClassifier

    fullpipeline = Pipeline([

    ('preprocessor', preprocessor),

    ('classifier', GradientBoostingClassifier(nestimators=100, randomstate=42)),

    ])


    Best Practices

  • Always split data before feature engineering. Fit transformers on training data only, then transform validation and test sets to prevent data leakage.
  • Document every feature. Maintain a feature registry with descriptions, data types, expected ranges, and the logic behind each feature.
  • Monitor feature distributions in production. Feature drift (distribution shift between training and serving) is a leading cause of model degradation. Set up alerts for distribution changes.
  • Start simple, iterate. Begin with raw features and a simple model to establish a baseline. Then add features incrementally, measuring the impact of each.
  • Handle missing values intentionally. Missingness itself is often a signal. Create boolean "ismissing" flags before imputing.
  • Be mindful of cardinality. Features with thousands of unique values (like ZIP codes) need special encoding strategies like target encoding or hashing.
  • Use cross-validation for feature selection. Selecting features on the entire dataset and then evaluating on a test split leaks information. Perform feature selection inside the cross-validation loop.

  • Conclusion

    Feature engineering remains the most effective lever for improving machine learning model performance. In this masterclass, you learned:

    • Numerical transformations: scaling, log/power transforms, binning, polynomial features
    • Categorical encoding: one-hot, target, frequency, binary, and ordinal encoding
    • Datetime feature extraction: cyclical encoding, temporal aggregations, part-of-day flags
    • Text features: TF-IDF, statistical features, and dense sentence embeddings
    • Feature selection: filter methods, correlation analysis, and model-based importance
    • Automated feature engineering with Featuretools

    The key takeaway is that feature engineering is not a one-time activity. It is an iterative process driven by domain knowledge, experimentation, and continuous evaluation. Master these techniques, and you will consistently build models that outperform those that rely on raw data alone.

    Related Articles

    Complete Comet ML Tutorial: MLOps Platform for Experiment Tracking and Model Management

    Tutorial Lengkap Comet ML: Platform MLOps untuk Experiment Tracking dan Model Management Dalam dunia machine learning mo...

    MLX Tutorial: Apple's Machine Learning Framework for Apple Silicon

    Tutorial MLX: Framework Machine Learning Apple untuk Apple Silicon MLX adalah framework machine learning open-source dar...

    Marimo Tutorial: Reactive and Reproducible Python Notebooks

    Marimo: Notebook Python yang Reaktif dan Reproducible Marimo adalah notebook Python yang menyimpan isinya sebagai berkas...

    Kedro Tutorial: Reproducible and Maintainable Data Science Pipelines

    Kedro: Pipeline Data Science yang Reproducible dan Mudah Dirawat Sebagian besar proyek data science dimulai dari satu no...