Tutorial 14: Feature Engineering Masterclass
Table of Contents
Introduction
Feature engineering is widely regarded as the most impactful activity in applied machine learning. While model architectures and hyperparameter tuning receive significant attention, the quality and expressiveness of your input features are what ultimately determine model performance. A well-engineered feature can make a simple linear model outperform a poorly-featured deep neural network.
This masterclass covers the full spectrum of feature engineering techniques: from fundamental numerical transformations and categorical encoding strategies to advanced datetime extraction, text vectorization, and automated feature generation. Each technique is demonstrated with production-ready Python code and practical guidance on when and how to apply it.
Prerequisites
- Python 3.9 or higher
- Solid understanding of pandas and NumPy
- Basic machine learning knowledge (classification, regression)
- Install required packages:
pip install pandas numpy scikit-learn categoryencoders featuretools scipy nltk sentence-transformers
Why Feature Engineering Matters
Consider a dataset with a dateofbirth column. Feeding the raw date string to a model is meaningless. But engineering features like age, birthmonth, isweekendbirth, or generationcohort from that single column creates highly informative signals. This is the essence of feature engineering: transforming raw data into representations that capture the underlying patterns your model needs to learn.
Key principles:
- Domain knowledge is your greatest asset. Understanding the business context helps you create features that no automated tool can discover.
- Simplicity wins. A few well-chosen features often outperform hundreds of noisy ones.
- Avoid data leakage. Never use information from the future or the target variable during feature creation.
Numerical Feature Transformations
Scaling and Normalization
Many algorithms (SVM, KNN, neural networks, regularized regression) are sensitive to feature scales. Scaling ensures all features contribute equally.
import pandas as pd
import numpy as np
from sklearn.preprocessing import StandardScaler, MinMaxScaler, RobustScaler
Sample data
df = pd.DataFrame({
'income': [30000, 45000, 120000, 55000, 250000, 80000],
'age': [22, 35, 45, 28, 60, 40],
'transactionscount': [5, 120, 450, 30, 800, 200],
})
StandardScaler: zero mean, unit variance (best for normally distributed data)
standardscaler = StandardScaler()
df['incomestandard'] = standardscaler.fittransform(df[['income']])
MinMaxScaler: scale to [0, 1] range (best when you need bounded values)
minmaxscaler = MinMaxScaler()
df['incomeminmax'] = minmaxscaler.fittransform(df[['income']])
RobustScaler: uses median and IQR (best for data with outliers)
robustscaler = RobustScaler()
df['incomerobust'] = robustscaler.fittransform(df[['income']])
print(df[['income', 'incomestandard', 'incomeminmax', 'incomerobust']])
Log and Power Transforms
Skewed distributions are common in real-world data (income, prices, counts). Log and power transforms reduce skewness, making data more normally distributed.
from sklearn.preprocessing import PowerTransformer
Log transform (for positive values with right skew)
df['incomelog'] = np.log1p(df['income']) # log1p handles zero values
Square root transform (milder than log)
df['transactionssqrt'] = np.sqrt(df['transactionscount'])
Box-Cox transform (finds optimal power parameter automatically)
ptboxcox = PowerTransformer(method='box-cox') # Requires strictly positive data
df['incomeboxcox'] = ptboxcox.fittransform(df[['income']])
Yeo-Johnson transform (handles zero and negative values)
ptyeojohnson = PowerTransformer(method='yeo-johnson')
df['incomeyeojohnson'] = ptyeojohnson.fittransform(df[['income']])
Binning and Discretization
Converting continuous variables into discrete bins can capture non-linear relationships and reduce the impact of outliers.
from sklearn.preprocessing import KBinsDiscretizer
Equal-width binning
df['agebinequal'] = pd.cut(df['age'], bins=4, labels=['young', 'mid', 'senior', 'elderly'])
Quantile-based binning (equal frequency in each bin)
df['incomequantile'] = pd.qcut(df['income'], q=4, labels=['Q1', 'Q2', 'Q3', 'Q4'])
KBinsDiscretizer with different strategies
kbd = KBinsDiscretizer(nbins=5, encode='ordinal', strategy='kmeans')
df['incomekmeansbin'] = kbd.fittransform(df[['income']])
Custom domain-driven bins
def categorizeincome(income):
if income < 40000:
return 'low'
elif income < 80000:
return 'medium'
elif income < 150000:
return 'high'
else:
return 'veryhigh'
df['incomecategory'] = df['income'].apply(categorizeincome)
Interaction and Polynomial Features
Combining features can reveal relationships that individual features cannot express.
from sklearn.preprocessing import PolynomialFeatures
Manual interaction features
df['incomepertransaction'] = df['income'] / (df['transactionscount'] + 1)
df['ageincomeinteraction'] = df['age'] df['income']
Polynomial features (degree 2)
poly = PolynomialFeatures(degree=2, includebias=False, interactiononly=False)
numericcols = ['income', 'age', 'transactionscount']
polyfeatures = poly.fittransform(df[numericcols])
polynames = poly.getfeaturenamesout(numericcols)
dfpoly = pd.DataFrame(polyfeatures, columns=polynames)
print(f"Original features: {len(numericcols)}, Polynomial features: {len(polynames)}")
Categorical Feature Encoding
One-Hot Encoding
Suitable for nominal categories with low cardinality.
from sklearn.preprocessing import OneHotEncoder
dfcat = pd.DataFrame({
'color': ['red', 'blue', 'green', 'red', 'blue', 'green'],
'size': ['S', 'M', 'L', 'M', 'L', 'S'],
'target': [1, 0, 1, 0, 1, 0],
})
pandas getdummies (quick and simple)
dfencoded = pd.getdummies(dfcat, columns=['color', 'size'], dropfirst=True)
scikit-learn OneHotEncoder (better for pipelines)
ohe = OneHotEncoder(sparseoutput=False, drop='first', handleunknown='ignore')
encoded = ohe.fittransform(dfcat[['color', 'size']])
encodeddf = pd.DataFrame(encoded, columns=ohe.getfeaturenamesout())
Target Encoding
Maps categories to the mean of the target variable. Powerful for high-cardinality features but requires careful handling to avoid overfitting.
import categoryencoders as ce
dftarget = pd.DataFrame({
'city': ['NYC', 'LA', 'Chicago', 'NYC', 'LA', 'Chicago', 'NYC', 'LA'],
'target': [1, 0, 1, 1, 0, 0, 1, 1],
})
Target encoding with smoothing to prevent overfitting
targetencoder = ce.TargetEncoder(cols=['city'], smoothing=1.0)
dftarget['citytargetencoded'] = targetencoder.fittransform(
dftarget['city'], dftarget['target']
)
Manual implementation with k-fold to prevent leakage
from sklearn.modelselection import KFold
def targetencodekfold(df, column, target, nsplits=5, smoothing=10):
"""Target encoding with k-fold cross-validation to prevent leakage."""
globalmean = df[target].mean()
encoded = pd.Series(index=df.index, dtype=float)
kf = KFold(nsplits=nsplits, shuffle=True, randomstate=42)
for trainidx, validx in kf.split(df):
train = df.iloc[trainidx]
stats = train.groupby(column)[target].agg(['mean', 'count'])
smoothed = (stats['count'] stats['mean'] + smoothing globalmean) / (
stats['count'] + smoothing
)
encoded.iloc[validx] = df.iloc[validx][column].map(smoothed).fillna(globalmean)
return encoded
Frequency and Count Encoding
Replace categories with their occurrence frequency. Simple yet effective.
def frequencyencode(df, column):
"""Replace category values with their frequency in the dataset."""
freq
map = df[column].valuecounts(normalize=True).todict()
return df[column].map(freqmap)
def countencode(df, column):
"""Replace category values with their count in the dataset."""
countmap = df[column].valuecounts().todict()
return df[column].map(countmap)
dfcat['colorfreq'] = frequencyencode(dfcat, 'color')
dfcat['colorcount'] = countencode(dfcat, 'color')
Ordinal and Binary Encoding
import categoryencoders as ce
Ordinal encoding (for naturally ordered categories)
ordinal
map = {'S': 1, 'M': 2, 'L': 3, 'XL': 4}
dfcat['sizeordinal'] = dfcat['size'].map(ordinalmap)
Binary encoding (efficient for high-cardinality)
binaryencoder = ce.BinaryEncoder(cols=['color'])
dfbinary = binaryencoder.fittransform(dfcat[['color']])
print(dfbinary)
Datetime Feature Engineering
Datetime columns contain rich temporal signals. Extracting them properly is critical for time-series and event-based models.
import pandas as pd
import numpy as np
Create sample datetime data
dftime = pd.DataFrame({
'eventtime': pd.daterange('2025-01-01', periods=1000, freq='H'),
'userid': np.random.randint(1, 50, 1000),
'amount': np.random.uniform(10, 500, 1000),
})
Basic temporal components
dftime['hour'] = dftime['eventtime'].dt.hour
dftime['dayofweek'] = dftime['eventtime'].dt.dayofweek # 0=Monday
dftime['dayofmonth'] = dftime['eventtime'].dt.day
dftime['month'] = dftime['eventtime'].dt.month
dftime['quarter'] = dftime['eventtime'].dt.quarter
dftime['year'] = dftime['eventtime'].dt.year
dftime['weekofyear'] = dftime['eventtime'].dt.isocalendar().week.astype(int)
Boolean flags
dftime['isweekend'] = dftime['dayofweek'].isin([5, 6]).astype(int)
dftime['ismonthstart'] = dftime['eventtime'].dt.ismonthstart.astype(int)
dftime['ismonthend'] = dftime['eventtime'].dt.ismonthend.astype(int)
Cyclical encoding (important for models to understand that hour 23 is close to hour 0)
dftime['hoursin'] = np.sin(2 np.pi dftime['hour'] / 24)
dftime['hourcos'] = np.cos(2 np.pi dftime['hour'] / 24)
dftime['dowsin'] = np.sin(2 np.pi dftime['dayofweek'] / 7)
dftime['dowcos'] = np.cos(2 np.pi dftime['dayofweek'] / 7)
dftime['monthsin'] = np.sin(2 np.pi dftime['month'] / 12)
dftime['monthcos'] = np.cos(2 np.pi * dftime['month'] / 12)
Time-based aggregations per user
dftime = dftime.sortvalues(['userid', 'eventtime'])
dftime['timesincelastevent'] = dftime.groupby('userid')['eventtime'].diff().dt.totalseconds()
dftime['rollingavgamount24h'] = (
dftime.groupby('userid')['amount']
.transform(lambda x: x.rolling(window=24, minperiods=1).mean())
)
Part of day
def getpartofday(hour):
if 5 <= hour < 12:
return 'morning'
elif 12 <= hour < 17:
return 'afternoon'
elif 17 <= hour < 21:
return 'evening'
else:
return 'night'
dftime['partofday'] = dftime['hour'].apply(getpartofday)
Text Feature Engineering
TF-IDF Vectorization
from sklearn.featureextraction.text import TfidfVectorizer
documents = [
"machine learning is a subset of artificial intelligence",
"deep learning uses neural networks with many layers",
"natural language processing deals with text data",
"computer vision analyzes images and video",
"reinforcement learning trains agents through rewards",
]
Basic TF-IDF
tfidf = TfidfVectorizer(
max
features=100,
mindf=1,
maxdf=0.95,
ngramrange=(1, 2), # Unigrams and bigrams
stopwords='english',
sublineartf=True, # Apply log normalization
)
tfidfmatrix = tfidf.fittransform(documents)
featurenames = tfidf.getfeaturenamesout()
print(f"TF-IDF features: {tfidfmatrix.shape[1]}")
Convert to DataFrame for inspection
tfidfdf = pd.DataFrame(tfidfmatrix.toarray(), columns=featurenames)
Statistical Text Features
Simple but surprisingly effective features extracted from raw text:
import re
def extracttextfeatures(text):
"""Extract statistical features from text."""
words = text.split()
sentences = text.split('.')
return {
'charcount': len(text),
'wordcount': len(words),
'sentencecount': len([s for s in sentences if s.strip()]),
'avgwordlength': np.mean([len(w) for w in words]) if words else 0,
'maxwordlength': max([len(w) for w in words]) if words else 0,
'uniquewordratio': len(set(words)) / len(words) if words else 0,
'uppercaseratio': sum(1 for c in text if c.isupper()) / len(text) if text else 0,
'digitratio': sum(1 for c in text if c.isdigit()) / len(text) if text else 0,
'punctuationcount': len(re.findall(r'[^\w\s]', text)),
'exclamationcount': text.count('!'),
'questioncount': text.count('?'),
}
Apply to DataFrame
dftext = pd.DataFrame({'text': documents})
textfeatures = dftext['text'].apply(lambda x: pd.Series(extracttextfeatures(x)))
dftext = pd.concat([dftext, textfeatures], axis=1)
Sentence Embeddings
For tasks requiring semantic understanding, dense embeddings outperform sparse TF-IDF:
from sentencetransformers import SentenceTransformer
Load a pre-trained sentence transformer model
model = SentenceTransformer('all-MiniLM-L6-v2')
Generate embeddings (384-dimensional vectors)
embeddings = model.encode(documents, showprogressbar=True)
print(f"Embedding shape: {embeddings.shape}") # (5, 384)
Create a DataFrame with embedding features
embeddingcols = [f'emb{i}' for i in range(embeddings.shape[1])]
dfembeddings = pd.DataFrame(embeddings, columns=embeddingcols)
Optionally reduce dimensions with PCA
from sklearn.decomposition import PCA
pca = PCA(ncomponents=50, randomstate=42)
reducedembeddings = pca.fittransform(embeddings)
print(f"Explained variance: {pca.explainedvarianceratio.sum():.2%}")
Feature Selection Techniques
After engineering many features, selecting the most relevant ones reduces overfitting, improves interpretability, and speeds up training.
Filter Methods
from sklearn.featureselection import mutualinfoclassif, fclassif, SelectKBest
from sklearn.datasets import make
classification
Generate synthetic dataset
X, y = makeclassification(nsamples=1000, nfeatures=50, ninformative=10,
nredundant=10, randomstate=42)
featurenames = [f'feature{i}' for i in range(50)]
dfselect = pd.DataFrame(X, columns=featurenames)
Mutual Information (works for non-linear relationships)
miscores = mutualinfoclassif(X, y, randomstate=42)
miranking = pd.Series(miscores, index=featurenames).sortvalues(ascending=False)
print("Top 10 features by Mutual Information:")
print(miranking.head(10))
ANOVA F-test (best for linear relationships)
selector = SelectKBest(scorefunc=fclassif, k=10)
Xselected = selector.fittransform(X, y)
selectedmask = selector.getsupport()
selectedfeatures = [f for f, s in zip(featurenames, selectedmask) if s]
Correlation-Based Removal
def removecorrelatedfeatures(df, threshold=0.90):
"""Remove features that are highly correlated with each other."""
corrmatrix = df.corr().abs()
uppertriangle = corrmatrix.where(
np.triu(np.ones(corrmatrix.shape), k=1).astype(bool)
)
todrop = [col for col in uppertriangle.columns
if any(uppertriangle[col] > threshold)]
print(f"Removing {len(todrop)} correlated features: {todrop}")
return df.drop(columns=todrop)
Model-Based Selection (Feature Importance)
from sklearn.ensemble import RandomForestClassifier
from sklearn.featureselection import SelectFromModel
Train a Random Forest and use feature importances
rf = RandomForestClassifier(nestimators=100, randomstate=42)
rf.fit(X, y)
importances = pd.Series(rf.featureimportances, index=featurenames)
importances = importances.sortvalues(ascending=False)
print("Top 10 features by Random Forest importance:")
print(importances.head(10))
Automatic selection based on importance threshold
selector = SelectFromModel(rf, threshold='median')
Xselected = selector.fittransform(X, y)
print(f"Selected {Xselected.shape[1]} features out of {X.shape[1]}")
Automated Feature Engineering with Featuretools
Featuretools automates the generation of features from relational datasets using Deep Feature Synthesis (DFS).
import featuretools as ft
Create sample entities (tables)
customers = pd.DataFrame({
'customerid': range(1, 6),
'signupdate': pd.daterange('2024-01-01', periods=5, freq='30D'),
'country': ['US', 'UK', 'DE', 'US', 'FR'],
})
transactions = pd.DataFrame({
'transactionid': range(1, 21),
'customerid': np.random.choice(range(1, 6), 20),
'amount': np.random.uniform(10, 500, 20).round(2),
'category': np.random.choice(['food', 'electronics', 'clothing', 'travel'], 20),
'transactiondate': pd.daterange('2024-06-01', periods=20, freq='3D'),
})
Create an EntitySet
es = ft.EntitySet(id='ecommerce')
es = es.adddataframe(
dataframename='customers',
dataframe=customers,
index='customerid',
timeindex='signupdate',
)
es = es.adddataframe(
dataframename='transactions',
dataframe=transactions,
index='transactionid',
timeindex='transactiondate',
)
es = es.addrelationship('customers', 'customerid', 'transactions', 'customerid')
Run Deep Feature Synthesis
featurematrix, featuredefs = ft.dfs(
entityset=es,
targetdataframename='customers',
aggprimitives=['mean', 'sum', 'max', 'min', 'count', 'std', 'numunique'],
transprimitives=['month', 'weekday', 'year'],
maxdepth=2,
)
print(f"Generated {len(featuredefs)} features automatically")
print(featurematrix.columns.tolist())
Building a Complete Feature Pipeline
Combine all techniques into a reusable sklearn pipeline:
from sklearn.pipeline import Pipeline
from sklearn.compose import ColumnTransformer
from sklearn.impute import SimpleImputer
from sklearn.preprocessing import StandardScaler, OneHotEncoder, FunctionTransformer
Define column groups
numericfeatures = ['income', 'age', 'transactionscount']
categoricalfeatures = ['city', 'devicetype']
Numeric pipeline
numericpipeline = Pipeline([
('imputer', SimpleImputer(strategy='median')),
('logtransform', FunctionTransformer(np.log1p, validate=True)),
('scaler', StandardScaler()),
])
Categorical pipeline
categoricalpipeline = Pipeline([
('imputer', SimpleImputer(strategy='constant', fillvalue='unknown')),
('encoder', OneHotEncoder(handleunknown='ignore', sparseoutput=False)),
])
Combined preprocessor
preprocessor = ColumnTransformer([
('numeric', numericpipeline, numericfeatures),
('categorical', categoricalpipeline, categoricalfeatures),
])
Full pipeline with model
from sklearn.ensemble import GradientBoostingClassifier
fullpipeline = Pipeline([
('preprocessor', preprocessor),
('classifier', GradientBoostingClassifier(nestimators=100, randomstate=42)),
])
Best Practices
Conclusion
Feature engineering remains the most effective lever for improving machine learning model performance. In this masterclass, you learned:
- Numerical transformations: scaling, log/power transforms, binning, polynomial features
- Categorical encoding: one-hot, target, frequency, binary, and ordinal encoding
- Datetime feature extraction: cyclical encoding, temporal aggregations, part-of-day flags
- Text features: TF-IDF, statistical features, and dense sentence embeddings
- Feature selection: filter methods, correlation analysis, and model-based importance
- Automated feature engineering with Featuretools
The key takeaway is that feature engineering is not a one-time activity. It is an iterative process driven by domain knowledge, experimentation, and continuous evaluation. Master these techniques, and you will consistently build models that outperform those that rely on raw data alone.