Tutorial Lengkap AWS SageMaker Feature Store: Manajemen Fitur untuk ML
Amazon SageMaker Feature Store adalah repositori terkelola penuh untuk menyimpan, berbagi, dan mengelola fitur ML. Layanan ini menyediakan penyimpanan terpusat untuk fitur yang dapat digunakan di seluruh training dan inference, memastikan konsistensi dan reusabilitas.
Mengapa Feature Store?
Manfaat Utama:- Konsistensi: Fitur sama untuk training dan inference
- Reusabilitas: Berbagi fitur antar tim dan model
- Versioning: Lacak perubahan fitur seiring waktu
- Latensi rendah: Serving fitur real-time
- Penyimpanan offline: Data historis untuk training
- Feature Groups
- Online Store (real-time)
- Offline Store (batch/training)
- Feature Definitions
Prerequisites
pip install sagemaker boto3 pandas
SageMaker SDK >= 2.0
python -c "import sagemaker; print(sagemaker.version)"
Quick Start
1. Setup
import boto3
import sagemaker
from sagemaker.featurestore.featuregroup import FeatureGroup
import pandas as pd
import time
session = sagemaker.Session()
bucket = session.defaultbucket()
role = sagemaker.getexecutionrole()
region = session.botoregionname
featurestoresession = sagemaker.Session()
2. Siapkan Data
import pandas as pd
import numpy as np
from datetime import datetime
Buat sample data pelanggan
customerdata = pd.DataFrame({
"customerid": [f"C{i:04d}" for i in range(1, 101)],
"age": np.random.randint(18, 70, 100),
"income": np.random.randint(30000, 150000, 100),
"creditscore": np.random.randint(300, 850, 100),
"accountbalance": np.random.uniform(0, 50000, 100).round(2),
"numproducts": np.random.randint(1, 5, 100),
"isactive": np.random.choice([0, 1], 100)
})
Tambahkan event time (wajib untuk Feature Store)
customerdata["eventtime"] = datetime.now().strftime("%Y-%m-%dT%H:%M:%SZ")
print(customerdata.head())
Membuat Feature Groups
1. Definisikan Feature Group
from sagemaker.featurestore.featuredefinition import (
FeatureDefinition,
FeatureTypeEnum
)
Nama feature group
featuregroupname = "customer-features"
Buat feature group
customerfeaturegroup = FeatureGroup(
name=featuregroupname,
sagemakersession=featurestoresession
)
Load definisi fitur dari DataFrame
customerfeaturegroup.loadfeaturedefinitions(dataframe=customerdata)
Atau definisikan manual
featuredefinitions = [
FeatureDefinition(featurename="customerid", featuretype=FeatureTypeEnum.STRING),
FeatureDefinition(featurename="age", featuretype=FeatureTypeEnum.INTEGRAL),
FeatureDefinition(featurename="income", featuretype=FeatureTypeEnum.INTEGRAL),
FeatureDefinition(featurename="creditscore", featuretype=FeatureTypeEnum.INTEGRAL),
FeatureDefinition(featurename="accountbalance", featuretype=FeatureTypeEnum.FRACTIONAL),
FeatureDefinition(featurename="numproducts", featuretype=FeatureTypeEnum.INTEGRAL),
FeatureDefinition(featurename="isactive", featuretype=FeatureTypeEnum.INTEGRAL),
FeatureDefinition(featurename="eventtime", featuretype=FeatureTypeEnum.STRING)
]
2. Buat Feature Group
# Buat feature group dengan online dan offline stores
customerfeaturegroup.create(
s3uri=f"s3://{bucket}/feature-store/",
recordidentifiername="customerid",
eventtimefeaturename="eventtime",
rolearn=role,
enableonlinestore=True,
description="Fitur pelanggan untuk prediksi churn"
)
Tunggu feature group selesai dibuat
status = customerfeaturegroup.describe().get("FeatureGroupStatus")
while status == "Creating":
print(f"Status: {status}")
time.sleep(5)
status = customerfeaturegroup.describe().get("FeatureGroupStatus")
print(f"Status feature group: {status}")
3. Ingest Fitur
# Ingest data ke feature store
customerfeaturegroup.ingest(
dataframe=customerdata,
maxworkers=4,
wait=True
)
print("Ingestion fitur selesai!")
Online Store (Real-time)
1. Get Single Record
# Get single record berdasarkan identifier
record = customerfeaturegroup.getrecord(
recordidentifiervalueasstring="C0001"
)
print("Record untuk C0001:")
for feature in record:
print(f" {feature['FeatureName']}: {feature['ValueAsString']}")
2. Batch Get Records
from sagemaker.featurestore.featurestore import FeatureStore
feature
store = FeatureStore(sagemakersession=featurestoresession)
Batch get records
identifiers = [
{"FeatureGroupName": feature
groupname, "RecordIdentifiersValueAsString": ["C0001", "C0002", "C0003"]}
]
response = feature
store.batchgetrecord(identifiers=identifiers)
for record in response["Records"]:
customerid = next(f["ValueAsString"] for f in record["Record"] if f["FeatureName"] == "customerid")
print(f"Pelanggan: {customerid}")
3. Integrasi Real-time Inference
import json
def getfeaturesforinference(customerid):
"""Dapatkan fitur untuk real-time inference."""
record = customerfeaturegroup.getrecord(
recordidentifiervalueasstring=customerid
)
# Konversi ke feature vector
features = {f["FeatureName"]: f["ValueAsString"] for f in record}
return [
float(features["age"]),
float(features["income"]),
float(features["creditscore"]),
float(features["accountbalance"]),
float(features["numproducts"])
]
Gunakan dalam inference
features = getfeaturesforinference("C0001")
print(f"Fitur untuk inference: {features}")
Offline Store (Training)
1. Query dengan Athena
from sagemaker.featurestore.featuregroup import AthenaQuery
Buat Athena query
athena
query = customerfeaturegroup.athenaquery()
Dapatkan nama tabel
table
name = athenaquery.tablename
databasename = athenaquery.database
print(f"Database: {databasename}")
print(f"Tabel: {tablename}")
Jalankan query
querystring = f"""
SELECT *
FROM "{tablename}"
WHERE isactive = 1
LIMIT 100
"""
athenaquery.run(
querystring=querystring,
outputlocation=f"s3://{bucket}/athena-results/"
)
Tunggu dan dapatkan hasil
athenaquery.wait()
df = athenaquery.asdataframe()
print(df.head())
2. Join Multiple Feature Groups
# Buat feature group lain untuk transaksi
transactiondata = pd.DataFrame({
"customerid": [f"C{i:04d}" for i in np.random.randint(1, 101, 500)],
"transactionid": [f"T{i:06d}" for i in range(1, 501)],
"amount": np.random.uniform(10, 1000, 500).round(2),
"transactiontype": np.random.choice(["purchase", "refund", "transfer"], 500),
"eventtime": datetime.now().strftime("%Y-%m-%dT%H:%M:%SZ")
})
Buat dan ingest transaction feature group
transactionfeaturegroup = FeatureGroup(
name="transaction-features",
sagemakersession=featurestoresession
)
transactionfeaturegroup.loadfeaturedefinitions(dataframe=transactiondata)
transactionfeaturegroup.create(
s3uri=f"s3://{bucket}/feature-store/",
recordidentifiername="transactionid",
eventtimefeaturename="eventtime",
rolearn=role,
enableonlinestore=True
)
Tunggu pembuatan
time.sleep(60)
transactionfeaturegroup.ingest(dataframe=transactiondata, wait=True)
# Query join
joinquery = f"""
SELECT
c.customerid,
c.age,
c.income,
c.creditscore,
t.amount,
t.transactiontype
FROM "{customerfeaturegroup.athenaquery().tablename}" c
JOIN "{transactionfeaturegroup.athenaquery().tablename}" t
ON c.customerid = t.customerid
WHERE c.isactive = 1
"""
athenaquery.run(
querystring=joinquery,
outputlocation=f"s3://{bucket}/athena-results/"
)
athenaquery.wait()
trainingdf = athenaquery.asdataframe()
3. Buat Training Dataset
from sagemaker.featurestore.datasetbuilder import DatasetBuilder
Bangun training dataset
builder = DatasetBuilder(
sagemakersession=featurestoresession,
base=customerfeaturegroup,
outputpath=f"s3://{bucket}/training-data/"
)
Tambahkan point-in-time join
builder = builder.withfeaturegroup(
featuregroup=transactionfeaturegroup,
targetfeaturenameinbase="customerid",
includedfeaturenames=["amount", "transactiontype"]
)
Bangun dataset
trainingdataset = builder.todataframe()
print(f"Shape training dataset: {trainingdataset.shape}")
Update Fitur
1. Update Records
# Update data pelanggan yang ada
updateddata = pd.DataFrame({
"customerid": ["C0001"],
"age": [35],
"income": [85000],
"creditscore": [750],
"accountbalance": [25000.00],
"numproducts": [3],
"isactive": [1],
"eventtime": datetime.now().strftime("%Y-%m-%dT%H:%M:%SZ")
})
Ingest data yang diupdate
customerfeaturegroup.ingest(
dataframe=updateddata,
maxworkers=1,
wait=True
)
print("Record diupdate!")
Verifikasi update
record = customerfeaturegroup.getrecord(
recordidentifiervalueasstring="C0001"
)
print("Record yang diupdate:", record)
2. Hapus Records
# Hapus record (soft delete dengan deletion marker)
customerfeaturegroup.deleterecord(
recordidentifiervalueasstring="C0099",
eventtime=datetime.now().strftime("%Y-%m-%dT%H:%M:%SZ")
)
Feature Store SDK
1. Feature Processor
from sagemaker.featurestore.featureprocessor import (
FeatureProcessor,
CSVDataSource,
FeatureGroupDataSource
)
Buat feature processor
@feature
processor(
inputs=[
CSVDataSource(s3uri=f"s3://{bucket}/raw-data/")
],
output=featuregroupname
)
def processfeatures(inputdf):
# Feature engineering
df = inputdf.copy()
# Tambah derived features
df["incomeperproduct"] = df["income"] / df["numproducts"]
df["creditincomeratio"] = df["creditscore"] / (df["income"] / 10000)
return df
2. Scheduled Ingestion
from sagemaker.featurestore.featureprocessor import (
FeatureProcessorPipelineEvents,
topipeline
)
Buat pipeline untuk scheduled ingestion
pipeline = topipeline(
pipelinename="feature-ingestion-pipeline",
step=processfeatures,
role=role
)
Jadwalkan eksekusi
pipeline.puttriggers(
triggers=[
FeatureProcessorPipelineEvents(
eventpattern={
"source": ["aws.s3"],
"detail-type": ["Object Created"],
"detail": {
"bucket": {"name": [bucket]},
"object": {"key": [{"prefix": "raw-data/"}]}
}
}
)
]
)
Monitoring dan Manajemen
1. Describe Feature Group
# Dapatkan detail feature group
description = customerfeaturegroup.describe()
print(f"Nama: {description['FeatureGroupName']}")
print(f"Status: {description['FeatureGroupStatus']}")
print(f"Record Identifier: {description['RecordIdentifierFeatureName']}")
print(f"Event Time: {description['EventTimeFeatureName']}")
List definisi fitur
for feature in description['FeatureDefinitions']:
print(f" {feature['FeatureName']}: {feature['FeatureType']}")
2. List Feature Groups
# List semua feature groups
sagemakerclient = boto3.client("sagemaker")
response = sagemakerclient.listfeaturegroups()
for fg in response["FeatureGroupSummaries"]:
print(f"{fg['FeatureGroupName']}: {fg['FeatureGroupStatus']}")
3. Hapus Feature Group
# Hapus feature group
customerfeaturegroup.delete()
Catatan: Data offline store di S3 tidak dihapus otomatis
Best Practices
1. Konvensi Penamaan Fitur
# Bagus: Penamaan deskriptif dan konsisten
featuredefinitions = [
FeatureDefinition("customerid", FeatureTypeEnum.STRING),
FeatureDefinition("customerageyears", FeatureTypeEnum.INTEGRAL),
FeatureDefinition("customertotalspendusd", FeatureTypeEnum.FRACTIONAL),
FeatureDefinition("customerispremiumflag", FeatureTypeEnum.INTEGRAL)
]
2. Manajemen Event Time
from datetime import datetime, timezone
def geteventtime():
"""Generate event time yang konsisten."""
return datetime.now(timezone.utc).strftime("%Y-%m-%dT%H:%M:%SZ")
Selalu sertakan eventtime dalam data Anda
data["eventtime"] = geteventtime()
3. Batch Ingestion
def batchingest(featuregroup, df, batchsize=1000):
"""Ingest data dalam batch untuk reliabilitas lebih baik."""
for i in range(0, len(df), batch
size):
batch = df.iloc[i:i+batchsize]
featuregroup.ingest(
dataframe=batch,
maxworkers=4,
wait=True
)
print(f"Batch {i//batch_size + 1} di-ingest")
Kesimpulan
SageMaker Feature Store menyediakan:
Key takeaways:
- Rancang fitur dengan reusabilitas dalam pikiran
- Gunakan online store untuk real-time inference
- Manfaatkan Athena untuk offline training queries
- Implementasikan manajemen event time yang proper
- Monitor kesegaran dan kualitas fitur