Tutorial AWS SageMaker Feature Store: Manajemen Feature untuk ML

# Tutorial Lengkap AWS SageMaker Feature Store: Manajemen Fitur untuk ML Amazon SageMaker Feature Store adalah repositori terkelola penuh untuk menyimpan, berbagi, dan mengelola fitur ML. Layanan ini...

By Ruby Abdullah · · tutorial
AWSSageMakerFeature StoreFeature EngineeringMLOpsData Management

Tutorial Lengkap AWS SageMaker Feature Store: Manajemen Fitur untuk ML

Amazon SageMaker Feature Store adalah repositori terkelola penuh untuk menyimpan, berbagi, dan mengelola fitur ML. Layanan ini menyediakan penyimpanan terpusat untuk fitur yang dapat digunakan di seluruh training dan inference, memastikan konsistensi dan reusabilitas.

Mengapa Feature Store?

Manfaat Utama:
  • Konsistensi: Fitur sama untuk training dan inference
  • Reusabilitas: Berbagi fitur antar tim dan model
  • Versioning: Lacak perubahan fitur seiring waktu
  • Latensi rendah: Serving fitur real-time
  • Penyimpanan offline: Data historis untuk training

Komponen:
  • Feature Groups
  • Online Store (real-time)
  • Offline Store (batch/training)
  • Feature Definitions

Prerequisites

pip install sagemaker boto3 pandas

SageMaker SDK >= 2.0

python -c "import sagemaker; print(sagemaker.version)"

Quick Start

1. Setup

import boto3

import sagemaker

from sagemaker.featurestore.featuregroup import FeatureGroup

import pandas as pd

import time

session = sagemaker.Session()

bucket = session.defaultbucket()

role = sagemaker.getexecutionrole()

region = session.botoregionname

featurestoresession = sagemaker.Session()

2. Siapkan Data

import pandas as pd

import numpy as np

from datetime import datetime

Buat sample data pelanggan

customerdata = pd.DataFrame({

"customerid": [f"C{i:04d}" for i in range(1, 101)],

"age": np.random.randint(18, 70, 100),

"income": np.random.randint(30000, 150000, 100),

"creditscore": np.random.randint(300, 850, 100),

"accountbalance": np.random.uniform(0, 50000, 100).round(2),

"numproducts": np.random.randint(1, 5, 100),

"isactive": np.random.choice([0, 1], 100)

})

Tambahkan event time (wajib untuk Feature Store)

customerdata["eventtime"] = datetime.now().strftime("%Y-%m-%dT%H:%M:%SZ")

print(customerdata.head())

Membuat Feature Groups

1. Definisikan Feature Group

from sagemaker.featurestore.featuredefinition import (

FeatureDefinition,

FeatureTypeEnum

)

Nama feature group

featuregroupname = "customer-features"

Buat feature group

customerfeaturegroup = FeatureGroup(

name=featuregroupname,

sagemakersession=featurestoresession

)

Load definisi fitur dari DataFrame

customerfeaturegroup.loadfeaturedefinitions(dataframe=customerdata)

Atau definisikan manual

featuredefinitions = [

FeatureDefinition(featurename="customerid", featuretype=FeatureTypeEnum.STRING),

FeatureDefinition(featurename="age", featuretype=FeatureTypeEnum.INTEGRAL),

FeatureDefinition(featurename="income", featuretype=FeatureTypeEnum.INTEGRAL),

FeatureDefinition(featurename="creditscore", featuretype=FeatureTypeEnum.INTEGRAL),

FeatureDefinition(featurename="accountbalance", featuretype=FeatureTypeEnum.FRACTIONAL),

FeatureDefinition(featurename="numproducts", featuretype=FeatureTypeEnum.INTEGRAL),

FeatureDefinition(featurename="isactive", featuretype=FeatureTypeEnum.INTEGRAL),

FeatureDefinition(featurename="eventtime", featuretype=FeatureTypeEnum.STRING)

]

2. Buat Feature Group

# Buat feature group dengan online dan offline stores

customerfeaturegroup.create(

s3uri=f"s3://{bucket}/feature-store/",

recordidentifiername="customerid",

eventtimefeaturename="eventtime",

rolearn=role,

enableonlinestore=True,

description="Fitur pelanggan untuk prediksi churn"

)

Tunggu feature group selesai dibuat

status = customerfeaturegroup.describe().get("FeatureGroupStatus")

while status == "Creating":

print(f"Status: {status}")

time.sleep(5)

status = customerfeaturegroup.describe().get("FeatureGroupStatus")

print(f"Status feature group: {status}")

3. Ingest Fitur

# Ingest data ke feature store

customerfeaturegroup.ingest(

dataframe=customerdata,

maxworkers=4,

wait=True

)

print("Ingestion fitur selesai!")

Online Store (Real-time)

1. Get Single Record

# Get single record berdasarkan identifier

record = customerfeaturegroup.getrecord(

recordidentifiervalueasstring="C0001"

)

print("Record untuk C0001:")

for feature in record:

print(f" {feature['FeatureName']}: {feature['ValueAsString']}")

2. Batch Get Records

from sagemaker.featurestore.featurestore import FeatureStore

featurestore = FeatureStore(sagemakersession=featurestoresession)

Batch get records

identifiers = [

{"FeatureGroupName": featuregroupname, "RecordIdentifiersValueAsString": ["C0001", "C0002", "C0003"]}

]

response = featurestore.batchgetrecord(identifiers=identifiers)

for record in response["Records"]:

customerid = next(f["ValueAsString"] for f in record["Record"] if f["FeatureName"] == "customerid")

print(f"Pelanggan: {customerid}")

3. Integrasi Real-time Inference

import json

def getfeaturesforinference(customerid):

"""Dapatkan fitur untuk real-time inference."""

record = customerfeaturegroup.getrecord(

recordidentifiervalueasstring=customerid

)

# Konversi ke feature vector

features = {f["FeatureName"]: f["ValueAsString"] for f in record}

return [

float(features["age"]),

float(features["income"]),

float(features["creditscore"]),

float(features["accountbalance"]),

float(features["numproducts"])

]

Gunakan dalam inference

features = getfeaturesforinference("C0001")

print(f"Fitur untuk inference: {features}")

Offline Store (Training)

1. Query dengan Athena

from sagemaker.featurestore.featuregroup import AthenaQuery

Buat Athena query

athenaquery = customerfeaturegroup.athenaquery()

Dapatkan nama tabel

tablename = athenaquery.tablename

databasename = athenaquery.database

print(f"Database: {databasename}")

print(f"Tabel: {tablename}")

Jalankan query

querystring = f"""

SELECT *

FROM "{tablename}"

WHERE isactive = 1

LIMIT 100

"""

athenaquery.run(

querystring=querystring,

outputlocation=f"s3://{bucket}/athena-results/"

)

Tunggu dan dapatkan hasil

athenaquery.wait()

df = athenaquery.asdataframe()

print(df.head())

2. Join Multiple Feature Groups

# Buat feature group lain untuk transaksi

transactiondata = pd.DataFrame({

"customerid": [f"C{i:04d}" for i in np.random.randint(1, 101, 500)],

"transactionid": [f"T{i:06d}" for i in range(1, 501)],

"amount": np.random.uniform(10, 1000, 500).round(2),

"transactiontype": np.random.choice(["purchase", "refund", "transfer"], 500),

"eventtime": datetime.now().strftime("%Y-%m-%dT%H:%M:%SZ")

})

Buat dan ingest transaction feature group

transactionfeaturegroup = FeatureGroup(

name="transaction-features",

sagemakersession=featurestoresession

)

transactionfeaturegroup.loadfeaturedefinitions(dataframe=transactiondata)

transactionfeaturegroup.create(

s3uri=f"s3://{bucket}/feature-store/",

recordidentifiername="transactionid",

eventtimefeaturename="eventtime",

rolearn=role,

enableonlinestore=True

)

Tunggu pembuatan

time.sleep(60)

transactionfeaturegroup.ingest(dataframe=transactiondata, wait=True)

# Query join

joinquery = f"""

SELECT

c.customerid,

c.age,

c.income,

c.creditscore,

t.amount,

t.transactiontype

FROM "{customerfeaturegroup.athenaquery().tablename}" c

JOIN "{transactionfeaturegroup.athenaquery().tablename}" t

ON c.customerid = t.customerid

WHERE c.isactive = 1

"""

athenaquery.run(

querystring=joinquery,

outputlocation=f"s3://{bucket}/athena-results/"

)

athenaquery.wait()

trainingdf = athenaquery.asdataframe()

3. Buat Training Dataset

from sagemaker.featurestore.datasetbuilder import DatasetBuilder

Bangun training dataset

builder = DatasetBuilder(

sagemakersession=featurestoresession,

base=customerfeaturegroup,

outputpath=f"s3://{bucket}/training-data/"

)

Tambahkan point-in-time join

builder = builder.withfeaturegroup(

featuregroup=transactionfeaturegroup,

targetfeaturenameinbase="customerid",

includedfeaturenames=["amount", "transactiontype"]

)

Bangun dataset

trainingdataset = builder.todataframe()

print(f"Shape training dataset: {trainingdataset.shape}")

Update Fitur

1. Update Records

# Update data pelanggan yang ada

updateddata = pd.DataFrame({

"customerid": ["C0001"],

"age": [35],

"income": [85000],

"creditscore": [750],

"accountbalance": [25000.00],

"numproducts": [3],

"isactive": [1],

"eventtime": datetime.now().strftime("%Y-%m-%dT%H:%M:%SZ")

})

Ingest data yang diupdate

customerfeaturegroup.ingest(

dataframe=updateddata,

maxworkers=1,

wait=True

)

print("Record diupdate!")

Verifikasi update

record = customerfeaturegroup.getrecord(

recordidentifiervalueasstring="C0001"

)

print("Record yang diupdate:", record)

2. Hapus Records

# Hapus record (soft delete dengan deletion marker)

customerfeaturegroup.deleterecord(

recordidentifiervalueasstring="C0099",

eventtime=datetime.now().strftime("%Y-%m-%dT%H:%M:%SZ")

)

Feature Store SDK

1. Feature Processor

from sagemaker.featurestore.featureprocessor import (

FeatureProcessor,

CSVDataSource,

FeatureGroupDataSource

)

Buat feature processor

@featureprocessor(

inputs=[

CSVDataSource(s3uri=f"s3://{bucket}/raw-data/")

],

output=featuregroupname

)

def processfeatures(inputdf):

# Feature engineering

df = inputdf.copy()

# Tambah derived features

df["incomeperproduct"] = df["income"] / df["numproducts"]

df["creditincomeratio"] = df["creditscore"] / (df["income"] / 10000)

return df

2. Scheduled Ingestion

from sagemaker.featurestore.featureprocessor import (

FeatureProcessorPipelineEvents,

topipeline

)

Buat pipeline untuk scheduled ingestion

pipeline = topipeline(

pipelinename="feature-ingestion-pipeline",

step=processfeatures,

role=role

)

Jadwalkan eksekusi

pipeline.puttriggers(

triggers=[

FeatureProcessorPipelineEvents(

eventpattern={

"source": ["aws.s3"],

"detail-type": ["Object Created"],

"detail": {

"bucket": {"name": [bucket]},

"object": {"key": [{"prefix": "raw-data/"}]}

}

}

)

]

)

Monitoring dan Manajemen

1. Describe Feature Group

# Dapatkan detail feature group

description = customerfeaturegroup.describe()

print(f"Nama: {description['FeatureGroupName']}")

print(f"Status: {description['FeatureGroupStatus']}")

print(f"Record Identifier: {description['RecordIdentifierFeatureName']}")

print(f"Event Time: {description['EventTimeFeatureName']}")

List definisi fitur

for feature in description['FeatureDefinitions']:

print(f" {feature['FeatureName']}: {feature['FeatureType']}")

2. List Feature Groups

# List semua feature groups

sagemakerclient = boto3.client("sagemaker")

response = sagemakerclient.listfeaturegroups()

for fg in response["FeatureGroupSummaries"]:

print(f"{fg['FeatureGroupName']}: {fg['FeatureGroupStatus']}")

3. Hapus Feature Group

# Hapus feature group

customerfeaturegroup.delete()

Catatan: Data offline store di S3 tidak dihapus otomatis

Best Practices

1. Konvensi Penamaan Fitur

# Bagus: Penamaan deskriptif dan konsisten

featuredefinitions = [

FeatureDefinition("customerid", FeatureTypeEnum.STRING),

FeatureDefinition("customerageyears", FeatureTypeEnum.INTEGRAL),

FeatureDefinition("customertotalspendusd", FeatureTypeEnum.FRACTIONAL),

FeatureDefinition("customerispremiumflag", FeatureTypeEnum.INTEGRAL)

]

2. Manajemen Event Time

from datetime import datetime, timezone

def geteventtime():

"""Generate event time yang konsisten."""

return datetime.now(timezone.utc).strftime("%Y-%m-%dT%H:%M:%SZ")

Selalu sertakan eventtime dalam data Anda

data["eventtime"] = geteventtime()

3. Batch Ingestion

def batchingest(featuregroup, df, batchsize=1000):

"""Ingest data dalam batch untuk reliabilitas lebih baik."""

for i in range(0, len(df), batchsize):

batch = df.iloc[i:i+batchsize]

featuregroup.ingest(

dataframe=batch,

maxworkers=4,

wait=True

)

print(f"Batch {i//batch_size + 1} di-ingest")

Kesimpulan

SageMaker Feature Store menyediakan:

  • Penyimpanan terpusat: Satu sumber kebenaran untuk fitur
  • Dual stores: Online (real-time) dan offline (training)
  • Konsistensi: Fitur sama di training dan inference
  • Reusabilitas: Berbagi fitur antar tim
  • Time travel: Nilai fitur historis
  • Key takeaways:

    • Rancang fitur dengan reusabilitas dalam pikiran
    • Gunakan online store untuk real-time inference
    • Manfaatkan Athena untuk offline training queries
    • Implementasikan manajemen event time yang proper
    • Monitor kesegaran dan kualitas fitur

    Artikel Terkait

    Tutorial Vertex AI Feature Store: Manajemen Feature Terpusat

    Tutorial Lengkap Vertex AI Feature Store: Manajemen Fitur Terpusat Vertex AI Feature Store adalah repositori terpusat un...

    Tutorial AWS SageMaker Model Monitor: Monitoring Model Produksi

    Tutorial Lengkap AWS SageMaker Model Monitor: Monitoring Model ML di Production Amazon SageMaker Model Monitor secara ot...

    Tutorial AWS SageMaker Pipelines: ML Pipeline Automation

    Tutorial Lengkap AWS SageMaker Pipelines: Automasi ML Workflows SageMaker Pipelines adalah layanan CI/CD yang dibuat khu...

    Tutorial Lengkap AWS SageMaker: Machine Learning di Cloud

    Tutorial Lengkap AWS SageMaker: End-to-End ML Pipeline Amazon SageMaker adalah layanan machine learning terkelola penuh ...