Weaviate: Database Vektor dengan AI Modules Terintegrasi
Weaviate adalah database vektor open-source yang dirancang untuk menyimpan objek data beserta vektor embeddingnya. Yang membuat Weaviate unik adalah kemampuannya mengintegrasikan AI modules langsung ke dalam database, memungkinkan auto-vectorization, pencarian semantik, dan bahkan generative AI tanpa infrastruktur tambahan.
Dalam tutorial ini, kita akan mempelajari cara menggunakan Weaviate mulai dari instalasi, definisi schema, pencarian vektor dan hybrid, hingga membangun mesin pencari produk semantik dengan auto-vectorization dan jawaban generatif.
Mengapa Weaviate?
Weaviate menawarkan beberapa keunggulan dibanding database vektor lainnya:
- AI Modules Terintegrasi: Vectorizer dan generative modules langsung di dalam database
- Auto-Vectorization: Data otomatis di-vektorisasi saat dimasukkan tanpa preprocessing manual
- Hybrid Search: Kombinasi pencarian vektor (semantic) dan keyword (BM25)
- GraphQL API: Interface query yang fleksibel dan powerful
- Multi-Tenancy: Isolasi data untuk aplikasi multi-tenant
- Skalabilitas: Mendukung horizontal scaling untuk dataset besar
- Ekosistem Integrasi: Kompatibel dengan LangChain, LlamaIndex, dan framework AI lainnya
Instalasi
Menggunakan Docker (Rekomendasi untuk Development)
Buat file docker-compose.yml:
version: '3.4'
services:
weaviate:
image: cr.weaviate.io/semitechnologies/weaviate:1.25.0
restart: on-failure:0
ports:
- "8080:8080"
- "50051:50051"
environment:
QUERYDEFAULTSLIMIT: 25
AUTHENTICATIONANONYMOUSACCESSENABLED: 'true'
PERSISTENCEDATAPATH: '/var/lib/weaviate'
DEFAULTVECTORIZERMODULE: 'text2vec-openai'
ENABLEMODULES: 'text2vec-openai,generative-openai'
OPENAIAPIKEY: 'sk-your-openai-api-key'
CLUSTERHOSTNAME: 'node1'
volumes:
- weaviatedata:/var/lib/weaviate
volumes:
weaviatedata:
Jalankan Weaviate:
docker-compose up -d
Menggunakan Weaviate Cloud (WCD)
Untuk deployment produksi, Anda bisa menggunakan Weaviate Cloud:
Instalasi Python Client
pip install weaviate-client
Koneksi ke Weaviate
import weaviate
from weaviate.classes.init import Auth
Koneksi ke instance lokal (Docker)
client = weaviate.connecttolocal()
Koneksi ke Weaviate Cloud
client = weaviate.connecttoweaviatecloud(
clusterurl="https://your-cluster.weaviate.network",
authcredentials=Auth.apikey("your-wcd-api-key"),
headers={
"X-OpenAI-Api-Key": "sk-your-openai-api-key"
}
)
Verifikasi koneksi
print(client.isready()) # True jika berhasil
Schema Definition
Schema di Weaviate mendefinisikan struktur data, termasuk properti dan konfigurasi vectorizer.
Membuat Collection (Class)
import weaviate
import weaviate.classes.config as wc
client = weaviate.connecttolocal()
Membuat collection sederhana
client.collections.create(
name="Article",
description="Koleksi artikel blog",
vectorizerconfig=wc.Configure.Vectorizer.text2vecopenai(
model="text-embedding-3-small",
),
generativeconfig=wc.Configure.Generative.openai(
model="gpt-4",
),
properties=[
wc.Property(
name="title",
datatype=wc.DataType.TEXT,
description="Judul artikel",
),
wc.Property(
name="content",
datatype=wc.DataType.TEXT,
description="Isi artikel",
),
wc.Property(
name="author",
datatype=wc.DataType.TEXT,
description="Nama penulis",
skipvectorization=True, # Tidak di-vektorisasi
),
wc.Property(
name="publisheddate",
datatype=wc.DataType.DATE,
description="Tanggal publikasi",
skipvectorization=True,
),
wc.Property(
name="tags",
datatype=wc.DataType.TEXTARRAY,
description="Tag artikel",
),
wc.Property(
name="viewcount",
datatype=wc.DataType.INT,
description="Jumlah views",
skipvectorization=True,
),
],
)
print("Collection 'Article' berhasil dibuat!")
Konfigurasi Vectorizer Alternatif
# Menggunakan text2vec-transformers (self-hosted)
client.collections.create(
name="Document",
vectorizerconfig=wc.Configure.Vectorizer.text2vectransformers(),
properties=[
wc.Property(name="text", datatype=wc.DataType.TEXT),
wc.Property(name="source", datatype=wc.DataType.TEXT),
],
)
Menggunakan text2vec-cohere
client.collections.create(
name="SearchIndex",
vectorizerconfig=wc.Configure.Vectorizer.text2veccohere(
model="embed-multilingual-v3.0",
),
properties=[
wc.Property(name="content", datatype=wc.DataType.TEXT),
wc.Property(name="language", datatype=wc.DataType.TEXT),
],
)
Melihat dan Menghapus Collection
# Melihat semua collections
collections = client.collections.listall()
for name, config in collections.items():
print(f"Collection: {name}")
Menghapus collection
client.collections.delete("Article")
Data Import (Batch)
Weaviate mendukung batch import untuk memasukkan data secara efisien.
Import Data Satu per Satu
articles = client.collections.get("Article")
Insert satu objek
articleuuid = articles.data.insert(
properties={
"title": "Pengenalan Machine Learning",
"content": "Machine learning adalah cabang AI yang memungkinkan komputer belajar dari data...",
"author": "Ruby Abdullah",
"publisheddate": "2024-01-15T00:00:00Z",
"tags": ["machine-learning", "ai", "tutorial"],
"viewcount": 1500,
}
)
print(f"Inserted with UUID: {articleuuid}")
Batch Import
articles = client.collections.get("Article")
Data sampel
sampledata = [
{
"title": "Deep Learning dengan PyTorch",
"content": "PyTorch adalah framework deep learning yang populer dikembangkan oleh Meta AI...",
"author": "Ruby Abdullah",
"publisheddate": "2024-02-10T00:00:00Z",
"tags": ["deep-learning", "pytorch", "tutorial"],
"viewcount": 2300,
},
{
"title": "Natural Language Processing untuk Pemula",
"content": "NLP adalah bidang AI yang berfokus pada interaksi antara komputer dan bahasa manusia...",
"author": "Ruby Abdullah",
"publisheddate": "2024-03-05T00:00:00Z",
"tags": ["nlp", "ai", "beginner"],
"viewcount": 1800,
},
{
"title": "Computer Vision dengan OpenCV",
"content": "OpenCV adalah library open-source untuk computer vision dan image processing...",
"author": "Ruby Abdullah",
"publisheddate": "2024-04-20T00:00:00Z",
"tags": ["computer-vision", "opencv", "tutorial"],
"viewcount": 2100,
},
{
"title": "Reinforcement Learning Dasar",
"content": "Reinforcement learning adalah paradigma machine learning di mana agen belajar melalui trial and error...",
"author": "Ruby Abdullah",
"publisheddate": "2024-05-15T00:00:00Z",
"tags": ["reinforcement-learning", "ai", "tutorial"],
"viewcount": 950,
},
]
Batch insert
with articles.batch.dynamic() as batch:
for item in sampledata:
batch.addobject(properties=item)
Cek hasil batch
failed = articles.batch.failedobjects
if failed:
print(f"Failed objects: {len(failed)}")
for obj in failed:
print(f" Error: {obj.message}")
else:
print(f"Semua {len(sampledata)} objek berhasil dimasukkan!")
Import dengan Custom Vector
import numpy as np
articles = client.collections.get("Article")
Insert dengan vector custom
articles.data.insert(
properties={
"title": "Custom Vector Article",
"content": "Artikel dengan vector yang disediakan manual...",
"author": "Ruby Abdullah",
},
vector=np.random.rand(1536).tolist(), # Custom embedding
)
Vector Search
Weaviate menyediakan beberapa metode pencarian vektor.
nearText Search
Pencarian berdasarkan kesamaan semantik dengan teks query.
articles = client.collections.get("Article")
Pencarian semantik
response = articles.query.neartext(
query="tutorial belajar kecerdasan buatan",
limit=3,
returnmetadata=wc.query.MetadataQuery(
distance=True,
certainty=True,
),
)
for obj in response.objects:
print(f"Title: {obj.properties['title']}")
print(f"Distance: {obj.metadata.distance:.4f}")
print(f"Certainty: {obj.metadata.certainty:.4f}")
print()
nearVector Search
Pencarian berdasarkan vektor embedding yang sudah ada.
import numpy as np
articles = client.collections.get("Article")
Generate atau ambil vector dari sumber lain
queryvector = np.random.rand(1536).tolist()
response = articles.query.nearvector(
nearvector=queryvector,
limit=5,
returnmetadata=wc.query.MetadataQuery(distance=True),
)
for obj in response.objects:
print(f"Title: {obj.properties['title']}")
print(f"Distance: {obj.metadata.distance:.4f}")
nearObject Search
Pencarian objek yang mirip dengan objek lain yang sudah ada di database.
articles = client.collections.get("Article")
Cari objek yang mirip dengan objek tertentu
response = articles.query.nearobject(
nearobject=targetuuid, # UUID objek referensi
limit=5,
returnmetadata=wc.query.MetadataQuery(distance=True),
)
for obj in response.objects:
print(f"Title: {obj.properties['title']}")
print(f"Distance: {obj.metadata.distance:.4f}")
Hybrid Search
Hybrid search menggabungkan pencarian vektor (semantic) dan keyword (BM25) untuk hasil yang lebih akurat.
import weaviate.classes.query as wq
articles = client.collections.get("Article")
Hybrid search
response = articles.query.hybrid(
query="deep learning framework python",
alpha=0.5, # 0 = murni BM25, 1 = murni vector search
limit=5,
returnmetadata=wq.MetadataQuery(score=True, explainscore=True),
)
for obj in response.objects:
print(f"Title: {obj.properties['title']}")
print(f"Score: {obj.metadata.score:.4f}")
print(f"Explain: {obj.metadata.explainscore}")
print()
Mengatur Bobot Alpha
# Lebih menekankan keyword matching
responsekeyword = articles.query.hybrid(
query="PyTorch tutorial",
alpha=0.25, # 75% BM25, 25% vector
limit=5,
)
Lebih menekankan semantic similarity
responsesemantic = articles.query.hybrid(
query="cara membuat model AI",
alpha=0.75, # 25% BM25, 75% vector
limit=5,
)
Filters
Weaviate mendukung berbagai filter untuk mempersempit hasil pencarian.
import weaviate.classes.query as wq
articles = client.collections.get("Article")
Filter berdasarkan properti
response = articles.query.neartext(
query="tutorial AI",
limit=10,
filters=wq.Filter.byproperty("author").equal("Ruby Abdullah"),
)
Filter dengan operator perbandingan
response = articles.query.neartext(
query="machine learning",
limit=10,
filters=wq.Filter.byproperty("viewcount").greaterthan(1000),
)
Filter dengan AND
response = articles.query.neartext(
query="deep learning",
limit=10,
filters=(
wq.Filter.byproperty("author").equal("Ruby Abdullah") &
wq.Filter.byproperty("viewcount").greaterthan(1000)
),
)
Filter dengan OR
response = articles.query.neartext(
query="programming",
limit=10,
filters=(
wq.Filter.byproperty("tags").containsany(["python", "tutorial"]) |
wq.Filter.byproperty("viewcount").greaterthan(2000)
),
)
Filter berdasarkan tanggal
from datetime import datetime
response = articles.query.neartext(
query="AI tutorial",
limit=10,
filters=wq.Filter.byproperty("publisheddate").greaterthan(
datetime(2024, 3, 1)
),
)
Generative Modules
Generative modules memungkinkan Anda menggunakan LLM langsung di dalam query Weaviate.
Single Prompt (Per Objek)
articles = client.collections.get("Article")
response = articles.generate.neartext(
query="machine learning untuk pemula",
limit=3,
singleprompt="Buatkan ringkasan singkat (2-3 kalimat) dari artikel berikut: {title} - {content}",
)
for obj in response.objects:
print(f"Title: {obj.properties['title']}")
print(f"Summary: {obj.generated}")
print()
Grouped Task (Semua Objek)
articles = client.collections.get("Article")
response = articles.generate.neartext(
query="tutorial programming",
limit=5,
groupedtask="Berdasarkan artikel-artikel berikut, buatkan rekomendasi learning path untuk seorang pemula yang ingin belajar AI. Urutkan dari yang paling dasar.",
)
Hasil generatif dari semua objek yang ditemukan
print("Learning Path Recommendation:")
print(response.generated)
Generative dengan Hybrid Search
articles = client.collections.get("Article")
response = articles.generate.hybrid(
query="python data science",
alpha=0.5,
limit=3,
singleprompt="Jelaskan mengapa artikel '{title}' relevan untuk data scientist pemula.",
groupedtask="Bandingkan ketiga artikel di atas dan tentukan mana yang paling cocok untuk pemula.",
)
for obj in response.objects:
print(f"Title: {obj.properties['title']}")
print(f"Relevance: {obj.generated}")
print()
print(f"\nComparison: {response.generated}")
Multi-Tenancy
Multi-tenancy memungkinkan isolasi data per tenant dalam satu collection.
Mengaktifkan Multi-Tenancy
import weaviate.classes.config as wc
Buat collection dengan multi-tenancy
client.collections.create(
name="CustomerData",
multitenancyconfig=wc.Configure.multitenancy(
enabled=True,
autotenantcreation=True,
),
vectorizerconfig=wc.Configure.Vectorizer.text2vecopenai(),
properties=[
wc.Property(name="name", datatype=wc.DataType.TEXT),
wc.Property(name="description", datatype=wc.DataType.TEXT),
wc.Property(name="category", datatype=wc.DataType.TEXT),
],
)
Mengelola Tenants
from weaviate.classes.tenants import Tenant, TenantActivityStatus
collection = client.collections.get("CustomerData")
Menambah tenants
collection.tenants.create([
Tenant(name="companya"),
Tenant(name="companyb"),
Tenant(name="companyc"),
])
Melihat semua tenants
tenants = collection.tenants.get()
for name, tenant in tenants.items():
print(f"Tenant: {name}, Status: {tenant.activitystatus}")
Operasi Data per Tenant
# Akses data untuk tenant tertentu
tenanta = client.collections.get("CustomerData").withtenant("companya")
Insert data untuk tenant A
tenanta.data.insert(
properties={
"name": "Product X",
"description": "Produk premium untuk Company A",
"category": "premium",
}
)
Query data hanya untuk tenant A
response = tenanta.query.neartext(
query="produk premium",
limit=5,
)
Data tenant B tidak akan muncul di query tenant A
tenantb = client.collections.get("CustomerData").withtenant("companyb")
tenantb.data.insert(
properties={
"name": "Product Y",
"description": "Produk basic untuk Company B",
"category": "basic",
}
)
Backup dan Restore
Membuat Backup
# Backup semua collections
result = client.backup.create(
backupid="backup-2024-01-15",
backend="filesystem",
waitforcompletion=True,
)
print(f"Backup status: {result.status}")
Backup collections tertentu
result = client.backup.create(
backupid="backup-articles-only",
backend="filesystem",
includecollections=["Article"],
waitforcompletion=True,
)
Restore dari Backup
# Restore semua collections
result = client.backup.restore(
backupid="backup-2024-01-15",
backend="filesystem",
waitforcompletion=True,
)
print(f"Restore status: {result.status}")
Restore collections tertentu
result = client.backup.restore(
backupid="backup-articles-only",
backend="filesystem",
includecollections=["Article"],
waitforcompletion=True,
)
Integrasi dengan LangChain
from langchainweaviate import WeaviateVectorStore
from langchainopenai import OpenAIEmbeddings
import weaviate
Koneksi ke Weaviate
client = weaviate.connecttolocal()
Buat vector store
embeddings = OpenAIEmbeddings(model="text-embedding-3-small")
vectorstore = WeaviateVectorStore(
client=client,
indexname="LangChainDocs",
textkey="content",
embedding=embeddings,
)
Tambah dokumen
from langchain.schema import Document
docs = [
Document(pagecontent="Python adalah bahasa pemrograman yang populer", metadata={"source": "intro"}),
Document(pagecontent="Machine learning menggunakan data untuk membuat prediksi", metadata={"source": "ml"}),
]
vectorstore.adddocuments(docs)
Pencarian
results = vectorstore.similaritysearch(
query="bahasa pemrograman untuk AI",
k=3,
)
for doc in results:
print(f"Content: {doc.pagecontent}")
print(f"Metadata: {doc.metadata}")
print()
Sebagai retriever untuk RAG
from langchainopenai import ChatOpenAI
from langchain.chains import RetrievalQA
llm = ChatOpenAI(model="gpt-4", temperature=0)
retriever = vectorstore.asretriever(searchkwargs={"k": 3})
qachain = RetrievalQA.fromchaintype(
llm=llm,
chaintype="stuff",
retriever=retriever,
)
answer = qachain.invoke("Apa itu machine learning?")
print(answer["result"])
Integrasi dengan LlamaIndex
from llamaindex.core import VectorStoreIndex, StorageContext
from llama
index.vectorstores.weaviate import WeaviateVectorStore
import weaviate
Koneksi
client = weaviate.connect
tolocal()
Buat vector store
vector
store = WeaviateVectorStore(
weaviateclient=client,
indexname="LlamaIndexDocs",
)
Buat storage context
storagecontext = StorageContext.fromdefaults(
vectorstore=vectorstore,
)
Buat index dari dokumen
from llamaindex.core import Document
documents = [
Document(text="Weaviate adalah database vektor yang powerful"),
Document(text="LlamaIndex memudahkan pembuatan aplikasi RAG"),
]
index = VectorStoreIndex.fromdocuments(
documents,
storagecontext=storagecontext,
)
Query
queryengine = index.asqueryengine()
response = queryengine.query("Apa itu Weaviate?")
print(response)
Contoh Praktis: Mesin Pencari Produk Semantik
Mari kita bangun mesin pencari produk dengan auto-vectorization dan jawaban generatif.
Setup Collection Produk
import weaviate
import weaviate.classes.config as wc
client = weaviate.connecttolocal()
Hapus collection jika sudah ada
if client.collections.exists("Product"):
client.collections.delete("Product")
Buat collection produk
client.collections.create(
name="Product",
description="Katalog produk e-commerce",
vectorizerconfig=wc.Configure.Vectorizer.text2vecopenai(
model="text-embedding-3-small",
),
generativeconfig=wc.Configure.Generative.openai(
model="gpt-4",
),
properties=[
wc.Property(
name="name",
datatype=wc.DataType.TEXT,
description="Nama produk",
),
wc.Property(
name="description",
datatype=wc.DataType.TEXT,
description="Deskripsi produk",
),
wc.Property(
name="category",
datatype=wc.DataType.TEXT,
description="Kategori produk",
),
wc.Property(
name="price",
datatype=wc.DataType.NUMBER,
description="Harga produk",
skipvectorization=True,
),
wc.Property(
name="brand",
datatype=wc.DataType.TEXT,
description="Merek produk",
),
wc.Property(
name="specs",
datatype=wc.DataType.TEXT,
description="Spesifikasi produk",
),
wc.Property(
name="rating",
datatype=wc.DataType.NUMBER,
description="Rating produk (1-5)",
skipvectorization=True,
),
wc.Property(
name="stock",
datatype=wc.DataType.INT,
description="Stok tersedia",
skipvectorization=True,
),
],
)
print("Collection 'Product' berhasil dibuat!")
Import Data Produk
products = client.collections.get("Product")
sampleproducts = [
{
"name": "Laptop Gaming ProMax X15",
"description": "Laptop gaming performa tinggi dengan layar 15.6 inci 144Hz, cocok untuk gaming AAA dan content creation. Dilengkapi sistem pendingin canggih.",
"category": "Laptop",
"price": 18500000,
"brand": "ProMax",
"specs": "Intel i7-13700H, RTX 4060, 16GB DDR5, 512GB NVMe SSD, 15.6\" FHD 144Hz",
"rating": 4.5,
"stock": 25,
},
{
"name": "Headphone Wireless NoiseBlock Pro",
"description": "Headphone wireless premium dengan Active Noise Cancellation terbaik di kelasnya. Battery tahan hingga 30 jam pemakaian.",
"category": "Audio",
"price": 3200000,
"brand": "NoiseBlock",
"specs": "ANC, Bluetooth 5.3, 30 jam battery, driver 40mm, codec LDAC/AAC",
"rating": 4.7,
"stock": 50,
},
{
"name": "Smartphone UltraVision 5G",
"description": "Smartphone flagship dengan kamera 200MP dan layar AMOLED 6.7 inci. Mendukung 5G untuk konektivitas super cepat.",
"category": "Smartphone",
"price": 12000000,
"brand": "UltraVision",
"specs": "Snapdragon 8 Gen 3, 12GB RAM, 256GB Storage, 200MP Camera, 6.7\" AMOLED 120Hz",
"rating": 4.6,
"stock": 100,
},
{
"name": "Mechanical Keyboard RGB TypeMaster",
"description": "Keyboard mekanikal full-size dengan switch Cherry MX Blue, hot-swappable, dan RGB per-key lighting. Ideal untuk programmer dan gamer.",
"category": "Accessories",
"price": 1500000,
"brand": "TypeMaster",
"specs": "Cherry MX Blue, Hot-swap, RGB per-key, PBT keycaps, USB-C, N-key rollover",
"rating": 4.4,
"stock": 75,
},
{
"name": "Monitor 4K UltraWide ScreenPro",
"description": "Monitor ultrawide 34 inci dengan resolusi 4K untuk produktivitas maksimal. Panel IPS dengan akurasi warna tinggi untuk desainer.",
"category": "Monitor",
"price": 8500000,
"brand": "ScreenPro",
"specs": "34\" IPS UltraWide, 3440x1440, 100% sRGB, USB-C PD 65W, HDR400",
"rating": 4.3,
"stock": 30,
},
{
"name": "Tablet CreativeTab Pro 12",
"description": "Tablet dengan stylus pen untuk digital drawing dan note-taking. Layar 12 inci dengan teknologi paper-like display.",
"category": "Tablet",
"price": 7500000,
"brand": "CreativeTab",
"specs": "12\" Paper-like display, Stylus 4096 levels, 8GB RAM, 128GB, Android 14",
"rating": 4.2,
"stock": 40,
},
]
Batch import
with products.batch.dynamic() as batch:
for product in sampleproducts:
batch.addobject(properties=product)
print(f"Berhasil import {len(sampleproducts)} produk!")
Pencarian Semantik Produk
import weaviate.classes.query as wq
products = client.collections.get("Product")
Pencarian: user mencari dengan bahasa natural
queries = [
"laptop untuk main game berat",
"earphone anti bising untuk kerja dari cafe",
"HP dengan kamera bagus untuk foto",
"keyboard enak untuk coding",
"layar besar untuk desain grafis",
]
for query in queries:
print(f"\nQuery: '{query}'")
print("-" * 50)
response = products.query.neartext(
query=query,
limit=2,
returnmetadata=wq.MetadataQuery(distance=True),
)
for obj in response.objects:
p = obj.properties
print(f" {p['name']} - Rp {p['price']:,.0f}")
print(f" Distance: {obj.metadata.distance:.4f}")
Hybrid Search dengan Filter
products = client.collections.get("Product")
Hybrid search + filter harga
response = products.query.hybrid(
query="perangkat untuk produktivitas kerja remote",
alpha=0.6,
limit=5,
filters=wq.Filter.byproperty("price").lessthan(10000000),
returnmetadata=wq.MetadataQuery(score=True),
)
print("Produk untuk kerja remote (budget < 10 juta):")
for obj in response.objects:
p = obj.properties
print(f" {p['name']} - Rp {p['price']:,.0f} (score: {obj.metadata.score:.4f})")
Rekomendasi Generatif
products = client.collections.get("Product")
Pencarian + rekomendasi personal
response = products.generate.neartext(
query="setup lengkap untuk programmer freelance",
limit=4,
singleprompt="Jelaskan dalam 1-2 kalimat mengapa produk '{name}' cocok untuk programmer freelance. Harga: Rp {price}",
groupedtask="""Berdasarkan produk-produk di atas, buatkan rekomendasi setup lengkap
untuk programmer freelance dengan budget 30 juta.
Sertakan total harga dan alasan pemilihan setiap item.""",
)
print("=== Rekomendasi Per Produk ===\n")
for obj in response.objects:
p = obj.properties
print(f"Produk: {p['name']}")
print(f"Harga: Rp {p['price']:,.0f}")
print(f"Rekomendasi: {obj.generated}")
print()
print("=== Rekomendasi Setup Lengkap ===\n")
print(response.generated)
Kelas Pencarian Produk Lengkap
import weaviate
import weaviate.classes.query as wq
from dataclasses import dataclass
from typing import Optional
@dataclass
class SearchResult:
name: str
description: str
category: str
price: float
brand: str
rating: float
score: float = 0.0
recommendation: str = ""
class ProductSearchEngine:
def init(self, client: weaviate.WeaviateClient):
self.client = client
self.products = client.collections.get("Product")
def semanticsearch(
self,
query: str,
limit: int = 5,
minrating: Optional[float] = None,
maxprice: Optional[float] = None,
category: Optional[str] = None,
) -> list[SearchResult]:
"""Pencarian semantik dengan filter opsional."""
filters = []
if minrating:
filters.append(
wq.Filter.byproperty("rating").greaterorequal(minrating)
)
if maxprice:
filters.append(
wq.Filter.byproperty("price").lessorequal(maxprice)
)
if category:
filters.append(
wq.Filter.byproperty("category").equal(category)
)
combinedfilter = None
if filters:
combinedfilter = filters[0]
for f in filters[1:]:
combinedfilter = combinedfilter & f
response = self.products.query.neartext(
query=query,
limit=limit,
filters=combinedfilter,
returnmetadata=wq.MetadataQuery(distance=True),
)
results = []
for obj in response.objects:
p = obj.properties
results.append(SearchResult(
name=p["name"],
description=p["description"],
category=p["category"],
price=p["price"],
brand=p["brand"],
rating=p["rating"],
score=1 - obj.metadata.distance,
))
return results
def hybridsearch(
self,
query: str,
alpha: float = 0.5,
limit: int = 5,
) -> list[SearchResult]:
"""Hybrid search (semantic + keyword)."""
response = self.products.query.hybrid(
query=query,
alpha=alpha,
limit=limit,
returnmetadata=wq.MetadataQuery(score=True),
)
results = []
for obj in response.objects:
p = obj.properties
results.append(SearchResult(
name=p["name"],
description=p["description"],
category=p["category"],
price=p["price"],
brand=p["brand"],
rating=p["rating"],
score=obj.metadata.score,
))
return results
def askrecommendation(
self,
query: str,
limit: int = 3,
) -> tuple[list[SearchResult], str]:
"""Pencarian dengan rekomendasi generatif."""
response = self.products.generate.neartext(
query=query,
limit=limit,
groupedtask=f"Berdasarkan produk yang ditemukan, berikan rekomendasi terbaik untuk: {query}. Jelaskan alasannya.",
)
results = []
for obj in response.objects:
p = obj.properties
results.append(SearchResult(
name=p["name"],
description=p["description"],
category=p["category"],
price=p["price"],
brand=p["brand"],
rating=p["rating"],
))
return results, response.generated
def findsimilar(self, productuuid: str, limit: int = 3) -> list[SearchResult]:
"""Cari produk serupa."""
response = self.products.query.nearobject(
nearobject=productuuid,
limit=limit,
returnmetadata=wq.MetadataQuery(distance=True),
)
results = []
for obj in response.objects:
p = obj.properties
results.append(SearchResult(
name=p["name"],
description=p["description"],
category=p["category"],
price=p["price"],
brand=p["brand"],
rating=p["rating"],
score=1 - obj.metadata.distance,
))
return results
Penggunaan
if name == "main":
client = weaviate.connecttolocal()
engine = ProductSearchEngine(client)
# Pencarian semantik
print("=== Semantic Search ===")
results = engine.semanticsearch(
query="perangkat untuk gaming",
maxprice=20000000,
minrating=4.0,
)
for r in results:
print(f" {r.name} - Rp {r.price:,.0f} (rating: {r.rating})")
# Hybrid search
print("\n=== Hybrid Search ===")
results = engine.hybridsearch(
query="keyboard mechanical RGB",
alpha=0.3,
)
for r in results:
print(f" {r.name} - Score: {r.score:.4f}")
# Rekomendasi generatif
print("\n=== AI Recommendation ===")
results, recommendation = engine.askrecommendation(