Semantic Search Engine from Scratch Tutorial: Embeddings and Vector Search

# Membangun Mesin Pencari Semantik dari Nol ## Daftar Isi 1. [Pendahuluan](#pendahuluan) 2. [Prasyarat](#prasyarat) 3. [Memahami Pencarian Semantik](#memahami-pencarian-semantik) 4. [Text Embedding...

By Ruby Abdullah · · tutorial
Semantic SearchEmbeddingsFAISSVector SearchSentence TransformersFastAPI

Semantic Search Engine from Scratch

Table of Contents

  • Introduction
  • Prerequisites
  • Understanding Semantic Search
  • Text Embeddings with Sentence-Transformers
  • Vector Indexing with FAISS
  • Vector Indexing with Annoy
  • Building the Search Pipeline
  • Filtering and Metadata
  • Reranking for Improved Relevance
  • Hybrid Search: Combining Semantic and Keyword Search
  • Building the API with FastAPI
  • Evaluation Metrics
  • Best Practices
  • Conclusion

  • Introduction

    Traditional keyword-based search engines match documents by exact or fuzzy word matches. Semantic search goes further by understanding the meaning behind queries and documents. When a user searches for "how to fix a broken pipe," a semantic search engine can also return results about "plumbing repair" or "pipe leak solutions" -- even if those exact words are not in the query.

    This tutorial guides you through building a complete semantic search engine from scratch. You will learn how to generate text embeddings, build vector indices with FAISS and Annoy, implement filtering and reranking, combine semantic and keyword search into a hybrid system, expose everything through a FastAPI REST API, and measure search quality with standard evaluation metrics.


    Prerequisites

    • Python 3.9 or higher
    • Basic understanding of machine learning concepts
    • Familiarity with REST APIs

    pip install sentence-transformers faiss-cpu annoy numpy fastapi uvicorn rank-bm25 scikit-learn pydantic
    


    Semantic search works in three stages:

  • Indexing: Documents are converted into dense vector embeddings and stored in a vector index.
  • Querying: The user query is converted into an embedding using the same model.
  • Retrieval: The vector index finds the documents whose embeddings are closest to the query embedding.
  • The key insight is that semantically similar texts produce similar vectors, enabling meaning-based retrieval rather than keyword matching.

    User Query: "affordable electric cars"
    

    |

    v

    [Embedding Model] -> Query Vector [0.12, -0.45, 0.78, ...]

    |

    v

    [Vector Index] -> Nearest Neighbor Search

    |

    v

    Results:

  • "Budget-friendly EVs for 2025" (similarity: 0.92)
  • "Low-cost electric vehicles comparison" (similarity: 0.89)
  • "Tesla Model 3 pricing guide" (similarity: 0.84)

  • Text Embeddings with Sentence-Transformers

    Sentence-Transformers is a Python library that provides pre-trained models for generating high-quality text embeddings.

    Loading and Using Embedding Models

    from sentencetransformers import SentenceTransformer
    

    import numpy as np

    Load a pre-trained model

    'all-MiniLM-L6-v2' is a good balance of speed and quality

    model = SentenceTransformer('all-MiniLM-L6-v2')

    Generate embeddings for single texts

    text = "Machine learning is a subset of artificial intelligence."

    embedding = model.encode(text)

    print(f"Embedding shape: {embedding.shape}") # (384,)

    print(f"Embedding dtype: {embedding.dtype}") # float32

    Generate embeddings for multiple texts (batched for efficiency)

    documents = [

    "Python is a versatile programming language.",

    "Deep learning uses neural networks with many layers.",

    "FastAPI is a modern web framework for Python.",

    "Natural language processing deals with text understanding.",

    "Docker containers simplify application deployment.",

    ]

    docembeddings = model.encode(documents, showprogressbar=True, batchsize=32)

    print(f"Batch embeddings shape: {docembeddings.shape}") # (5, 384)

    Measuring Similarity

    from sentencetransformers.util import cossim
    
    

    Compare two sentences

    sent1 = "The cat sat on the mat."

    sent2 = "A feline rested on the rug."

    sent3 = "The stock market crashed yesterday."

    emb1 = model.encode(sent1)

    emb2 = model.encode(sent2)

    emb3 = model.encode(sent3)

    Cosine similarity

    sim12 = cossim(emb1, emb2).item()

    sim13 = cossim(emb1, emb3).item()

    print(f"'{sent1}' vs '{sent2}': {sim12:.4f}") # High similarity (~0.7+)

    print(f"'{sent1}' vs '{sent3}': {sim13:.4f}") # Low similarity (~0.1)

    Choosing the Right Model

    # Model comparison for different use cases
    

    MODELS = {

    "fastgeneral": "all-MiniLM-L6-v2", # 384 dims, fast, good quality

    "highquality": "all-mpnet-base-v2", # 768 dims, slower, best quality

    "multilingual": "paraphrase-multilingual-MiniLM-L12-v2", # 384 dims, 50+ languages

    "asymmetric": "msmarco-distilbert-base-v4", # 768 dims, optimized for search

    }

    def benchmarkmodels(query: str, documents: list[str]):

    """Compare retrieval quality across different models."""

    results = {}

    for name, modelname in MODELS.items():

    m = SentenceTransformer(modelname)

    qemb = m.encode(query)

    dembs = m.encode(documents)

    similarities = cossim(qemb, dembs)[0].tolist()

    ranked = sorted(zip(documents, similarities), key=lambda x: x[1], reverse=True)

    results[name] = ranked[:3]

    return results


    Vector Indexing with FAISS

    FAISS (Facebook AI Similarity Search) is the industry standard for efficient similarity search over large collections of vectors.

    Building a FAISS Index

    import faiss
    

    import numpy as np

    class FAISSIndex:

    """FAISS-based vector index with support for different index types."""

    def init(self, dimension: int, indextype: str = "flat"):

    self.dimension = dimension

    self.indextype = indextype

    self.index = self.createindex()

    self.documents = []

    self.metadata = []

    def createindex(self) -> faiss.Index:

    """Create the appropriate FAISS index type."""

    if self.indextype == "flat":

    # Exact search (brute force) - best accuracy, slowest for large datasets

    return faiss.IndexFlatIP(self.dimension) # Inner Product (cosine sim for normalized vectors)

    elif self.indextype == "ivf":

    # Inverted file index - good balance of speed and accuracy

    nlist = 100 # Number of clusters

    quantizer = faiss.IndexFlatIP(self.dimension)

    index = faiss.IndexIVFFlat(quantizer, self.dimension, nlist, faiss.METRICINNERPRODUCT)

    return index

    elif self.indextype == "hnsw":

    # Hierarchical Navigable Small World - fast approximate search

    index = faiss.IndexHNSWFlat(self.dimension, 32) # 32 connections per node

    index.hnsw.efConstruction = 200

    index.hnsw.efSearch = 64

    return index

    else:

    raise ValueError(f"Unknown index type: {self.indextype}")

    def add(self, embeddings: np.ndarray, documents: list[str], metadata: list[dict] = None):

    """Add documents and their embeddings to the index."""

    # Normalize embeddings for cosine similarity

    faiss.normalizeL2(embeddings)

    # Train the index if needed (for IVF)

    if self.indextype == "ivf" and not self.index.istrained:

    self.index.train(embeddings)

    self.index.add(embeddings)

    self.documents.extend(documents)

    if metadata:

    self.metadata.extend(metadata)

    else:

    self.metadata.extend([{}] len(documents))

    def search(self, queryembedding: np.ndarray, k: int = 10) -> list[dict]:

    """Search for the k most similar documents."""

    # Normalize query

    querynorm = queryembedding.copy().reshape(1, -1)

    faiss.normalizeL2(querynorm)

    scores, indices = self.index.search(querynorm, k)

    results = []

    for score, idx in zip(scores[0], indices[0]):

    if idx == -1: # FAISS returns -1 for missing results

    continue

    results.append({

    "document": self.documents[idx],

    "score": float(score),

    "metadata": self.metadata[idx],

    "index": int(idx),

    })

    return results

    def save(self, path: str):

    """Save the index to disk."""

    faiss.writeindex(self.index, f"{path}.faiss")

    import json

    with open(f"{path}.meta", "w") as f:

    json.dump({"documents": self.documents, "metadata": self.metadata}, f)

    def load(self, path: str):

    """Load the index from disk."""

    self.index = faiss.readindex(f"{path}.faiss")

    import json

    with open(f"{path}.meta", "r") as f:

    data = json.load(f)

    self.documents = data["documents"]

    self.metadata = data["metadata"]

    Usage

    model = SentenceTransformer('all-MiniLM-L6-v2')

    documents = [

    "Python is great for data science and machine learning.",

    "JavaScript is the language of the web.",

    "Docker helps with containerization and deployment.",

    "Kubernetes orchestrates container workloads.",

    "PostgreSQL is a powerful relational database.",

    "Redis is an in-memory key-value store for caching.",

    "Git is essential for version control.",

    "CI/CD pipelines automate testing and deployment.",

    "REST APIs are the backbone of modern web services.",

    "GraphQL provides a flexible query language for APIs.",

    ]

    metadata = [

    {"category": "language", "level": "beginner"},

    {"category": "language", "level": "beginner"},

    {"category": "devops", "level": "intermediate"},

    {"category": "devops", "level": "advanced"},

    {"category": "database", "level": "intermediate"},

    {"category": "database", "level": "intermediate"},

    {"category": "tools", "level": "beginner"},

    {"category": "devops", "level": "intermediate"},

    {"category": "api", "level": "beginner"},

    {"category": "api", "level": "intermediate"},

    ]

    embeddings = model.encode(documents, converttonumpy=True)

    index = FAISSIndex(dimension=384, indextype="flat")

    index.add(embeddings, documents, metadata)

    Search

    query = "What database should I use for caching?"

    queryemb = model.encode(query, converttonumpy=True)

    results = index.search(queryemb, k=3)

    for r in results:

    print(f" [{r['score']:.4f}] {r['document']} (category: {r['metadata']['category']})")


    Vector Indexing with Annoy

    Annoy (Approximate Nearest Neighbors Oh Yeah) is a lightweight alternative to FAISS, particularly good for read-heavy workloads with static indices.

    from annoy import AnnoyIndex
    
    

    class AnnoySearchIndex:

    """Annoy-based vector index for fast approximate nearest neighbor search."""

    def init(self, dimension: int, metric: str = "angular", ntrees: int = 50):

    self.dimension = dimension

    self.metric = metric

    self.ntrees = ntrees

    self.index = AnnoyIndex(dimension, metric)

    self.documents = []

    self.metadata = []

    self.isbuilt = False

    def add(self, embeddings: np.ndarray, documents: list[str], metadata: list[dict] = None):

    """Add documents and their embeddings."""

    startidx = len(self.documents)

    for i, emb in enumerate(embeddings):

    self.index.additem(startidx + i, emb)

    self.documents.extend(documents)

    self.metadata.extend(metadata or [{}] len(documents))

    def build(self):

    """Build the index (must be called after adding all items)."""

    self.index.build(self.ntrees)

    self.isbuilt = True

    def search(self, queryembedding: np.ndarray, k: int = 10) -> list[dict]:

    """Search for the k nearest neighbors."""

    if not self.isbuilt:

    raise RuntimeError("Index must be built before searching. Call build() first.")

    indices, distances = self.index.getnnsbyvector(

    queryembedding, k, includedistances=True

    )

    results = []

    for idx, dist in zip(indices, distances):

    # Convert angular distance to cosine similarity

    similarity = 1 - (dist * 2) / 2

    results.append({

    "document": self.documents[idx],

    "score": float(similarity),

    "metadata": self.metadata[idx],

    "index": idx,

    })

    return results

    def save(self, path: str):

    """Save index to disk."""

    self.index.save(f"{path}.annoy")

    import json

    with open(f"{path}.meta", "w") as f:

    json.dump({"documents": self.documents, "metadata": self.metadata}, f)

    Usage

    annoyindex = AnnoySearchIndex(dimension=384, ntrees=50)

    annoyindex.add(embeddings, documents, metadata)

    annoyindex.build()

    results = annoyindex.search(queryemb, k=3)

    for r in results:

    print(f" [{r['score']:.4f}] {r['document']}")


    Building the Search Pipeline

    Now let us combine everything into a cohesive search pipeline.

    from dataclasses import dataclass, field
    

    from typing import Optional, Callable

    import time

    @dataclass

    class SearchResult:

    document: str

    score: float

    metadata: dict

    rank: int

    @dataclass

    class SearchResponse:

    query: str

    results: list[SearchResult]

    totalresults: int

    searchtimems: float

    class SemanticSearchEngine:

    """Complete semantic search engine with embedding, indexing, and retrieval."""

    def init(self, modelname: str = "all-MiniLM-L6-v2", indextype: str = "flat"):

    self.model = SentenceTransformer(modelname)

    self.dimension = self.model.getsentenceembeddingdimension()

    self.index = FAISSIndex(dimension=self.dimension, indextype=indextype)

    self.documentcount = 0

    def indexdocuments(self, documents: list[str], metadata: list[dict] = None,

    batchsize: int = 64):

    """Index a batch of documents."""

    embeddings = self.model.encode(

    documents,

    converttonumpy=True,

    batchsize=batchsize,

    showprogressbar=True,

    )

    self.index.add(embeddings, documents, metadata)

    self.documentcount += len(documents)

    def search(self, query: str, k: int = 10,

    filterfn: Optional[Callable[[dict], bool]] = None) -> SearchResponse:

    """Search for documents matching the query."""

    start = time.time()

    queryemb = self.model.encode(query, converttonumpy=True)

    # Fetch more results than needed if filtering

    fetchk = k 3 if filterfn else k

    rawresults = self.index.search(queryemb, k=fetchk)

    # Apply filters

    if filterfn:

    rawresults = [r for r in rawresults if filterfn(r["metadata"])]

    # Trim to requested k

    rawresults = rawresults[:k]

    results = [

    SearchResult(

    document=r["document"],

    score=r["score"],

    metadata=r["metadata"],

    rank=i + 1,

    )

    for i, r in enumerate(rawresults)

    ]

    elapsed = (time.time() - start) 1000

    return SearchResponse(

    query=query,

    results=results,

    totalresults=len(results),

    searchtimems=round(elapsed, 2),

    )

    Build the search engine

    engine = SemanticSearchEngine()

    engine.indexdocuments(documents, metadata)

    Simple search

    response = engine.search("containerization tools")

    print(f"Query: {response.query} ({response.searchtimems}ms)")

    for r in response.results[:3]:

    print(f" {r.rank}. [{r.score:.4f}] {r.document}")

    Filtered search - only DevOps documents

    response = engine.search(

    "deployment automation",

    filterfn=lambda m: m.get("category") == "devops"

    )

    print(f"\nFiltered results (devops only):")

    for r in response.results:

    print(f" {r.rank}. [{r.score:.4f}] {r.document}")


    Filtering and Metadata

    Effective metadata filtering is essential for production search systems.

    class MetadataFilter:
    

    """Composable metadata filters for search results."""

    @staticmethod

    def equals(field: str, value) -> Callable[[dict], bool]:

    return lambda m: m.get(field) == value

    @staticmethod

    def inlist(field: str, values: list) -> Callable[[dict], bool]:

    return lambda m: m.get(field) in values

    @staticmethod

    def rangefilter(field: str, minval=None, maxval=None) -> Callable[[dict], bool]:

    def check(m):

    val = m.get(field)

    if val is None:

    return False

    if minval is not None and val < minval:

    return False

    if maxval is not None and val > maxval:

    return False

    return True

    return check

    @staticmethod

    def combineand(filters: Callable[[dict], bool]) -> Callable[[dict], bool]:

    return lambda m: all(f(m) for f in filters)

    @staticmethod

    def combineor(filters: Callable[[dict], bool]) -> Callable[[dict], bool]:

    return lambda m: any(f(m) for f in filters)

    Usage examples

    f = MetadataFilter()

    Single filter

    beginnerfilter = f.equals("level", "beginner")

    Combined filters

    devopsintermediate = f.combineand(

    f.equals("category", "devops"),

    f.inlist("level", ["intermediate", "advanced"])

    )

    response = engine.search("best tools for deployment", filterfn=devopsintermediate)


    Reranking for Improved Relevance

    Initial retrieval is fast but approximate. A reranker can refine the order of results for better relevance.

    from sentencetransformers import CrossEncoder
    
    

    class Reranker:

    """Cross-encoder based reranker for improving search relevance."""

    def init(self, modelname: str = "cross-encoder/ms-marco-MiniLM-L-6-v2"):

    self.model = CrossEncoder(modelname)

    def rerank(self, query: str, results: list[SearchResult], topk: int = None) -> list[SearchResult]:

    """Rerank search results using a cross-encoder model."""

    if not results:

    return results

    # Prepare query-document pairs

    pairs = [(query, r.document) for r in results]

    # Score each pair with the cross-encoder

    scores = self.model.predict(pairs)

    # Combine with original results

    reranked = []

    for result, score in zip(results, scores):

    reranked.append(SearchResult(

    document=result.document,

    score=float(score),

    metadata=result.metadata,

    rank=0, # Will be updated below

    ))

    # Sort by new score and update ranks

    reranked.sort(key=lambda r: r.score, reverse=True)

    for i, r in enumerate(reranked):

    r.rank = i + 1

    if topk:

    reranked = reranked[:topk]

    return reranked

    Usage

    reranker = Reranker()

    Get initial results (retrieve more than needed)

    initial = engine.search("how to deploy a Python application", k=10)

    Rerank the results

    reranked = reranker.rerank("how to deploy a Python application", initial.results, topk=5)

    print("After reranking:")

    for r in reranked:

    print(f" {r.rank}. [{r.score:.4f}] {r.document}")


    Hybrid search combines the strengths of keyword-based retrieval (BM25) with semantic search for the best of both worlds.

    from rankbm25 import BM25Okapi
    

    import re

    class HybridSearchEngine:

    """Combines BM25 keyword search with semantic search."""

    def init(self, modelname: str = "all-MiniLM-L6-v2", alpha: float = 0.5):

    self.semanticengine = SemanticSearchEngine(modelname)

    self.alpha = alpha # Weight for semantic vs keyword (0=keyword only, 1=semantic only)

    self.bm25 = None

    self.tokenizeddocs = []

    self.documents = []

    self.metadata = []

    def tokenize(self, text: str) -> list[str]:

    """Simple tokenization for BM25."""

    return re.findall(r'\w+', text.lower())

    def indexdocuments(self, documents: list[str], metadata: list[dict] = None):

    """Index documents for both semantic and keyword search."""

    self.documents = documents

    self.metadata = metadata or [{}] len(documents)

    # Semantic indexing

    self.semanticengine.indexdocuments(documents, metadata)

    # BM25 indexing

    self.tokenizeddocs = [self.tokenize(doc) for doc in documents]

    self.bm25 = BM25Okapi(self.tokenizeddocs)

    def search(self, query: str, k: int = 10,

    filterfn: Optional[Callable[[dict], bool]] = None) -> SearchResponse:

    """Hybrid search combining semantic and keyword results."""

    start = time.time()

    # Semantic search

    semanticresults = self.semanticengine.search(query, k=k 2, filterfn=filterfn)

    # BM25 keyword search

    tokenizedquery = self.tokenize(query)

    bm25scores = self.bm25.getscores(tokenizedquery)

    # Normalize scores to [0, 1]

    semanticscores = {}

    for r in semanticresults.results:

    semanticscores[r.document] = r.score

    maxbm25 = max(bm25scores) if max(bm25scores) > 0 else 1

    bm25normalized = {self.documents[i]: s / maxbm25 for i, s in enumerate(bm25scores)}

    # Combine scores using weighted fusion

    combined = {}

    alldocs = set(semanticscores.keys()) | set(d for d, s in bm25normalized.items() if s > 0)

    for doc in alldocs:

    semscore = semanticscores.get(doc, 0)

    kwscore = bm25normalized.get(doc, 0)

    combined[doc] = self.alpha semscore + (1 - self.alpha) kwscore

    # Sort and build results

    sorteddocs = sorted(combined.items(), key=lambda x: x[1], reverse=True)[:k]

    results = []

    for rank, (doc, score) in enumerate(sorteddocs, 1):

    idx = self.documents.index(doc)

    meta = self.metadata[idx] if filterfn is None or filterfn(self.metadata[idx]) else {}

    results.append(SearchResult(document=doc, score=score, metadata=meta, rank=rank))

    elapsed = (time.time() - start) 1000

    return SearchResponse(

    query=query,

    results=results,

    totalresults=len(results),

    searchtimems=round(elapsed, 2),

    )

    Usage

    hybrid = HybridSearchEngine(alpha=0.6) # 60% semantic, 40% keyword

    hybrid.indexdocuments(documents, metadata)

    response = hybrid.search("PostgreSQL database performance")

    for r in response.results[:5]:

    print(f" {r.rank}. [{r.score:.4f}] {r.document}")


    Building the API with FastAPI

    from fastapi import FastAPI, HTTPException, Query
    

    from pydantic import BaseModel, Field

    from typing import Optional

    import uvicorn

    app = FastAPI(title="Semantic Search API", version="1.0.0")

    Initialize the search engine at startup

    searchengine = HybridSearchEngine(alpha=0.6)

    class DocumentInput(BaseModel):

    text: str = Field(description="Document text to index")

    metadata: dict = Field(defaultfactory=dict, description="Optional metadata")

    class BulkIndexRequest(BaseModel):

    documents: list[DocumentInput]

    class SearchRequest(BaseModel):

    query: str = Field(description="Search query")

    k: int = Field(default=10, ge=1, le=100, description="Number of results")

    alpha: Optional[float] = Field(default=None, ge=0, le=1, description="Semantic weight")

    category: Optional[str] = Field(default=None, description="Filter by category")

    class SearchResultResponse(BaseModel):

    document: str

    score: float

    metadata: dict

    rank: int

    class SearchResponseModel(BaseModel):

    query: str

    results: list[SearchResultResponse]

    totalresults: int

    searchtimems: float

    @app.post("/index", statuscode=201)

    async def indexdocuments(request: BulkIndexRequest):

    """Index a batch of documents."""

    texts = [d.text for d in request.documents]

    meta = [d.metadata for d in request.documents]

    searchengine.indexdocuments(texts, meta)

    return {"indexed": len(texts), "total": searchengine.semanticengine.documentcount}

    @app.post("/search", responsemodel=SearchResponseModel)

    async def search(request: SearchRequest):

    """Search for documents matching the query."""

    if searchengine.semanticengine.documentcount == 0:

    raise HTTPException(statuscode=400, detail="No documents indexed")

    filterfn = None

    if request.category:

    filterfn = lambda m: m.get("category") == request.category

    if request.alpha is not None:

    searchengine.alpha = request.alpha

    response = searchengine.search(request.query, k=request.k, filterfn=filterfn)

    return SearchResponseModel(

    query=response.query,

    results=[

    SearchResultResponse(

    document=r.document, score=r.score,

    metadata=r.metadata, rank=r.rank

    )

    for r in response.results

    ],

    totalresults=response.totalresults,

    searchtimems=response.searchtimems,

    )

    @app.get("/health")

    async def healthcheck():

    return {"status": "healthy", "documentsindexed": searchengine.semanticengine.documentcount}

    if name == "main":

    uvicorn.run(app, host="0.0.0.0", port=8000)


    Evaluation Metrics

    Measuring search quality is critical for iterating and improving your system.

    from sklearn.metrics import ndcgscore
    

    import numpy as np

    class SearchEvaluator:

    """Evaluate search engine quality with standard IR metrics."""

    @staticmethod

    def precisionatk(relevant: set[str], retrieved: list[str], k: int) -> float:

    """Proportion of retrieved documents that are relevant."""

    retrievedk = retrieved[:k]

    relevantretrieved = sum(1 for doc in retrievedk if doc in relevant)

    return relevantretrieved / k if k > 0 else 0.0

    @staticmethod

    def recallatk(relevant: set[str], retrieved: list[str], k: int) -> float:

    """Proportion of relevant documents that are retrieved."""

    retrievedk = set(retrieved[:k])

    relevantretrieved = len(relevant & retrievedk)

    return relevantretrieved / len(relevant) if relevant else 0.0

    @staticmethod

    def meanreciprocalrank(relevant: set[str], retrieved: list[str]) -> float:

    """Reciprocal of the rank of the first relevant document."""

    for i, doc in enumerate(retrieved, 1):

    if doc in relevant:

    return 1.0 / i

    return 0.0

    @staticmethod

    def averageprecision(relevant: set[str], retrieved: list[str]) -> float:

    """Average of precision values at each relevant document position."""

    precisions = []

    relevantcount = 0

    for i, doc in enumerate(retrieved, 1):

    if doc in relevant:

    relevantcount += 1

    precisions.append(relevantcount / i)

    return sum(precisions) / len(relevant) if relevant else 0.0

    def evaluate(self, engine, testqueries: list[dict], k: int = 10) -> dict:

    """Run full evaluation on a set of test queries.

    testqueries: list of {"query": str, "relevantdocs": set[str]}

    """

    metrics = {"precision@k": [], "recall@k": [], "mrr": [], "map": []}

    for tq in testqueries:

    response = engine.search(tq["query"], k=k)

    retrieved = [r.document for r in response.results]

    relevant = tq["relevantdocs"]

    metrics["precision@k"].append(self.precisionatk(relevant, retrieved, k))

    metrics["recall@k"].append(self.recallatk(relevant, retrieved, k))

    metrics["mrr"].append(self.meanreciprocalrank(relevant, retrieved))

    metrics["map"].append(self.averageprecision(relevant, retrieved))

    return {

    name: {"mean": np.mean(values), "std": np.std(values)}

    for name, values in metrics.items()

    }

    Example evaluation

    evaluator = SearchEvaluator()

    testqueries = [

    {

    "query": "programming languages",

    "relevantdocs": {

    "Python is great for data science and machine learning.",

    "JavaScript is the language of the web.",

    }

    },

    {

    "query": "container orchestration",

    "relevantdocs": {

    "Docker helps with containerization and deployment.",

    "Kubernetes orchestrates container workloads.",

    }

    },

    ]

    results = evaluator.evaluate(hybrid, testqueries, k=5)

    for metric, values in results.items():

    print(f" {metric}: {values['mean']:.4f} (+/- {values['std']:.4f})")


    Best Practices

  • Choose the right embedding model for your domain. Fine-tuned models outperform general ones.
  • Normalize embeddings before indexing for consistent cosine similarity computation.
  • Use batch encoding for indexing -- it is significantly faster than encoding one-by-one.
  • Implement hybrid search from the start. Pure semantic search misses exact keyword matches.
  • Add a reranker for top results. The small latency cost greatly improves relevance.
  • Pre-filter before vector search when possible to reduce the search space.
  • Chunk long documents into paragraphs or sentences for more precise retrieval.
  • Cache embeddings for frequently repeated queries.
  • Monitor search quality continuously with evaluation metrics and user feedback.
  • Use HNSW or IVF indices for datasets larger than 100K documents.

  • Conclusion

    Building a semantic search engine involves multiple components working together: embedding models for text representation, vector indices for efficient retrieval, rerankers for precision, and hybrid approaches for comprehensive coverage.

    Key takeaways:

    • Sentence-Transformers provides excellent pre-trained models for text embeddings.
    • FAISS and Annoy offer different tradeoffs for vector indexing (accuracy vs. speed vs. memory).
    • Hybrid search combining BM25 and semantic retrieval outperforms either approach alone.
    • Cross-encoder reranking significantly improves result relevance for top positions.
    • Systematic evaluation with standard IR metrics is essential for measuring and improving quality.

    With these building blocks, you can create a search system that truly understands user intent and delivers relevant results at scale.

    Related Articles

    Sentence Transformers Tutorial: Embeddings, Similarity, and Rerankers

    Sentence Transformers: Embedding, Kemiripan Semantik, dan Reranker Sentence Transformers (sering disebut SBERT) adalah p...

    FAISS Tutorial: Efficient Vector Similarity Search at Scale

    FAISS: Pencarian Kemiripan Vektor yang Efisien dalam Skala Besar FAISS (Facebook AI Similarity Search) adalah library C+...

    Complete txtai Tutorial: All-in-One Embeddings Database for Semantic Search and LLM Workflows

    Tutorial Lengkap txtai: Database Embeddings All-in-One untuk Semantic Search dan LLM Workflows txtai adalah framework Py...

    DINOv2: A Complete Guide to Meta AI's Vision Foundation Model for Label-Free Image Embeddings

    DINOv2: Panduan Lengkap Vision Foundation Model dari Meta AI untuk Embedding Gambar Tanpa Label Halo temen-temen, di tut...