Semantic Search Engine from Scratch
Table of Contents
Introduction
Traditional keyword-based search engines match documents by exact or fuzzy word matches. Semantic search goes further by understanding the meaning behind queries and documents. When a user searches for "how to fix a broken pipe," a semantic search engine can also return results about "plumbing repair" or "pipe leak solutions" -- even if those exact words are not in the query.
This tutorial guides you through building a complete semantic search engine from scratch. You will learn how to generate text embeddings, build vector indices with FAISS and Annoy, implement filtering and reranking, combine semantic and keyword search into a hybrid system, expose everything through a FastAPI REST API, and measure search quality with standard evaluation metrics.
Prerequisites
- Python 3.9 or higher
- Basic understanding of machine learning concepts
- Familiarity with REST APIs
pip install sentence-transformers faiss-cpu annoy numpy fastapi uvicorn rank-bm25 scikit-learn pydantic
Understanding Semantic Search
Semantic search works in three stages:
The key insight is that semantically similar texts produce similar vectors, enabling meaning-based retrieval rather than keyword matching.
User Query: "affordable electric cars"
|
v
[Embedding Model] -> Query Vector [0.12, -0.45, 0.78, ...]
|
v
[Vector Index] -> Nearest Neighbor Search
|
v
Results:
"Budget-friendly EVs for 2025" (similarity: 0.92)
"Low-cost electric vehicles comparison" (similarity: 0.89)
"Tesla Model 3 pricing guide" (similarity: 0.84)
Text Embeddings with Sentence-Transformers
Sentence-Transformers is a Python library that provides pre-trained models for generating high-quality text embeddings.
Loading and Using Embedding Models
from sentencetransformers import SentenceTransformer
import numpy as np
Load a pre-trained model
'all-MiniLM-L6-v2' is a good balance of speed and quality
model = SentenceTransformer('all-MiniLM-L6-v2')
Generate embeddings for single texts
text = "Machine learning is a subset of artificial intelligence."
embedding = model.encode(text)
print(f"Embedding shape: {embedding.shape}") # (384,)
print(f"Embedding dtype: {embedding.dtype}") # float32
Generate embeddings for multiple texts (batched for efficiency)
documents = [
"Python is a versatile programming language.",
"Deep learning uses neural networks with many layers.",
"FastAPI is a modern web framework for Python.",
"Natural language processing deals with text understanding.",
"Docker containers simplify application deployment.",
]
doc
embeddings = model.encode(documents, showprogressbar=True, batchsize=32)
print(f"Batch embeddings shape: {doc
embeddings.shape}") # (5, 384)
Measuring Similarity
from sentencetransformers.util import cossim
Compare two sentences
sent1 = "The cat sat on the mat."
sent2 = "A feline rested on the rug."
sent3 = "The stock market crashed yesterday."
emb1 = model.encode(sent1)
emb2 = model.encode(sent2)
emb3 = model.encode(sent3)
Cosine similarity
sim12 = cossim(emb1, emb2).item()
sim13 = cossim(emb1, emb3).item()
print(f"'{sent1}' vs '{sent2}': {sim12:.4f}") # High similarity (~0.7+)
print(f"'{sent1}' vs '{sent3}': {sim13:.4f}") # Low similarity (~0.1)
Choosing the Right Model
# Model comparison for different use cases
MODELS = {
"fastgeneral": "all-MiniLM-L6-v2", # 384 dims, fast, good quality
"highquality": "all-mpnet-base-v2", # 768 dims, slower, best quality
"multilingual": "paraphrase-multilingual-MiniLM-L12-v2", # 384 dims, 50+ languages
"asymmetric": "msmarco-distilbert-base-v4", # 768 dims, optimized for search
}
def benchmarkmodels(query: str, documents: list[str]):
"""Compare retrieval quality across different models."""
results = {}
for name, modelname in MODELS.items():
m = SentenceTransformer(modelname)
qemb = m.encode(query)
dembs = m.encode(documents)
similarities = cossim(qemb, dembs)[0].tolist()
ranked = sorted(zip(documents, similarities), key=lambda x: x[1], reverse=True)
results[name] = ranked[:3]
return results
Vector Indexing with FAISS
FAISS (Facebook AI Similarity Search) is the industry standard for efficient similarity search over large collections of vectors.
Building a FAISS Index
import faiss
import numpy as np
class FAISSIndex:
"""FAISS-based vector index with support for different index types."""
def init(self, dimension: int, indextype: str = "flat"):
self.dimension = dimension
self.indextype = indextype
self.index = self.createindex()
self.documents = []
self.metadata = []
def createindex(self) -> faiss.Index:
"""Create the appropriate FAISS index type."""
if self.indextype == "flat":
# Exact search (brute force) - best accuracy, slowest for large datasets
return faiss.IndexFlatIP(self.dimension) # Inner Product (cosine sim for normalized vectors)
elif self.indextype == "ivf":
# Inverted file index - good balance of speed and accuracy
nlist = 100 # Number of clusters
quantizer = faiss.IndexFlatIP(self.dimension)
index = faiss.IndexIVFFlat(quantizer, self.dimension, nlist, faiss.METRICINNERPRODUCT)
return index
elif self.indextype == "hnsw":
# Hierarchical Navigable Small World - fast approximate search
index = faiss.IndexHNSWFlat(self.dimension, 32) # 32 connections per node
index.hnsw.efConstruction = 200
index.hnsw.efSearch = 64
return index
else:
raise ValueError(f"Unknown index type: {self.indextype}")
def add(self, embeddings: np.ndarray, documents: list[str], metadata: list[dict] = None):
"""Add documents and their embeddings to the index."""
# Normalize embeddings for cosine similarity
faiss.normalizeL2(embeddings)
# Train the index if needed (for IVF)
if self.indextype == "ivf" and not self.index.istrained:
self.index.train(embeddings)
self.index.add(embeddings)
self.documents.extend(documents)
if metadata:
self.metadata.extend(metadata)
else:
self.metadata.extend([{}] len(documents))
def search(self, queryembedding: np.ndarray, k: int = 10) -> list[dict]:
"""Search for the k most similar documents."""
# Normalize query
querynorm = queryembedding.copy().reshape(1, -1)
faiss.normalizeL2(querynorm)
scores, indices = self.index.search(querynorm, k)
results = []
for score, idx in zip(scores[0], indices[0]):
if idx == -1: # FAISS returns -1 for missing results
continue
results.append({
"document": self.documents[idx],
"score": float(score),
"metadata": self.metadata[idx],
"index": int(idx),
})
return results
def save(self, path: str):
"""Save the index to disk."""
faiss.writeindex(self.index, f"{path}.faiss")
import json
with open(f"{path}.meta", "w") as f:
json.dump({"documents": self.documents, "metadata": self.metadata}, f)
def load(self, path: str):
"""Load the index from disk."""
self.index = faiss.readindex(f"{path}.faiss")
import json
with open(f"{path}.meta", "r") as f:
data = json.load(f)
self.documents = data["documents"]
self.metadata = data["metadata"]
Usage
model = SentenceTransformer('all-MiniLM-L6-v2')
documents = [
"Python is great for data science and machine learning.",
"JavaScript is the language of the web.",
"Docker helps with containerization and deployment.",
"Kubernetes orchestrates container workloads.",
"PostgreSQL is a powerful relational database.",
"Redis is an in-memory key-value store for caching.",
"Git is essential for version control.",
"CI/CD pipelines automate testing and deployment.",
"REST APIs are the backbone of modern web services.",
"GraphQL provides a flexible query language for APIs.",
]
metadata = [
{"category": "language", "level": "beginner"},
{"category": "language", "level": "beginner"},
{"category": "devops", "level": "intermediate"},
{"category": "devops", "level": "advanced"},
{"category": "database", "level": "intermediate"},
{"category": "database", "level": "intermediate"},
{"category": "tools", "level": "beginner"},
{"category": "devops", "level": "intermediate"},
{"category": "api", "level": "beginner"},
{"category": "api", "level": "intermediate"},
]
embeddings = model.encode(documents, converttonumpy=True)
index = FAISSIndex(dimension=384, indextype="flat")
index.add(embeddings, documents, metadata)
Search
query = "What database should I use for caching?"
queryemb = model.encode(query, converttonumpy=True)
results = index.search(queryemb, k=3)
for r in results:
print(f" [{r['score']:.4f}] {r['document']} (category: {r['metadata']['category']})")
Vector Indexing with Annoy
Annoy (Approximate Nearest Neighbors Oh Yeah) is a lightweight alternative to FAISS, particularly good for read-heavy workloads with static indices.
from annoy import AnnoyIndex
class AnnoySearchIndex:
"""Annoy-based vector index for fast approximate nearest neighbor search."""
def init(self, dimension: int, metric: str = "angular", ntrees: int = 50):
self.dimension = dimension
self.metric = metric
self.ntrees = ntrees
self.index = AnnoyIndex(dimension, metric)
self.documents = []
self.metadata = []
self.isbuilt = False
def add(self, embeddings: np.ndarray, documents: list[str], metadata: list[dict] = None):
"""Add documents and their embeddings."""
startidx = len(self.documents)
for i, emb in enumerate(embeddings):
self.index.additem(startidx + i, emb)
self.documents.extend(documents)
self.metadata.extend(metadata or [{}] len(documents))
def build(self):
"""Build the index (must be called after adding all items)."""
self.index.build(self.ntrees)
self.isbuilt = True
def search(self, queryembedding: np.ndarray, k: int = 10) -> list[dict]:
"""Search for the k nearest neighbors."""
if not self.isbuilt:
raise RuntimeError("Index must be built before searching. Call build() first.")
indices, distances = self.index.getnnsbyvector(
queryembedding, k, includedistances=True
)
results = []
for idx, dist in zip(indices, distances):
# Convert angular distance to cosine similarity
similarity = 1 - (dist * 2) / 2
results.append({
"document": self.documents[idx],
"score": float(similarity),
"metadata": self.metadata[idx],
"index": idx,
})
return results
def save(self, path: str):
"""Save index to disk."""
self.index.save(f"{path}.annoy")
import json
with open(f"{path}.meta", "w") as f:
json.dump({"documents": self.documents, "metadata": self.metadata}, f)
Usage
annoyindex = AnnoySearchIndex(dimension=384, ntrees=50)
annoyindex.add(embeddings, documents, metadata)
annoyindex.build()
results = annoyindex.search(queryemb, k=3)
for r in results:
print(f" [{r['score']:.4f}] {r['document']}")
Building the Search Pipeline
Now let us combine everything into a cohesive search pipeline.
from dataclasses import dataclass, field
from typing import Optional, Callable
import time
@dataclass
class SearchResult:
document: str
score: float
metadata: dict
rank: int
@dataclass
class SearchResponse:
query: str
results: list[SearchResult]
totalresults: int
searchtimems: float
class SemanticSearchEngine:
"""Complete semantic search engine with embedding, indexing, and retrieval."""
def init(self, modelname: str = "all-MiniLM-L6-v2", indextype: str = "flat"):
self.model = SentenceTransformer(modelname)
self.dimension = self.model.getsentenceembeddingdimension()
self.index = FAISSIndex(dimension=self.dimension, indextype=indextype)
self.documentcount = 0
def indexdocuments(self, documents: list[str], metadata: list[dict] = None,
batchsize: int = 64):
"""Index a batch of documents."""
embeddings = self.model.encode(
documents,
converttonumpy=True,
batchsize=batchsize,
showprogressbar=True,
)
self.index.add(embeddings, documents, metadata)
self.documentcount += len(documents)
def search(self, query: str, k: int = 10,
filterfn: Optional[Callable[[dict], bool]] = None) -> SearchResponse:
"""Search for documents matching the query."""
start = time.time()
queryemb = self.model.encode(query, converttonumpy=True)
# Fetch more results than needed if filtering
fetchk = k 3 if filterfn else k
rawresults = self.index.search(queryemb, k=fetchk)
# Apply filters
if filterfn:
rawresults = [r for r in rawresults if filterfn(r["metadata"])]
# Trim to requested k
rawresults = rawresults[:k]
results = [
SearchResult(
document=r["document"],
score=r["score"],
metadata=r["metadata"],
rank=i + 1,
)
for i, r in enumerate(rawresults)
]
elapsed = (time.time() - start) 1000
return SearchResponse(
query=query,
results=results,
totalresults=len(results),
searchtimems=round(elapsed, 2),
)
Build the search engine
engine = SemanticSearchEngine()
engine.indexdocuments(documents, metadata)
Simple search
response = engine.search("containerization tools")
print(f"Query: {response.query} ({response.searchtimems}ms)")
for r in response.results[:3]:
print(f" {r.rank}. [{r.score:.4f}] {r.document}")
Filtered search - only DevOps documents
response = engine.search(
"deployment automation",
filterfn=lambda m: m.get("category") == "devops"
)
print(f"\nFiltered results (devops only):")
for r in response.results:
print(f" {r.rank}. [{r.score:.4f}] {r.document}")
Filtering and Metadata
Effective metadata filtering is essential for production search systems.
class MetadataFilter:
"""Composable metadata filters for search results."""
@staticmethod
def equals(field: str, value) -> Callable[[dict], bool]:
return lambda m: m.get(field) == value
@staticmethod
def inlist(field: str, values: list) -> Callable[[dict], bool]:
return lambda m: m.get(field) in values
@staticmethod
def rangefilter(field: str, minval=None, maxval=None) -> Callable[[dict], bool]:
def check(m):
val = m.get(field)
if val is None:
return False
if minval is not None and val < minval:
return False
if maxval is not None and val > maxval:
return False
return True
return check
@staticmethod
def combineand(filters: Callable[[dict], bool]) -> Callable[[dict], bool]:
return lambda m: all(f(m) for f in filters)
@staticmethod
def combineor(filters: Callable[[dict], bool]) -> Callable[[dict], bool]:
return lambda m: any(f(m) for f in filters)
Usage examples
f = MetadataFilter()
Single filter
beginnerfilter = f.equals("level", "beginner")
Combined filters
devopsintermediate = f.combineand(
f.equals("category", "devops"),
f.inlist("level", ["intermediate", "advanced"])
)
response = engine.search("best tools for deployment", filterfn=devopsintermediate)
Reranking for Improved Relevance
Initial retrieval is fast but approximate. A reranker can refine the order of results for better relevance.
from sentencetransformers import CrossEncoder
class Reranker:
"""Cross-encoder based reranker for improving search relevance."""
def init(self, modelname: str = "cross-encoder/ms-marco-MiniLM-L-6-v2"):
self.model = CrossEncoder(modelname)
def rerank(self, query: str, results: list[SearchResult], topk: int = None) -> list[SearchResult]:
"""Rerank search results using a cross-encoder model."""
if not results:
return results
# Prepare query-document pairs
pairs = [(query, r.document) for r in results]
# Score each pair with the cross-encoder
scores = self.model.predict(pairs)
# Combine with original results
reranked = []
for result, score in zip(results, scores):
reranked.append(SearchResult(
document=result.document,
score=float(score),
metadata=result.metadata,
rank=0, # Will be updated below
))
# Sort by new score and update ranks
reranked.sort(key=lambda r: r.score, reverse=True)
for i, r in enumerate(reranked):
r.rank = i + 1
if topk:
reranked = reranked[:topk]
return reranked
Usage
reranker = Reranker()
Get initial results (retrieve more than needed)
initial = engine.search("how to deploy a Python application", k=10)
Rerank the results
reranked = reranker.rerank("how to deploy a Python application", initial.results, topk=5)
print("After reranking:")
for r in reranked:
print(f" {r.rank}. [{r.score:.4f}] {r.document}")
Hybrid Search: Combining Semantic and Keyword Search
Hybrid search combines the strengths of keyword-based retrieval (BM25) with semantic search for the best of both worlds.
from rankbm25 import BM25Okapi
import re
class HybridSearchEngine:
"""Combines BM25 keyword search with semantic search."""
def init(self, model
name: str = "all-MiniLM-L6-v2", alpha: float = 0.5):
self.semanticengine = SemanticSearchEngine(modelname)
self.alpha = alpha # Weight for semantic vs keyword (0=keyword only, 1=semantic only)
self.bm25 = None
self.tokenizeddocs = []
self.documents = []
self.metadata = []
def tokenize(self, text: str) -> list[str]:
"""Simple tokenization for BM25."""
return re.findall(r'\w+', text.lower())
def indexdocuments(self, documents: list[str], metadata: list[dict] = None):
"""Index documents for both semantic and keyword search."""
self.documents = documents
self.metadata = metadata or [{}] len(documents)
# Semantic indexing
self.semanticengine.indexdocuments(documents, metadata)
# BM25 indexing
self.tokenizeddocs = [self.tokenize(doc) for doc in documents]
self.bm25 = BM25Okapi(self.tokenizeddocs)
def search(self, query: str, k: int = 10,
filterfn: Optional[Callable[[dict], bool]] = None) -> SearchResponse:
"""Hybrid search combining semantic and keyword results."""
start = time.time()
# Semantic search
semanticresults = self.semanticengine.search(query, k=k 2, filterfn=filterfn)
# BM25 keyword search
tokenizedquery = self.tokenize(query)
bm25scores = self.bm25.getscores(tokenizedquery)
# Normalize scores to [0, 1]
semanticscores = {}
for r in semanticresults.results:
semanticscores[r.document] = r.score
maxbm25 = max(bm25scores) if max(bm25scores) > 0 else 1
bm25normalized = {self.documents[i]: s / maxbm25 for i, s in enumerate(bm25scores)}
# Combine scores using weighted fusion
combined = {}
alldocs = set(semanticscores.keys()) | set(d for d, s in bm25normalized.items() if s > 0)
for doc in alldocs:
semscore = semanticscores.get(doc, 0)
kwscore = bm25normalized.get(doc, 0)
combined[doc] = self.alpha semscore + (1 - self.alpha) kwscore
# Sort and build results
sorteddocs = sorted(combined.items(), key=lambda x: x[1], reverse=True)[:k]
results = []
for rank, (doc, score) in enumerate(sorteddocs, 1):
idx = self.documents.index(doc)
meta = self.metadata[idx] if filterfn is None or filterfn(self.metadata[idx]) else {}
results.append(SearchResult(document=doc, score=score, metadata=meta, rank=rank))
elapsed = (time.time() - start) 1000
return SearchResponse(
query=query,
results=results,
totalresults=len(results),
searchtimems=round(elapsed, 2),
)
Usage
hybrid = HybridSearchEngine(alpha=0.6) # 60% semantic, 40% keyword
hybrid.indexdocuments(documents, metadata)
response = hybrid.search("PostgreSQL database performance")
for r in response.results[:5]:
print(f" {r.rank}. [{r.score:.4f}] {r.document}")
Building the API with FastAPI
from fastapi import FastAPI, HTTPException, Query
from pydantic import BaseModel, Field
from typing import Optional
import uvicorn
app = FastAPI(title="Semantic Search API", version="1.0.0")
Initialize the search engine at startup
searchengine = HybridSearchEngine(alpha=0.6)
class DocumentInput(BaseModel):
text: str = Field(description="Document text to index")
metadata: dict = Field(defaultfactory=dict, description="Optional metadata")
class BulkIndexRequest(BaseModel):
documents: list[DocumentInput]
class SearchRequest(BaseModel):
query: str = Field(description="Search query")
k: int = Field(default=10, ge=1, le=100, description="Number of results")
alpha: Optional[float] = Field(default=None, ge=0, le=1, description="Semantic weight")
category: Optional[str] = Field(default=None, description="Filter by category")
class SearchResultResponse(BaseModel):
document: str
score: float
metadata: dict
rank: int
class SearchResponseModel(BaseModel):
query: str
results: list[SearchResultResponse]
totalresults: int
searchtimems: float
@app.post("/index", statuscode=201)
async def indexdocuments(request: BulkIndexRequest):
"""Index a batch of documents."""
texts = [d.text for d in request.documents]
meta = [d.metadata for d in request.documents]
searchengine.indexdocuments(texts, meta)
return {"indexed": len(texts), "total": searchengine.semanticengine.documentcount}
@app.post("/search", responsemodel=SearchResponseModel)
async def search(request: SearchRequest):
"""Search for documents matching the query."""
if searchengine.semanticengine.documentcount == 0:
raise HTTPException(statuscode=400, detail="No documents indexed")
filterfn = None
if request.category:
filterfn = lambda m: m.get("category") == request.category
if request.alpha is not None:
searchengine.alpha = request.alpha
response = searchengine.search(request.query, k=request.k, filterfn=filterfn)
return SearchResponseModel(
query=response.query,
results=[
SearchResultResponse(
document=r.document, score=r.score,
metadata=r.metadata, rank=r.rank
)
for r in response.results
],
totalresults=response.totalresults,
searchtimems=response.searchtimems,
)
@app.get("/health")
async def healthcheck():
return {"status": "healthy", "documentsindexed": searchengine.semanticengine.documentcount}
if name == "main":
uvicorn.run(app, host="0.0.0.0", port=8000)
Evaluation Metrics
Measuring search quality is critical for iterating and improving your system.
from sklearn.metrics import ndcgscore
import numpy as np
class SearchEvaluator:
"""Evaluate search engine quality with standard IR metrics."""
@staticmethod
def precisionatk(relevant: set[str], retrieved: list[str], k: int) -> float:
"""Proportion of retrieved documents that are relevant."""
retrievedk = retrieved[:k]
relevantretrieved = sum(1 for doc in retrievedk if doc in relevant)
return relevantretrieved / k if k > 0 else 0.0
@staticmethod
def recallatk(relevant: set[str], retrieved: list[str], k: int) -> float:
"""Proportion of relevant documents that are retrieved."""
retrievedk = set(retrieved[:k])
relevantretrieved = len(relevant & retrievedk)
return relevantretrieved / len(relevant) if relevant else 0.0
@staticmethod
def meanreciprocalrank(relevant: set[str], retrieved: list[str]) -> float:
"""Reciprocal of the rank of the first relevant document."""
for i, doc in enumerate(retrieved, 1):
if doc in relevant:
return 1.0 / i
return 0.0
@staticmethod
def averageprecision(relevant: set[str], retrieved: list[str]) -> float:
"""Average of precision values at each relevant document position."""
precisions = []
relevantcount = 0
for i, doc in enumerate(retrieved, 1):
if doc in relevant:
relevantcount += 1
precisions.append(relevantcount / i)
return sum(precisions) / len(relevant) if relevant else 0.0
def evaluate(self, engine, testqueries: list[dict], k: int = 10) -> dict:
"""Run full evaluation on a set of test queries.
testqueries: list of {"query": str, "relevantdocs": set[str]}
"""
metrics = {"precision@k": [], "recall@k": [], "mrr": [], "map": []}
for tq in testqueries:
response = engine.search(tq["query"], k=k)
retrieved = [r.document for r in response.results]
relevant = tq["relevantdocs"]
metrics["precision@k"].append(self.precisionatk(relevant, retrieved, k))
metrics["recall@k"].append(self.recallatk(relevant, retrieved, k))
metrics["mrr"].append(self.meanreciprocalrank(relevant, retrieved))
metrics["map"].append(self.averageprecision(relevant, retrieved))
return {
name: {"mean": np.mean(values), "std": np.std(values)}
for name, values in metrics.items()
}
Example evaluation
evaluator = SearchEvaluator()
testqueries = [
{
"query": "programming languages",
"relevantdocs": {
"Python is great for data science and machine learning.",
"JavaScript is the language of the web.",
}
},
{
"query": "container orchestration",
"relevantdocs": {
"Docker helps with containerization and deployment.",
"Kubernetes orchestrates container workloads.",
}
},
]
results = evaluator.evaluate(hybrid, testqueries, k=5)
for metric, values in results.items():
print(f" {metric}: {values['mean']:.4f} (+/- {values['std']:.4f})")
Best Practices
Conclusion
Building a semantic search engine involves multiple components working together: embedding models for text representation, vector indices for efficient retrieval, rerankers for precision, and hybrid approaches for comprehensive coverage.
Key takeaways:
- Sentence-Transformers provides excellent pre-trained models for text embeddings.
- FAISS and Annoy offer different tradeoffs for vector indexing (accuracy vs. speed vs. memory).
- Hybrid search combining BM25 and semantic retrieval outperforms either approach alone.
- Cross-encoder reranking significantly improves result relevance for top positions.
- Systematic evaluation with standard IR metrics is essential for measuring and improving quality.
With these building blocks, you can create a search system that truly understands user intent and delivers relevant results at scale.