Tutorial 10: Milvus - Distributed Vector Database for AI
Table of Contents
Introduction
As AI applications increasingly rely on semantic search, recommendation systems, and retrieval-augmented generation (RAG), the need for efficient vector storage and retrieval has become critical. Milvus is an open-source, distributed vector database purpose-built for handling billion-scale vector data with millisecond latency.
Unlike general-purpose databases that bolt on vector search as an afterthought, Milvus was designed from the ground up for similarity search. It supports multiple index types (IVFFLAT, HNSW, IVFPQ, and more), hybrid search combining vector similarity with scalar filtering, horizontal scaling across clusters, and seamless integration with ML frameworks.
This tutorial provides a comprehensive, hands-on guide to Milvus, from basic setup to production deployment.
Prerequisites
- Python 3.9 or higher
- Docker and Docker Compose (for local Milvus deployment)
- Basic understanding of vector embeddings and similarity search
- Familiarity with Python data structures
Install the required packages:
pip install pymilvus langchain langchain-openai numpy pandas
Milvus Architecture
Milvus uses a cloud-native, disaggregated architecture with four key layers:
Access Layer - Stateless proxy nodes that handle client connections, request routing, and result aggregation. These nodes are horizontally scalable and sit behind a load balancer. Coordinator Service - The brain of the cluster, responsible for metadata management, query coordination, and data coordination. It manages collection schemas, index building tasks, and query routing. Worker Nodes - Divided into three types:- Query Nodes: Execute search and query operations on loaded segments
- Data Nodes: Handle data insertion, deletion, and compaction
- Index Nodes: Build vector indexes in the background
Client Applications
|
[Access Layer - Proxy Nodes]
|
[Coordinator Service]
/ | \
[Query] [Data] [Index]
Nodes Nodes Nodes
|
[Object Storage + etcd]
Installation and Setup
Local Setup with Docker Compose
# Download the docker-compose file
wget https://github.com/milvus-io/milvus/releases/download/v2.4.0/milvus-standalone-docker-compose.yml -O docker-compose.yml
Start Milvus
docker compose up -d
Verify it is running
docker compose ps
Connecting from Python
from pymilvus import connections, utility
Connect to Milvus
connections.connect(
alias="default",
host="localhost",
port="19530"
)
Verify connection
print(f"Connected to Milvus. Server version: {utility.getserverversion()}")
List existing collections
collections = utility.listcollections()
print(f"Existing collections: {collections}")
Using Milvus Lite (In-Process)
from pymilvus import MilvusClient
Milvus Lite runs in-process, no server needed
client = MilvusClient("./milvusdemo.db")
This is ideal for development, testing, and small datasets
print("Milvus Lite client initialized successfully.")
Collection, Partition, and Index Management
Creating a Collection
from pymilvus import (
Collection, CollectionSchema, FieldSchema, DataType,
connections, utility
)
connections.connect(host="localhost", port="19530")
Define the schema
fields = [
FieldSchema(
name="id",
dtype=DataType.INT64,
isprimary=True,
autoid=True,
),
FieldSchema(
name="title",
dtype=DataType.VARCHAR,
maxlength=500,
),
FieldSchema(
name="category",
dtype=DataType.VARCHAR,
maxlength=100,
),
FieldSchema(
name="year",
dtype=DataType.INT32,
),
FieldSchema(
name="embedding",
dtype=DataType.FLOATVECTOR,
dim=1536, # OpenAI text-embedding-3-small dimension
),
]
schema = CollectionSchema(
fields=fields,
description="Document collection for semantic search",
enabledynamicfield=True, # Allow additional fields
)
Create the collection
collectionname = "documents"
if utility.hascollection(collectionname):
utility.dropcollection(collectionname)
collection = Collection(
name=collectionname,
schema=schema,
consistencylevel="Strong",
)
print(f"Collection '{collectionname}' created successfully.")
print(f"Schema: {collection.schema}")
Managing Partitions
# Create partitions for data organization
collection = Collection("documents")
Create partitions by category
partitions = ["technology", "science", "business", "health"]
for partitionname in partitions:
if not collection.haspartition(partitionname):
collection.createpartition(partitionname)
print(f"Partition '{partitionname}' created.")
List all partitions
print(f"All partitions: {[p.name for p in collection.partitions]}")
Insert data into a specific partition
import numpy as np
data = [
["Introduction to Machine Learning"], # title
["technology"], # category
[2024], # year
[np.random.rand(1536).tolist()], # embedding
]
collection.insert(data, partitionname="technology")
print("Data inserted into 'technology' partition.")
Index Management
# Create an index on the vector field
indexparams = {
"metrictype": "COSINE", # L2, IP, or COSINE
"indextype": "HNSW",
"params": {
"M": 16, # Number of bi-directional links
"efConstruction": 256, # Construction-time search width
},
}
collection.createindex(
fieldname="embedding",
indexparams=indexparams,
)
print("HNSW index created on 'embedding' field.")
Create scalar indexes for filtering
collection.createindex(
fieldname="category",
indexparams={"indextype": "Trie"},
)
collection.createindex(
fieldname="year",
indexparams={"indextype": "STLSORT"},
)
print("Scalar indexes created on 'category' and 'year' fields.")
Load collection into memory for searching
collection.load()
print("Collection loaded into memory.")
Vector Indexing Strategies
IVFFLAT: Balanced Speed and Accuracy
# IVFFLAT partitions vectors into clusters using k-means
Good for medium-sized datasets where accuracy matters
ivfflatparams = {
"metrictype": "COSINE",
"indextype": "IVFFLAT",
"params": {
"nlist": 1024, # Number of cluster centroids
},
}
Search parameters for IVFFLAT
searchparamsivf = {
"metrictype": "COSINE",
"params": {
"nprobe": 64, # Number of clusters to search (higher = more accurate)
},
}
When to use:
- Dataset size: 1M - 10M vectors
- Need good recall with reasonable speed
- Cannot tolerate compression artifacts
HNSW: High Recall with Graph-Based Search
# HNSW builds a multi-layer navigable graph
Best for high-recall requirements
hnswparams = {
"metrictype": "COSINE",
"indextype": "HNSW",
"params": {
"M": 32, # Higher M = better recall, more memory
"efConstruction": 512, # Higher = better graph quality, slower build
},
}
Search parameters for HNSW
searchparamshnsw = {
"metrictype": "COSINE",
"params": {
"ef": 256, # Search breadth (higher = better recall, slower)
},
}
When to use:
- Need top-tier recall (>99%)
- Have sufficient memory (HNSW is memory-intensive)
- Dataset size: up to ~50M vectors per node
IVFPQ: Memory-Efficient for Large Datasets
# IVFPQ combines inverted file index with product quantization
Trades accuracy for massive memory savings
ivf
pqparams = {
"metric
type": "L2",
"indextype": "IVFPQ",
"params": {
"nlist": 2048, # Number of clusters
"m": 16, # Number of sub-quantizers (must divide dimension evenly)
"nbits": 8, # Bits per sub-quantizer code
},
}
Search parameters for IVFPQ
searchparamspq = {
"metrictype": "L2",
"params": {
"nprobe": 128,
},
}
When to use:
- Billion-scale datasets
- Memory is the primary constraint
- Can tolerate ~5-10% recall drop
- Memory savings: ~32x compared to flat index for 1536-dim vectors
Choosing the Right Index
def recommendindex(numvectors: int, dimension: int,
memorygb: float, targetrecall: float) -> str:
"""Recommend an index type based on requirements."""
flatmemory = (numvectors dimension 4) / (1024*3) # GB
if numvectors < 100000:
return "FLAT - Small dataset, brute force is fine"
if targetrecall > 0.99 and flatmemory < memorygb 0.5:
return "HNSW - High recall requirement, sufficient memory"
if flatmemory > memorygb:
return "IVFPQ - Dataset too large for memory, use compression"
return "IVFFLAT - Good balance of speed, accuracy, and memory"
Examples
print(recommendindex(50000, 1536, 16, 0.95))
print(recommendindex(5000000, 1536, 32, 0.99))
print(recommendindex(500000000, 768, 64, 0.90))
Hybrid Search: Vector + Scalar Filtering
Basic Hybrid Search
from pymilvus import Collection
import numpy as np
collection = Collection("documents")
collection.load()
Generate a query embedding (in practice, use your embedding model)
queryembedding = np.random.rand(1536).tolist()
Search with scalar filtering
results = collection.search(
data=[queryembedding],
annsfield="embedding",
param={"metrictype": "COSINE", "params": {"ef": 128}},
limit=10,
expr='category == "technology" and year >= 2023',
outputfields=["title", "category", "year"],
)
for hits in results:
for hit in hits:
print(f"ID: {hit.id}")
print(f" Score: {hit.score:.4f}")
print(f" Title: {hit.entity.get('title')}")
print(f" Category: {hit.entity.get('category')}")
print(f" Year: {hit.entity.get('year')}")
print()
Complex Filter Expressions
# Boolean expressions for filtering
filters = [
'category == "technology"',
'year >= 2023 and year <= 2025',
'category in ["technology", "science"]',
'title like "Machine%"',
'year > 2020 and category != "business"',
]
for filterexpr in filters:
results = collection.search(
data=[queryembedding],
annsfield="embedding",
param={"metrictype": "COSINE", "params": {"ef": 64}},
limit=5,
expr=filterexpr,
outputfields=["title", "category", "year"],
)
count = sum(len(hits) for hits in results)
print(f"Filter: {filterexpr} -> {count} results")
Multi-Vector Search (Reranking)
from pymilvus import AnnSearchRequest, RRFRanker
Search with multiple query vectors and rerank
request1 = AnnSearchRequest(
data=[queryembedding],
annsfield="embedding",
param={"metrictype": "COSINE", "params": {"ef": 128}},
limit=20,
expr='category == "technology"',
)
Second query vector (e.g., from a different representation)
queryembedding2 = np.random.rand(1536).tolist()
request2 = AnnSearchRequest(
data=[queryembedding2],
annsfield="embedding",
param={"metrictype": "COSINE", "params": {"ef": 128}},
limit=20,
)
Combine results using Reciprocal Rank Fusion
reranker = RRFRanker(k=60)
results = collection.hybridsearch(
reqs=[request1, request2],
ranker=reranker,
limit=10,
outputfields=["title", "category"],
)
for hits in results:
for rank, hit in enumerate(hits, 1):
print(f"Rank {rank}: {hit.entity.get('title')} (score: {hit.score:.4f})")
Batch Operations
Bulk Insert
import numpy as np
from pymilvus import Collection
collection = Collection("documents")
Generate batch data
batchsize = 10000
titles = [f"Document {i}" for i in range(batchsize)]
categories = np.random.choice(
["technology", "science", "business", "health"],
size=batchsize
).tolist()
years = np.random.randint(2020, 2026, size=batchsize).tolist()
embeddings = np.random.rand(batchsize, 1536).tolist()
Insert in batches to avoid memory issues
chunksize = 2000
for i in range(0, batchsize, chunksize):
end = min(i + chunksize, batchsize)
data = [
titles[i:end],
categories[i:end],
years[i:end],
embeddings[i:end],
]
collection.insert(data)
print(f"Inserted batch {i // chunksize + 1}: "
f"records {i} to {end}")
Flush to ensure data is persisted
collection.flush()
print(f"Total entities: {collection.numentities}")
Bulk Delete
# Delete by expression
collection.delete(expr='year < 2022')
print("Deleted all documents before 2022.")
Delete by primary keys
idstodelete = [1, 2, 3, 4, 5]
collection.delete(expr=f"id in {idstodelete}")
print(f"Deleted {len(idstodelete)} documents by ID.")
Compact to reclaim space
collection.compact()
print("Compaction started.")
Upsert Operations
# Upsert (insert or update) data
upsertdata = [
[100, 101, 102], # id (explicit)
["Updated Doc 1", "Updated Doc 2", "Updated Doc 3"], # title
["technology", "science", "business"], # category
[2025, 2025, 2025], # year
[np.random.rand(1536).tolist() for in range(3)], # embedding
]
Note: for upsert, autoid must be False and IDs must be provided
collection.upsert(upsertdata)
Scaling with Kubernetes
Helm Chart Deployment
# Add the Milvus Helm repository
helm repo add milvus https://zilliztech.github.io/milvus-helm/
helm repo update
Install Milvus cluster
helm install my-milvus milvus/milvus \
--set cluster.enabled=true \
--set queryNode.replicas=3 \
--set dataNode.replicas=2 \
--set indexNode.replicas=2 \
--set proxy.replicas=2 \
--namespace milvus \
--create-namespace
Resource Configuration
# values.yaml for production deployment
cluster:
enabled: true
queryNode:
replicas: 3
resources:
requests:
cpu: "4"
memory: "16Gi"
limits:
cpu: "8"
memory: "32Gi"
dataNode:
replicas: 2
resources:
requests:
cpu: "2"
memory: "8Gi"
limits:
cpu: "4"
memory: "16Gi"
indexNode:
replicas: 2
resources:
requests:
cpu: "4"
memory: "16Gi"
limits:
cpu: "8"
memory: "32Gi"
proxy:
replicas: 2
resources:
requests:
cpu: "1"
memory: "4Gi"
limits:
cpu: "2"
memory: "8Gi"
minio:
persistence:
size: 500Gi
resources:
requests:
cpu: "2"
memory: "4Gi"
etcd:
replicaCount: 3
persistence:
size: 20Gi
Comparison vs Qdrant, ChromaDB, and Pinecone
"""
Vector Database Comparison Matrix
| Feature | Milvus | Qdrant | ChromaDB | Pinecone |
|----------------------|----------------|----------------|----------------|----------------|
| Deployment | Self-hosted/ | Self-hosted/ | Self-hosted/ | Fully managed |
| | Zilliz Cloud | Qdrant Cloud | Embedded | |
| Max Vectors | Billions | Billions | Millions | Billions |
| Index Types | IVF, HNSW, PQ, | HNSW | HNSW | Proprietary |
| | DiskANN, +more | | | |
| Hybrid Search | Yes | Yes | Limited | Yes |
| Scalar Filtering | Full SQL-like | JSON payload | Metadata | Metadata |
| Multi-tenancy | Partitions | Collections | Collections | Namespaces |
| GPU Acceleration | Yes | No | No | N/A |
| Replication | Yes | Yes | No | Managed |
| License | Apache 2.0 | Apache 2.0 | Apache 2.0 | Proprietary |
| Best For | Large-scale | Mid-scale, | Prototyping, | Zero-ops, |
| | production | developer- | small apps | enterprise |
| | | friendly | | |
"""
Quick comparison in code
def compareinsertperformance():
"""Compare basic operations across vector databases."""
# Milvus
from pymilvus import Collection
import time
import numpy as np
collection = Collection("documents")
vectors = np.random.rand(1000, 1536).tolist()
titles = [f"Doc {i}" for i in range(1000)]
categories = ["tech"] 1000
years = [2024] 1000
start = time.time()
collection.insert([titles, categories, years, vectors])
collection.flush()
milvustime = time.time() - start
print(f"Milvus insert 1000 vectors: {milvustime:.3f}s")
Integration with LangChain
Using Milvus as a LangChain Vector Store
from langchainopenai import OpenAIEmbeddings, ChatOpenAI
from langchain
community.vectorstores import Milvus
from langchaincore.documents import Document
from langchaincore.prompts import ChatPromptTemplate
from langchaincore.outputparsers import StrOutputParser
Initialize embeddings
embeddings = OpenAIEmbeddings(model="text-embedding-3-small")
Create documents
documents = [
Document(
pagecontent="Milvus is an open-source vector database designed for AI applications.",
metadata={"source": "docs", "category": "database", "year": 2024},
),
Document(
pagecontent="HNSW is a graph-based index that provides high recall for vector search.",
metadata={"source": "docs", "category": "indexing", "year": 2024},
),
Document(
pagecontent="RAG combines retrieval and generation for knowledge-grounded AI responses.",
metadata={"source": "blog", "category": "architecture", "year": 2024},
),
Document(
pagecontent="Vector quantization reduces memory usage by compressing vector representations.",
metadata={"source": "paper", "category": "optimization", "year": 2023},
),
Document(
pagecontent="Hybrid search combines dense vector similarity with sparse keyword matching.",
metadata={"source": "docs", "category": "search", "year": 2024},
),
]
Create Milvus vector store
vectorstore = Milvus.fromdocuments(
documents=documents,
embedding=embeddings,
collectionname="langchaindocs",
connectionargs={"host": "localhost", "port": "19530"},
dropold=True,
)
print("Vector store created with LangChain integration.")
RAG Pipeline with Milvus
# Create a retriever
retriever = vectorstore.asretriever(
searchtype="similarity",
searchkwargs={"k": 3},
)
Build RAG chain
llm = ChatOpenAI(model="gpt-4o", temperature=0)
prompt = ChatPromptTemplate.fromtemplate(
"Answer the question based on the provided context.\n\n"
"Context:\n{context}\n\n"
"Question: {question}\n\n"
"Answer:"
)
def formatdocs(docs):
return "\n\n".join(doc.pagecontent for doc in docs)
chain = (
{"context": retriever | formatdocs, "question": lambda x: x}
| prompt
| llm
| StrOutputParser()
)
Query the RAG pipeline
answer = chain.invoke("What is hybrid search and how does it work?")
print(f"Answer: {answer}")
Search with metadata filtering
filteredresults = vectorstore.similaritysearch(
"vector indexing techniques",
k=3,
expr='category == "indexing" or category == "optimization"',
)
for doc in filteredresults:
print(f"Content: {doc.pagecontent[:80]}...")
print(f"Metadata: {doc.metadata}")
print()
Production Deployment
Health Monitoring
from pymilvus import connections, utility
def checkmilvushealth() -> dict:
"""Comprehensive health check for Milvus."""
try:
connections.connect(host="localhost", port="19530")
health = {
"connected": True,
"version": utility.getserverversion(),
"collections": [],
}
for name in utility.listcollections():
from pymilvus import Collection
col = Collection(name)
colinfo = {
"name": name,
"numentities": col.numentities,
"loaded": utility.loadstate(name).name,
"partitions": len(col.partitions),
"indexes": [idx.todict() for idx in col.indexes],
}
health["collections"].append(colinfo)
return health
except Exception as e:
return {"connected": False, "error": str(e)}
Run health check
status = checkmilvushealth()
print(f"Milvus Status: {'Healthy' if status['connected'] else 'Unhealthy'}")
if status["connected"]:
print(f"Version: {status['version']}")
for col in status["collections"]:
print(f" Collection '{col['name']}': {col['numentities']} entities, "
f"state: {col['loaded']}")
Backup and Restore
# Using milvus-backup tool
Install: https://github.com/zilliztech/milvus-backup
import subprocess
def backupcollection(collectionname: str, backupname: str):
"""Backup a Milvus collection."""
cmd = [
"milvus-backup", "create",
"-n", backupname,
"-c", collectionname,
"--config", "/path/to/backup.yaml"
]
result = subprocess.run(cmd, captureoutput=True, text=True)
print(f"Backup result: {result.stdout}")
def restorecollection(backupname: str):
"""Restore a Milvus collection from backup."""
cmd = [
"milvus-backup", "restore",
"-n", backupname,
"--config", "/path/to/backup.yaml"
]
result = subprocess.run(cmd, captureoutput=True, text=True)
print(f"Restore result: {result.stdout}")
Performance Tuning
# Optimize search performance
def tunesearchparams(collectionname: str, queryvectors: list,
groundtruth: list) -> dict:
"""Find optimal search parameters by testing different configurations."""
from pymilvus import Collection
import time
collection = Collection(collectionname)
collection.load()
configs = [
{"ef": 64},
{"ef": 128},
{"ef": 256},
{"ef": 512},
]
results = []
for config in configs:
start = time.time()
searchresults = collection.search(
data=queryvectors,
annsfield="embedding",
param={"metrictype": "COSINE", "params": config},
limit=10,
)
elapsed = time.time() - start
results.append({
"params": config,
"latencyms": elapsed * 1000,
"resultscount": sum(len(hits) for hits in searchresults),
})
print("Search Parameter Tuning Results:")
for r in results:
print(f" ef={r['params']['ef']}: "
f"latency={r['latencyms']:.1f}ms, "
f"results={r['resultscount']}")
return results
Best Practices
Conclusion
Milvus provides a robust, production-ready foundation for AI applications that require vector similarity search at scale. Its disaggregated architecture allows independent scaling of compute and storage, while its support for multiple index types gives you the flexibility to optimize for your specific latency, recall, and memory requirements.
The key decisions when deploying Milvus are choosing the right index type for your scale and accuracy needs, designing your partition strategy for multi-tenancy and efficient filtering, and properly sizing your cluster for your workload. With the techniques covered in this tutorial, you have the knowledge to make these decisions confidently and build high-performance vector search into your AI applications.
Start with Milvus Lite for development, graduate to standalone Docker for staging, and deploy a full Kubernetes cluster for production. The API remains the same across all deployment modes, making the transition seamless.