LanceDB: Serverless Vector Database for Multimodal AI Applications

# LanceDB: Database Vektor Serverless untuk Aplikasi AI Multimodal Database vektor telah menjadi komponen fundamental dalam aplikasi AI modern, mulai dari pencarian semantik hingga Retrieval-Augmente...

By Ruby Abdullah · · tutorial
LanceDBVector DatabaseMultimodalEmbeddingPython

LanceDB: Serverless Vector Database for Multimodal AI Applications

Vector databases have become a fundamental component in modern AI applications, from semantic search to Retrieval-Augmented Generation (RAG). However, many vector database solutions require complex server infrastructure that is expensive to maintain. LanceDB offers a lightweight, serverless alternative with multimodal search support.

LanceDB is an open-source vector database that runs embedded, meaning it requires no separate server. Built on the Lance data format optimized for vector operations, LanceDB delivers high performance with a minimal footprint. What makes it special is its native support for multimodal data, including text, images, audio, and video.

In this tutorial, we will learn how to use LanceDB from installation, basic operations, to building a complete multimodal search engine.

Prerequisites

Before starting, make sure you have:

  • Python 3.9 or later
  • pip package manager
  • Basic understanding of Python and vector embedding concepts
  • (Optional) OpenAI API key for embedding functions

Installation

Basic Installation

pip install lancedb

Installation with Embedding Functions

pip install lancedb sentence-transformers

pip install lancedb open-clip-torch Pillow

Verify Installation

import lancedb

print(f"LanceDB version: {lancedb.version}")

Creating Databases and Tables

LanceDB uses an embedded approach, so a database is simply created as a local directory.

Creating a Database

import lancedb

Create database connection (local directory)

db = lancedb.connect("./mylancedb")

print("Database created successfully!")

print(f"Location: ./mylancedb")

Creating a Table with Data

import lancedb

import numpy as np

db = lancedb.connect("./mylancedb")

Create data with embeddings

data = [

{

"id": 1,

"text": "Python is a popular programming language for AI",

"vector": np.random.randn(128).tolist(),

"category": "programming",

},

{

"id": 2,

"text": "Machine learning uses data to make predictions",

"vector": np.random.randn(128).tolist(),

"category": "ai",

},

{

"id": 3,

"text": "Deep learning is a subset of machine learning",

"vector": np.random.randn(128).tolist(),

"category": "ai",

},

]

Create table

table = db.createtable("articles", data=data)

print(f"Table 'articles' created with {len(table)} rows")

Using Pydantic Models

import lancedb

from lancedb.pydantic import LanceModel, Vector

import numpy as np

Define schema using Pydantic

class Article(LanceModel):

id: int

title: str

content: str

vector: Vector(384) # Embedding dimension

category: str

published: bool = True

db = lancedb.connect("./mylancedb")

Create table with schema

table = db.createtable("articlesv2", schema=Article)

Add data

articles = [

Article(

id=1,

title="Introduction to LanceDB",

content="LanceDB is a serverless vector database",

vector=np.random.randn(384).tolist(),

category="database",

),

Article(

id=2,

title="RAG Tutorial",

content="RAG combines retrieval with generation",

vector=np.random.randn(384).tolist(),

category="ai",

),

]

table.add([a.dict() for a in articles])

print(f"Added {len(articles)} articles")

Adding Data

Adding Data to an Existing Table

import lancedb

import numpy as np

db = lancedb.connect("./mylancedb")

table = db.opentable("articles")

Add new data

newdata = [

{

"id": 4,

"text": "Natural Language Processing processes human language",

"vector": np.random.randn(128).tolist(),

"category": "nlp",

},

{

"id": 5,

"text": "Computer Vision enables computers to understand images",

"vector": np.random.randn(128).tolist(),

"category": "cv",

},

]

table.add(newdata)

print(f"Total rows now: {len(table)}")

Updating Data

import lancedb

db = lancedb.connect("./mylancedb")

table = db.opentable("articles")

Update data based on condition

table.update(

where="id = 1",

values={"category": "python"},

)

print("Data updated successfully")

Deleting Data

import lancedb

db = lancedb.connect("./mylancedb")

table = db.opentable("articles")

Delete data based on condition

table.delete("id = 5")

print(f"Total rows after deletion: {len(table)}")

Vector search is LanceDB's primary feature. Here is how to perform various types of searches.

L2 Search (Euclidean Distance)

import lancedb

import numpy as np

db = lancedb.connect("./mylancedb")

table = db.opentable("articles")

Query vector

queryvector = np.random.randn(128).tolist()

Vector search (default: L2 distance)

results = (

table.search(queryvector)

.limit(5)

.topandas()

)

print("Search results (L2):")

print(results[["id", "text", "distance"]])

import lancedb

import numpy as np

db = lancedb.connect("./mylancedb")

table = db.opentable("articles")

queryvector = np.random.randn(128).tolist()

Search with cosine similarity

results = (

table.search(queryvector)

.metric("cosine")

.limit(5)

.topandas()

)

print("Search results (Cosine):")

print(results[["id", "text", "distance"]])

import lancedb

import numpy as np

db = lancedb.connect("./mylancedb")

table = db.opentable("articles")

queryvector = np.random.randn(128).tolist()

Search with dot product

results = (

table.search(queryvector)

.metric("dot")

.limit(5)

.topandas()

)

print("Search results (Dot Product):")

print(results[["id", "text", "distance"]])

LanceDB also supports full-text search using Tantivy.

import lancedb

db = lancedb.connect("./mylancedb")

Create table with FTS index

data = [

{"id": 1, "text": "Python programming language for AI development"},

{"id": 2, "text": "Machine learning algorithms and deep learning"},

{"id": 3, "text": "Natural language processing with transformers"},

{"id": 4, "text": "Computer vision and image recognition"},

{"id": 5, "text": "Reinforcement learning for game AI"},

]

table = db.createtable("ftsarticles", data=data, mode="overwrite")

Create full-text search index

table.createftsindex("text")

Text search

results = (

table.search("machine learning", querytype="fts")

.limit(3)

.topandas()

)

print("Full-Text Search results:")

print(results)

Hybrid search combines vector search and full-text search for more accurate results.

import lancedb

import numpy as np

db = lancedb.connect("./mylancedb")

Create data with vectors and text

data = [

{

"id": i,

"text": text,

"vector": np.random.randn(128).tolist(),

}

for i, text in enumerate([

"Introduction to machine learning algorithms",

"Deep learning with neural networks",

"Natural language processing fundamentals",

"Computer vision and object detection",

"Reinforcement learning in robotics",

"Transfer learning for NLP tasks",

"Generative AI and large language models",

"Vector databases for semantic search",

])

]

table = db.createtable("hybridarticles", data=data, mode="overwrite")

Create FTS index

table.createftsindex("text")

Hybrid search

queryvector = np.random.randn(128).tolist()

results = (

table.search(queryvector, querytype="hybrid")

.text("neural networks")

.limit(5)

.topandas()

)

print("Hybrid Search results:")

print(results[["id", "text", "relevancescore"]])

Filtering

LanceDB supports SQL-like filtering on vector searches.

import lancedb

import numpy as np

db = lancedb.connect("./mylancedb")

data = [

{

"id": i,

"title": f"Article {i}",

"text": f"Content of article {i}",

"vector": np.random.randn(128).tolist(),

"category": ["ai", "database", "web"][i % 3],

"year": 2024 + (i % 3),

}

for i in range(20)

]

table = db.createtable("filteredarticles", data=data, mode="overwrite")

queryvector = np.random.randn(128).tolist()

Search with filter

results = (

table.search(queryvector)

.where("category = 'ai' AND year >= 2025")

.limit(5)

.topandas()

)

print("Search results with filter:")

print(results[["id", "title", "category", "year", "distance"]])

Filter with IN clause

resultsin = (

table.search(queryvector)

.where("category IN ('ai', 'database')")

.limit(5)

.topandas()

)

print("\nResults with IN filter:")

print(resultsin[["id", "title", "category", "distance"]])

Indexing (IVFPQ)

For large datasets, LanceDB supports IVFPQ indexing to speed up searches.

import lancedb

import numpy as np

db = lancedb.connect("./mylancedb")

Create large dataset

largedata = [

{

"id": i,

"text": f"Document {i}",

"vector": np.random.randn(128).tolist(),

}

for i in range(10000)

]

table = db.createtable("largedataset", data=largedata, mode="overwrite")

Create IVFPQ index

table.createindex(

metric="cosine",

numpartitions=16,

numsubvectors=8,

indextype="IVFPQ",

)

print("IVFPQ index created successfully!")

Search with index

queryvector = np.random.randn(128).tolist()

results = (

table.search(queryvector)

.metric("cosine")

.nprobes(8) # Number of partitions to check

.limit(10)

.topandas()

)

print(f"\nTop 10 results:")

print(results[["id", "text", "distance"]])

Embedding Functions

LanceDB provides built-in embedding functions that simplify automatic embedding generation.

Using Sentence Transformers

import lancedb

from lancedb.pydantic import LanceModel, Vector

from lancedb.embeddings import getregistry

Get embedding function

sentencetransformers = getregistry().get("sentence-transformers")

embeddingfunc = sentencetransformers.create(

name="all-MiniLM-L6-v2",

device="cpu",

)

Define model with automatic embedding

class Document(LanceModel):

text: str = embeddingfunc.SourceField()

vector: Vector(embeddingfunc.ndims()) = embeddingfunc.VectorField()

category: str = ""

db = lancedb.connect("./mylancedb")

table = db.createtable("autoembed", schema=Document, mode="overwrite")

Add data - embeddings are created automatically

documents = [

{"text": "Python is a popular programming language", "category": "programming"},

{"text": "Machine learning for data prediction", "category": "ai"},

{"text": "Vector databases for semantic search", "category": "database"},

{"text": "Deep learning uses neural networks", "category": "ai"},

{"text": "API development with FastAPI", "category": "web"},

]

table.add(documents)

print(f"Added {len(documents)} documents with automatic embedding")

Search - query is also embedded automatically

results = (

table.search("how to use artificial intelligence")

.limit(3)

.topandas()

)

print("\nSearch results:")

for , row in results.iterrows():

print(f" [{row['distance']:.4f}] {row['text']}")

Using OpenAI Embeddings

import lancedb

from lancedb.pydantic import LanceModel, Vector

from lancedb.embeddings import getregistry

import os

Set API key

os.environ["OPENAIAPIKEY"] = "sk-your-api-key"

Use OpenAI embedding

openaiembed = getregistry().get("openai")

embeddingfunc = openaiembed.create(

name="text-embedding-3-small",

)

class Document(LanceModel):

text: str = embeddingfunc.SourceField()

vector: Vector(embeddingfunc.ndims()) = embeddingfunc.VectorField()

db = lancedb.connect("./mylancedb")

table = db.createtable("openaiembed", schema=Document, mode="overwrite")

Data will be embedded automatically using OpenAI

table.add([

{"text": "Artificial intelligence is transforming industries"},

{"text": "Vector databases enable semantic search"},

{"text": "Large language models understand natural language"},

])

results = table.search("how AI changes business").limit(3).topandas()

print(results[["text", "distance"]])

Using CLIP for Multimodal

import lancedb

from lancedb.pydantic import LanceModel, Vector

from lancedb.embeddings import getregistry

Use CLIP for multimodal embedding

clip = getregistry().get("open-clip")

embeddingfunc = clip.create(

name="ViT-B-32",

pretrained="laion2bs34bb79k",

)

class MultimodalItem(LanceModel):

text: str = embeddingfunc.SourceField()

imageuri: str = embeddingfunc.SourceField()

vector: Vector(embeddingfunc.ndims()) = embeddingfunc.VectorField()

label: str = ""

db = lancedb.connect("./mylancedb")

table = db.createtable(

"multimodal", schema=MultimodalItem, mode="overwrite"

)

print("Multimodal table ready to use!")

Multimodal Search (Text + Images)

LanceDB supports multimodal search that combines text and images.

import lancedb

from lancedb.pydantic import LanceModel, Vector

from lancedb.embeddings import getregistry

from pathlib import Path

from PIL import Image

import numpy as np

Setup CLIP embedding

clip = getregistry().get("open-clip")

embeddingfunc = clip.create(

name="ViT-B-32",

pretrained="laion2bs34bb79k",

)

class ImageDocument(LanceModel):

imageuri: str = embeddingfunc.SourceField()

vector: Vector(embeddingfunc.ndims()) = embeddingfunc.VectorField()

label: str

description: str = ""

db = lancedb.connect("./multimodaldb")

Assuming we have an images directory

imagedir = Path("images")

if imagedir.exists():

imagedata = []

for imgpath in imagedir.glob(".jpg"):

imagedata.append({

"imageuri": str(imgpath),

"label": imgpath.stem,

"description": f"Image: {imgpath.stem}",

})

table = db.createtable(

"imagesearch", data=imagedata, schema=ImageDocument, mode="overwrite"

)

# Search images using text query

results = (

table.search("a photo of a cat")

.limit(5)

.topandas()

)

print("Multimodal search results:")

for , row in results.iterrows():

print(f" [{row['distance']:.4f}] {row['label']}: {row['description']}")

# Search using an image as query

queryimage = "images/queryimage.jpg"

if Path(queryimage).exists():

results = (

table.search(Image.open(queryimage))

.limit(5)

.topandas()

)

print("\nSearch results with image query:")

for , row in results.iterrows():

print(f" [{row['distance']:.4f}] {row['label']}")

Integration with LangChain

LanceDB integrates with LangChain as a vector store.

pip install langchain-community lancedb langchain-openai

from langchaincommunity.vectorstores import LanceDB

from langchainopenai import OpenAIEmbeddings, ChatOpenAI

from langchain.chains import RetrievalQA

from langchain.schema import Document

import lancedb

Create LanceDB connection

db = lancedb.connect("./langchainlancedb")

Setup embeddings

embeddings = OpenAIEmbeddings(model="text-embedding-3-small")

Prepare documents

documents = [

Document(

pagecontent="LanceDB is a fast and lightweight serverless vector database",

metadata={"source": "docs", "topic": "database"},

),

Document(

pagecontent="Vector search enables searching based on semantic meaning",

metadata={"source": "docs", "topic": "search"},

),

Document(

pagecontent="RAG combines retrieval and generation for accurate answers",

metadata={"source": "tutorial", "topic": "rag"},

),

Document(

pagecontent="Embeddings convert text into numerical vector representations",

metadata={"source": "tutorial", "topic": "embedding"},

),

]

Create vector store

vectorstore = LanceDB.fromdocuments(

documents,

embeddings,

connection=db,

tablename="langchaindocs",

)

Similarity search

results = vectorstore.similaritysearch(

"how does vector search work?", k=3

)

for doc in results:

print(f"- {doc.pagecontent}")

print(f" Metadata: {doc.metadata}\n")

RAG chain

llm = ChatOpenAI(model="gpt-4o", temperature=0)

qachain = RetrievalQA.fromchaintype(

llm=llm,

chaintype="stuff",

retriever=vectorstore.asretriever(searchkwargs={"k": 3}),

)

response = qachain.invoke({"query": "What is LanceDB?"})

print(f"\nAnswer: {response['result']}")

Integration with LlamaIndex

LanceDB is also available as a vector store for LlamaIndex.

pip install llama-index-vector-stores-lancedb llama-index

from llamaindex.core import VectorStoreIndex, SimpleDirectoryReader

from llamaindex.vectorstores.lancedb import LanceDBVectorStore

from llamaindex.core import StorageContext

import lancedb

Create LanceDB connection

db = lancedb.connect("./llamaindexlancedb")

Configure LanceDB vector store

vectorstore = LanceDBVectorStore(

uri="./llamaindexlancedb",

tablename="llamadocs",

)

storagecontext = StorageContext.fromdefaults(

vectorstore=vectorstore

)

Load documents

documents = SimpleDirectoryReader("data/").loaddata()

Create index

index = VectorStoreIndex.fromdocuments(

documents,

storagecontext=storagecontext,

)

Query

queryengine = index.asqueryengine()

response = queryengine.query("What are the main topics in these documents?")

print(response)

Practical Example: Building a Multimodal Search Engine

Here is a complete example of building a multimodal search engine using LanceDB.

"""

Multimodal Search Engine with LanceDB

Supports text and image search in a single system

"""

import lancedb

from lancedb.pydantic import LanceModel, Vector

from lancedb.embeddings import getregistry

from pathlib import Path

from datetime import datetime

from typing import Optional

import json

class MultimodalSearchEngine:

"""Multimodal search engine using LanceDB."""

def init(

self,

dbpath: str = "./multimodalsearchdb",

embeddingmodel: str = "all-MiniLM-L6-v2",

):

self.db = lancedb.connect(dbpath)

# Setup text embedding

st = getregistry().get("sentence-transformers")

self.textembedfunc = st.create(

name=embeddingmodel,

device="cpu",

)

# Define schema

embedfunc = self.textembedfunc

class TextDocument(LanceModel):

text: str = embedfunc.SourceField()

vector: Vector(embedfunc.ndims()) = embedfunc.VectorField()

title: str = ""

source: str = ""

doctype: str = "text"

createdat: str = ""

metadatajson: str = "{}"

self.TextDocument = TextDocument

self.inittables()

def inittables(self):

"""Initialize tables if they don't exist."""

try:

self.texttable = self.db.opentable("textdocuments")

print("Table 'textdocuments' found")

except Exception:

self.texttable = self.db.createtable(

"textdocuments",

schema=self.TextDocument,

)

print("Table 'textdocuments' created")

def adddocuments(

self,

texts: list[str],

titles: list[str] = None,

sources: list[str] = None,

metadata: list[dict] = None,

):

"""Add text documents to the database."""

documents = []

for i, text in enumerate(texts):

doc = {

"text": text,

"title": titles[i] if titles else f"Document {i + 1}",

"source": sources[i] if sources else "unknown",

"doctype": "text",

"createdat": datetime.now().isoformat(),

"metadatajson": json.dumps(

metadata[i] if metadata else {}

),

}

documents.append(doc)

self.texttable.add(documents)

print(f"Added {len(documents)} text documents")

def searchtext(

self,

query: str,

limit: int = 10,

filtercondition: str = None,

metric: str = "cosine",

) -> list[dict]:

"""Semantic search on text documents."""

search = (

self.texttable

.search(query)

.metric(metric)

.limit(limit)

)

if filtercondition:

search = search.where(filtercondition)

results = search.topandas()

output = []

for , row in results.iterrows():

output.append({

"title": row.get("title", ""),

"text": row.get("text", ""),

"source": row.get("source", ""),

"score": 1 - row.get("distance", 0),

"distance": row.get("distance", 0),

})

return output

def searchfulltext(

self,

query: str,

limit: int = 10,

) -> list[dict]:

"""Full-text search on documents."""

try:

results = (

self.texttable

.search(query, querytype="fts")

.limit(limit)

.topandas()

)

output = []

for , row in results.iterrows():

output.append({

"title": row.get("title", ""),

"text": row.get("text", ""),

"source": row.get("source", ""),

"score": row.get("score", 0),

})

return output

except Exception as e:

print(f"FTS not indexed. Run createftsindex() first.")

print(f"Error: {e}")

return []

def createftsindex(self):

"""Create full-text search index."""

self.texttable.createftsindex("text")

print("FTS index created successfully")

def createvectorindex(self, numpartitions: int = 16):

"""Create vector index for faster search."""

rowcount = len(self.texttable)

if rowcount < 256:

print(

f"Dataset too small ({rowcount} rows) "

"for indexing. Minimum 256 rows."

)

return

self.texttable.createindex(

metric="cosine",

numpartitions=numpartitions,

numsubvectors=8,

indextype="IVFPQ",

)

print("Vector index IVFPQ created successfully")

def getstats(self) -> dict:

"""Get database statistics."""

return {

"totaltextdocuments": len(self.texttable),

"tables": self.db.tablenames(),

}

def deletebysource(self, source: str):

"""Delete documents by source."""

self.texttable.delete(f"source = '{source}'")

print(f"Documents from source '{source}' deleted successfully")

Usage

if name == "main":

# Initialize search engine

engine = MultimodalSearchEngine(

dbpath="./mysearchengine",

embeddingmodel="all-MiniLM-L6-v2",

)

# Add documents

texts = [

"LanceDB is a lightweight and fast serverless vector database",

"PostgreSQL is a powerful open source relational database",

"MongoDB is a document-based NoSQL database",

"Redis is an in-memory data store for caching",

"Elasticsearch provides powerful full-text search",

"ChromaDB is a vector database for AI applications",

"Pinecone is a managed vector database in the cloud",

"Weaviate supports vector search and hybrid search",

"Milvus is an open source vector database for large scale",

"FAISS is a library from Meta for similarity search",

]

titles = [

"LanceDB Overview", "PostgreSQL Guide", "MongoDB Basics",

"Redis Caching", "Elasticsearch Search", "ChromaDB AI",

"Pinecone Cloud", "Weaviate Search", "Milvus Scale",

"FAISS Library",

]

engine.adddocuments(

texts=texts,

titles=titles,

sources=["docs"] len(texts),

)

# Semantic search

print("\n=== Semantic Search ===")

results = engine.searchtext("database for AI applications", limit=5)

for r in results:

print(f" [{r['score']:.4f}] {r['title']}: {r['text'][:80]}...")

# Create FTS index and search

engine.createftsindex()

print("\n=== Full-Text Search ===")

ftsresults = engine.searchfulltext("vector database", limit=5)

for r in ftsresults:

print(f" [{r['score']:.4f}] {r['title']}: {r['text'][:80]}...")

# Statistics

print(f"\n=== Statistics ===")

stats = engine.getstats()

print(f"Total documents: {stats['totaltextdocuments']}")

print(f"Tables: {stats['tables']}")

Tips and Best Practices

  • Choose the right metric: Use cosine similarity for normalized text embeddings, L2 for raw embeddings, and dot product for models optimized for dot product.
  • Use Pydantic schemas: Define schemas with Pydantic for better data validation and cleaner code.
  • Leverage embedding functions: Use LanceDB's built-in embedding functions to avoid manual embedding management.
  • Index for large datasets: For datasets with more than 50,000 rows, create an IVF_PQ index to speed up searches.
  • Use filtering: Combine vector search with filters for more relevant results.
  • Back up your database: Since LanceDB stores data in a local directory, make sure to perform regular backups.
  • Mind embedding dimensions: Ensure vector dimensions are consistent between data and queries. Mismatches will produce errors.
  • Conclusion

    LanceDB offers a unique vector database solution with a serverless, embedded approach. Without needing to manage server infrastructure, you can immediately build semantic search, RAG, and multimodal search applications with good performance.

    LanceDB's advantages lie in its ease of use, native multimodal support, and seamless integration with the AI ecosystem including LangChain and LlamaIndex. For projects that need a vector database without operational overhead, LanceDB is an excellent choice.

    Start experimenting with LanceDB and discover how this serverless vector database can accelerate your AI application development!

    Related Articles

    Complete Pinecone Tutorial: Vector Database for AI and Semantic Search

    Tutorial Lengkap Pinecone: Vector Database untuk AI dan Semantic Search Pinecone adalah managed vector database yang dir...

    Weaviate: Vector Database with Integrated AI Modules

    Weaviate: Database Vektor dengan AI Modules Terintegrasi Weaviate adalah database vektor open-source yang dirancang untu...

    Milvus Tutorial: Distributed Vector Database for AI

    Tutorial 10: Milvus - Database Vektor Terdistribusi untuk AI Daftar Isi Pendahuluan Prasyarat Arsitektur Milvus [Instala...

    Complete Qdrant Tutorial: Vector Database for AI Applications

    Tutorial Lengkap Qdrant: Vector Database untuk Aplikasi AI Qdrant adalah vector database performa tinggi yang dirancang ...