Complete LlamaIndex Tutorial: Building RAG Applications with LLMs

# Tutorial Lengkap LlamaIndex: Membangun Aplikasi RAG dengan LLM LlamaIndex adalah framework data yang powerful untuk membangun aplikasi berbasis LLM. Library ini menyediakan tools untuk mengimpor, m...

By Ruby Abdullah · · tutorial
LlamaIndexRAGLLMVector DatabasePythonAI

Complete LlamaIndex Tutorial: Building RAG Applications with LLMs

LlamaIndex is a powerful data framework for building LLM-powered applications. It provides tools to ingest, structure, and access private or domain-specific data, making it perfect for building Retrieval-Augmented Generation (RAG) systems.

Why LlamaIndex?

LlamaIndex Advantages:
  • Easy data ingestion: Connect to 100+ data sources
  • Flexible indexing: Multiple index types for different use cases
  • Query engines: Natural language querying over your data
  • Agent capabilities: Build autonomous LLM agents
  • Production ready: Scalable and observable

Use Cases:
  • Question answering over documents
  • Chatbots with knowledge bases
  • Document summarization
  • Semantic search
  • Data analysis agents

Installation

pip install llama-index

With OpenAI

pip install llama-index-llms-openai llama-index-embeddings-openai

With local models

pip install llama-index-llms-ollama llama-index-embeddings-huggingface

Verify

python -c "import llamaindex; print(llamaindex.version)"

Quick Start

1. Basic RAG Pipeline

from llamaindex.core import VectorStoreIndex, SimpleDirectoryReader

from llamaindex.llms.openai import OpenAI

import os

os.environ["OPENAIAPIKEY"] = "your-api-key"

Load documents

documents = SimpleDirectoryReader("./data").loaddata()

Create index

index = VectorStoreIndex.fromdocuments(documents)

Query

queryengine = index.asqueryengine()

response = queryengine.query("What is the main topic of these documents?")

print(response)

2. With Custom LLM

from llamaindex.core import VectorStoreIndex, SimpleDirectoryReader, Settings

from llamaindex.llms.openai import OpenAI

from llamaindex.embeddings.openai import OpenAIEmbedding

Configure settings

Settings.llm = OpenAI(model="gpt-4", temperature=0.1)

Settings.embedmodel = OpenAIEmbedding(model="text-embedding-3-small")

Load and index

documents = SimpleDirectoryReader("./data").loaddata()

index = VectorStoreIndex.fromdocuments(documents)

Query with streaming

queryengine = index.asqueryengine(streaming=True)

response = queryengine.query("Summarize the key points")

for token in response.responsegen:

print(token, end="", flush=True)

3. With Local Models (Ollama)

from llamaindex.core import VectorStoreIndex, SimpleDirectoryReader, Settings

from llamaindex.llms.ollama import Ollama

from llamaindex.embeddings.huggingface import HuggingFaceEmbedding

Use local models

Settings.llm = Ollama(model="llama2", requesttimeout=300.0)

Settings.embedmodel = HuggingFaceEmbedding(modelname="BAAI/bge-small-en-v1.5")

Build RAG

documents = SimpleDirectoryReader("./data").loaddata()

index = VectorStoreIndex.fromdocuments(documents)

queryengine = index.asqueryengine()

response = queryengine.query("What does this document discuss?")

print(response)

Data Loading

1. File Readers

from llamaindex.core import SimpleDirectoryReader

Load from directory

documents = SimpleDirectoryReader(

inputdir="./data",

recursive=True,

requiredexts=[".pdf", ".txt", ".md"]

).loaddata()

Load specific files

documents = SimpleDirectoryReader(

inputfiles=["doc1.pdf", "doc2.txt"]

).loaddata()

With metadata

documents = SimpleDirectoryReader(

"./data",

filemetadata=lambda filename: {"source": filename}

).loaddata()

2. Web Readers

from llamaindex.readers.web import SimpleWebPageReader, BeautifulSoupWebReader

Simple web reader

reader = SimpleWebPageReader()

documents = reader.loaddata(["https://example.com/page1", "https://example.com/page2"])

BeautifulSoup reader

reader = BeautifulSoupWebReader()

documents = reader.loaddata(

urls=["https://example.com"],

customhostname="example.com"

)

3. Database Readers

from llamaindex.readers.database import DatabaseReader

SQL database

reader = DatabaseReader(

sqldatabase=SQLDatabase.fromuri("postgresql://user:pass@localhost/db")

)

documents = reader.loaddata(query="SELECT FROM articles")

4. API Readers

from llamaindex.readers.notion import NotionPageReader

from llamaindex.readers.slack import SlackReader

Notion

notionreader = NotionPageReader(integrationtoken="secretxxx")

documents = notionreader.loaddata(pageids=["pageid1", "pageid2"])

Slack

slackreader = SlackReader(slacktoken="xoxb-xxx")

documents = slackreader.loaddata(channelids=["C01234567"])

Index Types

1. VectorStoreIndex

from llamaindex.core import VectorStoreIndex, StorageContext

from llamaindex.vectorstores.chroma import ChromaVectorStore

import chromadb

In-memory

index = VectorStoreIndex.fromdocuments(documents)

With ChromaDB persistence

chromaclient = chromadb.PersistentClient(path="./chromadb")

chromacollection = chromaclient.getorcreatecollection("mycollection")

vectorstore = ChromaVectorStore(chromacollection=chromacollection)

storagecontext = StorageContext.fromdefaults(vectorstore=vectorstore)

index = VectorStoreIndex.fromdocuments(

documents,

storagecontext=storagecontext

)

2. SummaryIndex

from llamaindex.core import SummaryIndex

Good for summarization tasks

index = SummaryIndex.fromdocuments(documents)

queryengine = index.asqueryengine(

responsemode="treesummarize"

)

response = queryengine.query("Provide a comprehensive summary")

3. KeywordTableIndex

from llamaindex.core import KeywordTableIndex

Keyword-based retrieval

index = KeywordTableIndex.fromdocuments(documents)

queryengine = index.asqueryengine()

response = queryengine.query("Find documents about machine learning")

4. KnowledgeGraphIndex

from llamaindex.core import KnowledgeGraphIndex

Build knowledge graph

index = KnowledgeGraphIndex.fromdocuments(

documents,

maxtripletsperchunk=10,

includeembeddings=True

)

queryengine = index.asqueryengine(

includetext=True,

responsemode="treesummarize"

)

Query Engines

1. Basic Query Engine

from llamaindex.core import VectorStoreIndex

index = VectorStoreIndex.fromdocuments(documents)

Default query engine

queryengine = index.asqueryengine()

With parameters

queryengine = index.asqueryengine(

similaritytopk=5,

responsemode="compact",

streaming=True

)

response = queryengine.query("What are the main findings?")

2. Response Modes

# Refine: iteratively refine answer

queryengine = index.asqueryengine(responsemode="refine")

Compact: compact context before answering

queryengine = index.asqueryengine(responsemode="compact")

Tree summarize: build tree and summarize

queryengine = index.asqueryengine(responsemode="treesummarize")

Simple summarize: truncate and summarize

queryengine = index.asqueryengine(responsemode="simplesummarize")

No text: return retrieved nodes only

queryengine = index.asqueryengine(responsemode="notext")

3. Custom Query Engine

from llamaindex.core.queryengine import RetrieverQueryEngine

from llamaindex.core.retrievers import VectorIndexRetriever

from llamaindex.core.responsesynthesizers import getresponsesynthesizer

Custom retriever

retriever = VectorIndexRetriever(

index=index,

similaritytopk=10

)

Custom response synthesizer

responsesynthesizer = getresponsesynthesizer(

responsemode="treesummarize"

)

Custom query engine

queryengine = RetrieverQueryEngine(

retriever=retriever,

responsesynthesizer=responsesynthesizer

)

Chat Engines

1. Basic Chat Engine

from llamaindex.core import VectorStoreIndex

index = VectorStoreIndex.fromdocuments(documents)

Create chat engine

chatengine = index.aschatengine(

chatmode="condensequestion",

verbose=True

)

Chat

response = chatengine.chat("What is this document about?")

print(response)

response = chatengine.chat("Can you elaborate on that?")

print(response)

Reset conversation

chatengine.reset()

2. Chat Modes

# Condense question: reformulate question with context

chatengine = index.aschatengine(chatmode="condensequestion")

Context: always use context from index

chatengine = index.aschatengine(chatmode="context")

React: use ReAct agent

chatengine = index.aschatengine(chatmode="react")

OpenAI: use OpenAI function calling

chatengine = index.aschatengine(chatmode="openai")

3. Streaming Chat

chatengine = index.aschatengine(streaming=True)

response = chatengine.streamchat("Tell me about the main topics")

for token in response.responsegen:

print(token, end="", flush=True)

Agents

1. ReAct Agent

from llamaindex.core.agent import ReActAgent

from llamaindex.core.tools import QueryEngineTool, ToolMetadata

Create tools

queryenginetool = QueryEngineTool(

queryengine=index.asqueryengine(),

metadata=ToolMetadata(

name="documentsearch",

description="Search through documents for information"

)

)

Create agent

agent = ReActAgent.fromtools(

[queryenginetool],

llm=llm,

verbose=True

)

response = agent.chat("Find information about machine learning")

2. OpenAI Agent

from llamaindex.agent.openai import OpenAIAgent

from llamaindex.core.tools import FunctionTool

Define custom function

def multiply(a: int, b: int) -> int:

"""Multiply two numbers."""

return a b

def add(a: int, b: int) -> int:

"""Add two numbers."""

return a + b

Create tools

multiplytool = FunctionTool.fromdefaults(fn=multiply)

addtool = FunctionTool.fromdefaults(fn=add)

Create agent

agent = OpenAIAgent.fromtools(

[multiplytool, addtool, queryenginetool],

verbose=True

)

response = agent.chat("What is 5 * 3 + 2?")

3. Multi-Document Agent

from llamaindex.core import VectorStoreIndex, SummaryIndex

from llamaindex.core.tools import QueryEngineTool

Create indices for each document

docagents = {}

for doc in documents:

vectorindex = VectorStoreIndex.fromdocuments([doc])

summaryindex = SummaryIndex.fromdocuments([doc])

vectortool = QueryEngineTool(

queryengine=vectorindex.asqueryengine(),

metadata=ToolMetadata(

name=f"vector{doc.docid}",

description=f"Search {doc.docid}"

)

)

summarytool = QueryEngineTool(

queryengine=summaryindex.asqueryengine(),

metadata=ToolMetadata(

name=f"summary{doc.docid}",

description=f"Summarize {doc.docid}"

)

)

docagents[doc.docid] = [vectortool, summarytool]

Create top-level agent

alltools = [tool for tools in docagents.values() for tool in tools]

agent = ReActAgent.fromtools(alltools, verbose=True)

Advanced Features

1. Node Postprocessors

from llamaindex.core.postprocessor import (

SimilarityPostprocessor,

KeywordNodePostprocessor,

MetadataReplacementPostProcessor

)

Filter by similarity

similarityprocessor = SimilarityPostprocessor(similaritycutoff=0.7)

Filter by keywords

keywordprocessor = KeywordNodePostprocessor(

requiredkeywords=["machine learning"],

excludekeywords=["deprecated"]

)

Apply to query engine

queryengine = index.asqueryengine(

nodepostprocessors=[similarityprocessor, keywordprocessor]

)

from llamaindex.core.retrievers import (

VectorIndexRetriever,

KeywordTableSimpleRetriever

)

from llamaindex.core.queryengine import RetrieverQueryEngine

Vector retriever

vectorretriever = VectorIndexRetriever(index=vectorindex, similaritytopk=5)

Keyword retriever

keywordretriever = KeywordTableSimpleRetriever(index=keywordindex)

Combine with custom logic

class HybridRetriever:

def init(self, vectorretriever, keywordretriever):

self.vectorretriever = vectorretriever

self.keywordretriever = keywordretriever

def retrieve(self, query):

vectornodes = self.vectorretriever.retrieve(query)

keywordnodes = self.keywordretriever.retrieve(query)

# Combine and deduplicate

allnodes = {n.node.nodeid: n for n in vectornodes + keywordnodes}

return list(allnodes.values())

hybridretriever = HybridRetriever(vectorretriever, keywordretriever)

3. Query Transformations

from llamaindex.core.queryengine import TransformQueryEngine

from llamaindex.core.indices.query.querytransform import HyDEQueryTransform

HyDE (Hypothetical Document Embeddings)

hyde = HyDEQueryTransform(includeoriginal=True)

queryengine = TransformQueryEngine(

queryengine=index.asqueryengine(),

querytransform=hyde

)

response = queryengine.query("What are the benefits?")

Evaluation

1. Response Evaluation

from llamaindex.core.evaluation import (

FaithfulnessEvaluator,

RelevancyEvaluator,

CorrectnessEvaluator

)

Create evaluators

faithfulnessevaluator = FaithfulnessEvaluator()

relevancyevaluator = RelevancyEvaluator()

Evaluate response

query = "What is machine learning?"

response = queryengine.query(query)

faithfulnessresult = faithfulnessevaluator.evaluateresponse(response=response)

relevancyresult = relevancyevaluator.evaluateresponse(

query=query,

response=response

)

print(f"Faithfulness: {faithfulnessresult.passing}")

print(f"Relevancy: {relevancyresult.passing}")

2. Retrieval Evaluation

from llamaindex.core.evaluation import RetrieverEvaluator

Create dataset

evalquestions = [

"What is the main topic?",

"Who are the authors?",

"What methodology was used?"

]

Evaluate retriever

retriever = index.asretriever(similaritytopk=5)

evaluator = RetrieverEvaluator.frommetricnames(

["mrr", "hitrate"],

retriever=retriever

)

Run evaluation

results = await evaluator.aevaluatedataset(evalquestions)

Persistence

1. Save and Load Index

from llamaindex.core import VectorStoreIndex, StorageContext, loadindexfromstorage

Save index

index.storagecontext.persist(persistdir="./storage")

Load index

storagecontext = StorageContext.fromdefaults(persistdir="./storage")

index = loadindexfromstorage(storagecontext)

2. With Vector Store

from llamaindex.vectorstores.qdrant import QdrantVectorStore

from qdrantclient import QdrantClient

Create Qdrant client

client = QdrantClient(path="./qdrantdata")

Create vector store

vectorstore = QdrantVectorStore(

client=client,

collectionname="mycollection"

)

Build index with vector store

storagecontext = StorageContext.fromdefaults(vectorstore=vectorstore)

index = VectorStoreIndex.fromdocuments(

documents,

storagecontext=storagecontext

)

Load existing index

index = VectorStoreIndex.fromvectorstore(vectorstore)

Best Practices

1. Chunking Strategy

from llamaindex.core.nodeparser import (

SentenceSplitter,

SemanticSplitterNodeParser

)

Sentence splitter

splitter = SentenceSplitter(

chunksize=1024,

chunkoverlap=200

)

Semantic splitter

semanticsplitter = SemanticSplitterNodeParser(

buffersize=1,

breakpointpercentilethreshold=95,

embedmodel=embedmodel

)

nodes = splitter.getnodesfromdocuments(documents)

2. Metadata Filtering

from llamaindex.core.vectorstores import MetadataFilters, MetadataFilter

Add metadata to documents

for doc in documents:

doc.metadata["category"] = "technical"

doc.metadata["year"] = 2024

Filter by metadata

filters = MetadataFilters(

filters=[

MetadataFilter(key="category", value="technical"),

MetadataFilter(key="year", value=2024)

]

)

queryengine = index.asqueryengine(filters=filters)

Conclusion

LlamaIndex is essential for building RAG applications with:

  • Easy data ingestion: 100+ data connectors
  • Flexible indexing: Multiple index types
  • Powerful querying: Query engines and chat engines
  • Agent capabilities: ReAct and OpenAI agents
  • Production ready: Persistence and evaluation
  • Key takeaways:

    • Start with VectorStoreIndex for most use cases
    • Use chat engines for conversational applications
    • Leverage agents for complex multi-step tasks
    • Evaluate and iterate on retrieval quality
    • Persist indices for production deployment

    Related Articles

    Complete Qdrant Tutorial: Vector Database for AI Applications

    Tutorial Lengkap Qdrant: Vector Database untuk Aplikasi AI Qdrant adalah vector database performa tinggi yang dirancang ...

    Complete ChromaDB Tutorial: Simple Vector Database for AI

    Tutorial Lengkap ChromaDB: Vector Database Sederhana untuk AI ChromaDB adalah open-source vector database yang dirancang...

    DSPy: Stop Hand-Tuning Prompts, Let the Compiler Optimize Them

    DSPy: Berhenti Ngoprek Prompt Manual, Biarkan Compiler yang Optimasi Halo temen-temen, kali ini aku mau ngenalin satu li...

    Complete Pinecone Tutorial: Vector Database for AI and Semantic Search

    Tutorial Lengkap Pinecone: Vector Database untuk AI dan Semantic Search Pinecone adalah managed vector database yang dir...