RAGAS: Evaluation Framework for RAG Pipelines
Introduction
Retrieval-Augmented Generation (RAG) has become the standard architecture for building knowledge-grounded LLM applications. However, building a RAG pipeline alone is not enough. We need a systematic way to measure the quality of the pipeline's output. This is where RAGAS (Retrieval Augmented Generation Assessment) comes in.
RAGAS is an open-source framework specifically designed for automatically evaluating RAG pipelines. It provides a set of metrics that measure various quality aspects, from the relevance of retrieved context to the faithfulness of answers to source documents.
In this tutorial, we will learn how to install and use RAGAS, understand its core metrics, create evaluation datasets, integrate it with LangChain, and apply it to a practical example of evaluating and improving a document QA system.
Installation
Install RAGAS and required dependencies using pip:
pip install ragas langchain langchain-openai chromadb
Make sure you have your OpenAI API key configured as an environment variable:
export OPENAIAPIKEY="sk-your-api-key-here"
RAGAS uses an LLM (OpenAI by default) to compute some of its metrics, so this API key is required for evaluation.
Core RAGAS Metrics
RAGAS provides four main metrics, each measuring a different aspect of the RAG pipeline:
1. Faithfulness
Faithfulness measures how consistent the generated answer is with the provided context. This metric ensures that the LLM does not "hallucinate" or add information that is not present in the source documents.
from ragas.metrics import faithfulness
Score 1.0 = all claims in the answer are supported by the context
Score 0.0 = no claims are supported by the context
Faithfulness is calculated by:
- Breaking down the answer into individual claims
- Verifying each claim against the provided context
- Computing the ratio of supported claims to total claims
2. Answer Relevancy
Answer Relevancy measures how relevant the generated answer is to the question asked. Answers that are complete and directly address the question receive higher scores.
from ragas.metrics import answerrelevancy
High score = answer is highly relevant to the question
Low score = answer is irrelevant or too generic
3. Context Precision
Context Precision measures how precise the retrieved context is. This metric evaluates whether relevant document chunks are ranked higher than irrelevant ones.
from ragas.metrics import contextprecision
High score = relevant context appears at the top
Low score = relevant context is mixed with irrelevant ones
4. Context Recall
Context Recall measures how much of the information needed to answer the question was successfully retrieved from the documents. This metric requires ground truth (reference answers) for computation.
from ragas.metrics import contextrecall
High score = all necessary information was retrieved
Low score = important information was missed
Creating Evaluation Datasets
To run a RAGAS evaluation, you need to prepare a dataset in the correct format. The dataset must contain: questions, retrieved contexts, generated answers, and ground truth (optional, but required for context recall).
from datasets import Dataset
questions = [
"What is machine learning?",
"How do neural networks work?",
"What is the difference between supervised and unsupervised learning?"
]
groundtruths = [
"Machine learning is a branch of artificial intelligence that enables systems to learn from data without being explicitly programmed.",
"Neural networks work by mimicking how neurons in the human brain function, using layers of interconnected nodes to process information.",
"Supervised learning uses labeled data for training, while unsupervised learning finds patterns in data without labels."
]
contexts = [
["Machine learning is a subset of artificial intelligence that focuses on developing algorithms that can learn from data. ML systems can automatically improve their performance through experience."],
["Neural networks or artificial neural networks consist of input, hidden, and output layers. Each node processes input using activation functions and weights that are updated during training."],
["In supervised learning, models are trained using known input-output pairs. Unsupervised learning works with unlabeled data, searching for hidden structures like clusters."]
]
answers = [
"Machine learning is a branch of AI that allows computers to learn from data automatically without being explicitly programmed for each task.",
"Neural networks work by processing data through layers of interconnected nodes, similar to neurons in the human brain.",
"Supervised learning uses labeled data while unsupervised learning does not use labels on its data."
]
evaldataset = Dataset.fromdict({
"question": questions,
"groundtruth": groundtruths,
"contexts": contexts,
"answer": answers
})
print(evaldataset)
Running Evaluation
Once the dataset is ready, you can run the evaluation with your chosen metrics:
from ragas import evaluate
from ragas.metrics import (
faithfulness,
answerrelevancy,
contextprecision,
contextrecall
)
result = evaluate(
dataset=evaldataset,
metrics=[
faithfulness,
answerrelevancy,
contextprecision,
contextrecall
]
)
print("Evaluation Results:")
print(f"Faithfulness: {result['faithfulness']:.4f}")
print(f"Answer Relevancy: {result['answerrelevancy']:.4f}")
print(f"Context Precision: {result['contextprecision']:.4f}")
print(f"Context Recall: {result['contextrecall']:.4f}")
To see per-question details:
df = result.topandas()
print(df[['question', 'faithfulness', 'answerrelevancy',
'contextprecision', 'contextrecall']])
Integration with LangChain
RAGAS integrates well with LangChain, enabling direct evaluation of existing RAG chains:
from langchainopenai import ChatOpenAI, OpenAIEmbeddings
from langchaincommunity.vectorstores import Chroma
from langchain.textsplitter import RecursiveCharacterTextSplitter
from langchain.chains import RetrievalQA
from langchaincommunity.documentloaders import TextLoader
Load documents
loader = TextLoader("knowledgebase.txt")
documents = loader.load()
Split documents
textsplitter = RecursiveCharacterTextSplitter(
chunksize=500,
chunkoverlap=50
)
splits = textsplitter.splitdocuments(documents)
Create vector store
embeddings = OpenAIEmbeddings()
vectorstore = Chroma.fromdocuments(splits, embeddings)
Create RAG chain
llm = ChatOpenAI(model="gpt-4o-mini", temperature=0)
qachain = RetrievalQA.fromchaintype(
llm=llm,
chaintype="stuff",
retriever=vectorstore.asretriever(searchkwargs={"k": 3}),
returnsourcedocuments=True
)
Run chain and collect results for evaluation
questions = [
"What is the company's leave policy?",
"How do I submit a reimbursement request?",
"What health benefits are available?"
]
allanswers = []
allcontexts = []
for question in questions:
response = qachain.invoke({"query": question})
allanswers.append(response["result"])
allcontexts.append(
[doc.pagecontent for doc in response["sourcedocuments"]]
)
Create evaluation dataset from chain results
evaldataset = Dataset.fromdict({
"question": questions,
"answer": allanswers,
"contexts": allcontexts,
"groundtruth": groundtruths # Prepared beforehand
})
Evaluate
result = evaluate(
dataset=evaldataset,
metrics=[faithfulness, answerrelevancy, contextprecision, contextrecall]
)
print("LangChain RAG Chain Evaluation:")
print(result)
Custom Metrics
RAGAS allows you to create custom metrics tailored to your needs:
from ragas.metrics.base import MetricWithLLM
from dataclasses import dataclass, field
@dataclass
class ToneConsistencyMetric(MetricWithLLM):
name: str = "toneconsistency"
evaluationmode: str = "qa"
def init(self, runconfig=None):
pass
async def ascore(self, row, callbacks=None):
"""Evaluate the tone consistency of the answer."""
question = row["question"]
answer = row["answer"]
prompt = f"""Evaluate whether the following answer uses a professional
and consistent tone.
Question: {question}
Answer: {answer}
Provide a score from 0-1 where 1 is the most professional tone.
Reply with only a number."""
response = await self.llm.ageneratetext(prompt)
try:
score = float(response.generations[0][0].text.strip())
return min(max(score, 0), 1)
except ValueError:
return 0.5
Use custom metric
tonemetric = ToneConsistencyMetric()
result = evaluate(
dataset=evaldataset,
metrics=[faithfulness, answerrelevancy, tonemetric]
)
Automatic Test Set Generation
RAGAS provides a feature to automatically generate test sets from your documents:
from ragas.testset.generator import TestsetGenerator
from ragas.testset.evolutions import simple, reasoning, multicontext
from langchainopenai import ChatOpenAI, OpenAIEmbeddings
from langchaincommunity.documentloaders import DirectoryLoader
Load documents
loader = DirectoryLoader("./docs/", glob="*/.txt")
documents = loader.load()
Initialize generator
generatorllm = ChatOpenAI(model="gpt-4o-mini")
criticllm = ChatOpenAI(model="gpt-4o")
embeddings = OpenAIEmbeddings()
generator = TestsetGenerator.fromlangchain(
generatorllm=generatorllm,
criticllm=criticllm,
embeddings=embeddings
)
Generate test set
testset = generator.generatewithlangchaindocs(
documents=documents,
testsize=20,
distributions={
simple: 0.5, # Simple factual questions
reasoning: 0.25, # Questions requiring reasoning
multicontext: 0.25 # Multi-context questions
}
)
testdf = testset.topandas()
print(f"Generated {len(testdf)} test questions")
print(testdf[['question', 'evolutiontype']].head(10))
Practical Example: Evaluating and Improving a Document QA System
Let's build an end-to-end example where we evaluate a document QA system and improve it based on RAGAS metrics.
Step 1: Build a Baseline QA System
import os
from langchainopenai import ChatOpenAI, OpenAIEmbeddings
from langchaincommunity.vectorstores import Chroma
from langchain.textsplitter import RecursiveCharacterTextSplitter
from langchain.chains import RetrievalQA
Simulated knowledge base
knowledgebase = """
AI Assistant Pro Product Guide
Key Features
AI Assistant Pro is an enterprise AI platform supporting natural language
processing, sentiment analysis, and information extraction. The platform
supports 50+ languages and can be integrated via REST API.
Pricing
Starter Plan: $50/month for 10,000 requests.
Business Plan: $200/month for 100,000 requests.
Enterprise Plan: Custom pricing for unlimited requests.
Integrations
The platform supports integration with Slack, Microsoft Teams, WhatsApp,
and Telegram. SDKs are available for Python, JavaScript, and Java.
Security
Data is encrypted using AES-256 at rest and TLS 1.3 in transit.
The platform is compliant with ISO 27001 and SOC 2 Type II.
Supports SSO via SAML 2.0 and OAuth 2.0.
Support
24/7 support is available for Business and Enterprise plans.
Starter plan gets email support with a 48-hour SLA.
Full documentation is available at docs.aiassistantpro.com.
"""
Save knowledge base
with open("/tmp/knowledgebase.txt", "w") as f:
f.write(knowledgebase)
Setup RAG pipeline v1 (baseline)
from langchaincommunity.documentloaders import TextLoader
loader = TextLoader("/tmp/knowledgebase.txt")
documents = loader.load()
Aggressive chunking (small chunks)
splitterv1 = RecursiveCharacterTextSplitter(
chunksize=200,
chunkoverlap=20
)
splitsv1 = splitterv1.splitdocuments(documents)
embeddings = OpenAIEmbeddings()
vectorstorev1 = Chroma.fromdocuments(splitsv1, embeddings,
collectionname="v1")
qav1 = RetrievalQA.fromchaintype(
llm=ChatOpenAI(model="gpt-4o-mini", temperature=0),
retriever=vectorstorev1.asretriever(searchkwargs={"k": 2}),
returnsourcedocuments=True
)
Step 2: Prepare Evaluation Dataset
from datasets import Dataset
evalquestions = [
"How much does the Business plan cost?",
"What security features are available?",
"What integrations are supported?",
"What is the support SLA for the Starter plan?",
"What programming languages are supported by the SDK?"
]
evalgroundtruths = [
"The Business plan costs $200/month for 100,000 requests.",
"Security features include AES-256 encryption at rest, TLS 1.3 in transit, ISO 27001 and SOC 2 Type II compliance, and SSO via SAML 2.0 and OAuth 2.0.",
"The platform supports integration with Slack, Microsoft Teams, WhatsApp, and Telegram.",
"The Starter plan gets email support with a 48-hour SLA.",
"SDKs are available for Python, JavaScript, and Java."
]
Step 3: Evaluate Version 1
from ragas import evaluate
from ragas.metrics import (
faithfulness, answerrelevancy,
contextprecision, contextrecall
)
Collect answers from pipeline v1
answersv1 = []
contextsv1 = []
for q in evalquestions:
resp = qav1.invoke({"query": q})
answersv1.append(resp["result"])
contextsv1.append([d.pagecontent for d in resp["sourcedocuments"]])
datasetv1 = Dataset.fromdict({
"question": evalquestions,
"answer": answersv1,
"contexts": contextsv1,
"groundtruth": evalgroundtruths
})
resultv1 = evaluate(
dataset=datasetv1,
metrics=[faithfulness, answerrelevancy, contextprecision, contextrecall]
)
print("=== V1 Evaluation Results (Baseline) ===")
print(f"Faithfulness: {resultv1['faithfulness']:.4f}")
print(f"Answer Relevancy: {resultv1['answerrelevancy']:.4f}")
print(f"Context Precision: {resultv1['contextprecision']:.4f}")
print(f"Context Recall: {resultv1['contextrecall']:.4f}")
Step 4: Identify Issues and Improve
Based on the evaluation results, if context recall is low, it means chunks that are too small are breaking apart important information. Let's fix this:
# Pipeline V2: Larger chunks + more overlap + more context retrieved
splitterv2 = RecursiveCharacterTextSplitter(
chunksize=500, # Larger chunks
chunkoverlap=100 # More overlap
)
splitsv2 = splitterv2.splitdocuments(documents)
vectorstorev2 = Chroma.fromdocuments(splitsv2, embeddings,
collectionname="v2")
qav2 = RetrievalQA.fromchaintype(
llm=ChatOpenAI(model="gpt-4o-mini", temperature=0),
retriever=vectorstorev2.asretriever(searchkwargs={"k": 4}), # Retrieve more
returnsourcedocuments=True
)
Evaluate V2
answersv2 = []
contextsv2 = []
for q in evalquestions:
resp = qav2.invoke({"query": q})
answersv2.append(resp["result"])
contextsv2.append([d.pagecontent for d in resp["sourcedocuments"]])
datasetv2 = Dataset.fromdict({
"question": evalquestions,
"answer": answersv2,
"contexts": contextsv2,
"groundtruth": evalgroundtruths
})
resultv2 = evaluate(
dataset=datasetv2,
metrics=[faithfulness, answerrelevancy, contextprecision, contextrecall]
)
print("\n=== V2 Evaluation Results (Improved) ===")
print(f"Faithfulness: {resultv2['faithfulness']:.4f}")
print(f"Answer Relevancy: {resultv2['answerrelevancy']:.4f}")
print(f"Context Precision: {resultv2['contextprecision']:.4f}")
print(f"Context Recall: {resultv2['contextrecall']:.4f}")
Step 5: Compare Results
import pandas as pd
comparison = pd.DataFrame({
"Metric": ["Faithfulness", "Answer Relevancy",
"Context Precision", "Context Recall"],
"V1 (Baseline)": [
resultv1['faithfulness'],
resultv1['answerrelevancy'],
resultv1['contextprecision'],
resultv1['contextrecall']
],
"V2 (Improved)": [
resultv2['faithfulness'],
resultv2['answerrelevancy'],
resultv2['contextprecision'],
resultv2['contextrecall']
]
})
comparison["Improvement"] = comparison["V2 (Improved)"] - comparison["V1 (Baseline)"]
print("\n=== V1 vs V2 Comparison ===")
print(comparison.tostring(index=False))
Tips and Best Practices
Conclusion
RAGAS provides a comprehensive and easy-to-use framework for evaluating RAG pipelines. With metrics like faithfulness, answer relevancy, context precision, and context recall, you can identify specific weaknesses in your pipeline and make measurable improvements.
The key to effective evaluation is consistency. Run evaluations regularly, use representative datasets, and always compare results against your baseline. With this approach, you can build RAG pipelines that continuously improve over time.
For more information, visit the official RAGAS documentation at docs.ragas.io.