RAGAS: Evaluation Framework for RAG Pipelines

# RAGAS: Framework Evaluasi untuk Pipeline RAG ## Pendahuluan Retrieval-Augmented Generation (RAG) telah menjadi arsitektur standar untuk membangun aplikasi LLM yang berbasis pengetahuan. Namun, mem...

By Ruby Abdullah · · tutorial
RAGASRAGEvaluationLLMPython

RAGAS: Evaluation Framework for RAG Pipelines

Introduction

Retrieval-Augmented Generation (RAG) has become the standard architecture for building knowledge-grounded LLM applications. However, building a RAG pipeline alone is not enough. We need a systematic way to measure the quality of the pipeline's output. This is where RAGAS (Retrieval Augmented Generation Assessment) comes in.

RAGAS is an open-source framework specifically designed for automatically evaluating RAG pipelines. It provides a set of metrics that measure various quality aspects, from the relevance of retrieved context to the faithfulness of answers to source documents.

In this tutorial, we will learn how to install and use RAGAS, understand its core metrics, create evaluation datasets, integrate it with LangChain, and apply it to a practical example of evaluating and improving a document QA system.

Installation

Install RAGAS and required dependencies using pip:

pip install ragas langchain langchain-openai chromadb

Make sure you have your OpenAI API key configured as an environment variable:

export OPENAIAPIKEY="sk-your-api-key-here"

RAGAS uses an LLM (OpenAI by default) to compute some of its metrics, so this API key is required for evaluation.

Core RAGAS Metrics

RAGAS provides four main metrics, each measuring a different aspect of the RAG pipeline:

1. Faithfulness

Faithfulness measures how consistent the generated answer is with the provided context. This metric ensures that the LLM does not "hallucinate" or add information that is not present in the source documents.

from ragas.metrics import faithfulness

Score 1.0 = all claims in the answer are supported by the context

Score 0.0 = no claims are supported by the context

Faithfulness is calculated by:

  • Breaking down the answer into individual claims
  • Verifying each claim against the provided context
  • Computing the ratio of supported claims to total claims

2. Answer Relevancy

Answer Relevancy measures how relevant the generated answer is to the question asked. Answers that are complete and directly address the question receive higher scores.

from ragas.metrics import answerrelevancy

High score = answer is highly relevant to the question

Low score = answer is irrelevant or too generic

3. Context Precision

Context Precision measures how precise the retrieved context is. This metric evaluates whether relevant document chunks are ranked higher than irrelevant ones.

from ragas.metrics import contextprecision

High score = relevant context appears at the top

Low score = relevant context is mixed with irrelevant ones

4. Context Recall

Context Recall measures how much of the information needed to answer the question was successfully retrieved from the documents. This metric requires ground truth (reference answers) for computation.

from ragas.metrics import contextrecall

High score = all necessary information was retrieved

Low score = important information was missed

Creating Evaluation Datasets

To run a RAGAS evaluation, you need to prepare a dataset in the correct format. The dataset must contain: questions, retrieved contexts, generated answers, and ground truth (optional, but required for context recall).

from datasets import Dataset

questions = [

"What is machine learning?",

"How do neural networks work?",

"What is the difference between supervised and unsupervised learning?"

]

groundtruths = [

"Machine learning is a branch of artificial intelligence that enables systems to learn from data without being explicitly programmed.",

"Neural networks work by mimicking how neurons in the human brain function, using layers of interconnected nodes to process information.",

"Supervised learning uses labeled data for training, while unsupervised learning finds patterns in data without labels."

]

contexts = [

["Machine learning is a subset of artificial intelligence that focuses on developing algorithms that can learn from data. ML systems can automatically improve their performance through experience."],

["Neural networks or artificial neural networks consist of input, hidden, and output layers. Each node processes input using activation functions and weights that are updated during training."],

["In supervised learning, models are trained using known input-output pairs. Unsupervised learning works with unlabeled data, searching for hidden structures like clusters."]

]

answers = [

"Machine learning is a branch of AI that allows computers to learn from data automatically without being explicitly programmed for each task.",

"Neural networks work by processing data through layers of interconnected nodes, similar to neurons in the human brain.",

"Supervised learning uses labeled data while unsupervised learning does not use labels on its data."

]

evaldataset = Dataset.fromdict({

"question": questions,

"groundtruth": groundtruths,

"contexts": contexts,

"answer": answers

})

print(evaldataset)

Running Evaluation

Once the dataset is ready, you can run the evaluation with your chosen metrics:

from ragas import evaluate

from ragas.metrics import (

faithfulness,

answerrelevancy,

contextprecision,

contextrecall

)

result = evaluate(

dataset=evaldataset,

metrics=[

faithfulness,

answerrelevancy,

contextprecision,

contextrecall

]

)

print("Evaluation Results:")

print(f"Faithfulness: {result['faithfulness']:.4f}")

print(f"Answer Relevancy: {result['answerrelevancy']:.4f}")

print(f"Context Precision: {result['contextprecision']:.4f}")

print(f"Context Recall: {result['contextrecall']:.4f}")

To see per-question details:

df = result.topandas()

print(df[['question', 'faithfulness', 'answerrelevancy',

'contextprecision', 'contextrecall']])

Integration with LangChain

RAGAS integrates well with LangChain, enabling direct evaluation of existing RAG chains:

from langchainopenai import ChatOpenAI, OpenAIEmbeddings

from langchaincommunity.vectorstores import Chroma

from langchain.textsplitter import RecursiveCharacterTextSplitter

from langchain.chains import RetrievalQA

from langchaincommunity.documentloaders import TextLoader

Load documents

loader = TextLoader("knowledgebase.txt")

documents = loader.load()

Split documents

textsplitter = RecursiveCharacterTextSplitter(

chunksize=500,

chunkoverlap=50

)

splits = textsplitter.splitdocuments(documents)

Create vector store

embeddings = OpenAIEmbeddings()

vectorstore = Chroma.fromdocuments(splits, embeddings)

Create RAG chain

llm = ChatOpenAI(model="gpt-4o-mini", temperature=0)

qachain = RetrievalQA.fromchaintype(

llm=llm,

chaintype="stuff",

retriever=vectorstore.asretriever(searchkwargs={"k": 3}),

returnsourcedocuments=True

)

Run chain and collect results for evaluation

questions = [

"What is the company's leave policy?",

"How do I submit a reimbursement request?",

"What health benefits are available?"

]

allanswers = []

allcontexts = []

for question in questions:

response = qachain.invoke({"query": question})

allanswers.append(response["result"])

allcontexts.append(

[doc.pagecontent for doc in response["sourcedocuments"]]

)

Create evaluation dataset from chain results

evaldataset = Dataset.fromdict({

"question": questions,

"answer": allanswers,

"contexts": allcontexts,

"groundtruth": groundtruths # Prepared beforehand

})

Evaluate

result = evaluate(

dataset=evaldataset,

metrics=[faithfulness, answerrelevancy, contextprecision, contextrecall]

)

print("LangChain RAG Chain Evaluation:")

print(result)

Custom Metrics

RAGAS allows you to create custom metrics tailored to your needs:

from ragas.metrics.base import MetricWithLLM

from dataclasses import dataclass, field

@dataclass

class ToneConsistencyMetric(MetricWithLLM):

name: str = "toneconsistency"

evaluationmode: str = "qa"

def init(self, runconfig=None):

pass

async def ascore(self, row, callbacks=None):

"""Evaluate the tone consistency of the answer."""

question = row["question"]

answer = row["answer"]

prompt = f"""Evaluate whether the following answer uses a professional

and consistent tone.

Question: {question}

Answer: {answer}

Provide a score from 0-1 where 1 is the most professional tone.

Reply with only a number."""

response = await self.llm.ageneratetext(prompt)

try:

score = float(response.generations[0][0].text.strip())

return min(max(score, 0), 1)

except ValueError:

return 0.5

Use custom metric

tonemetric = ToneConsistencyMetric()

result = evaluate(

dataset=evaldataset,

metrics=[faithfulness, answerrelevancy, tonemetric]

)

Automatic Test Set Generation

RAGAS provides a feature to automatically generate test sets from your documents:

from ragas.testset.generator import TestsetGenerator

from ragas.testset.evolutions import simple, reasoning, multicontext

from langchainopenai import ChatOpenAI, OpenAIEmbeddings

from langchaincommunity.documentloaders import DirectoryLoader

Load documents

loader = DirectoryLoader("./docs/", glob="*/.txt")

documents = loader.load()

Initialize generator

generatorllm = ChatOpenAI(model="gpt-4o-mini")

criticllm = ChatOpenAI(model="gpt-4o")

embeddings = OpenAIEmbeddings()

generator = TestsetGenerator.fromlangchain(

generatorllm=generatorllm,

criticllm=criticllm,

embeddings=embeddings

)

Generate test set

testset = generator.generatewithlangchaindocs(

documents=documents,

testsize=20,

distributions={

simple: 0.5, # Simple factual questions

reasoning: 0.25, # Questions requiring reasoning

multicontext: 0.25 # Multi-context questions

}

)

testdf = testset.topandas()

print(f"Generated {len(testdf)} test questions")

print(testdf[['question', 'evolutiontype']].head(10))

Practical Example: Evaluating and Improving a Document QA System

Let's build an end-to-end example where we evaluate a document QA system and improve it based on RAGAS metrics.

Step 1: Build a Baseline QA System

import os

from langchainopenai import ChatOpenAI, OpenAIEmbeddings

from langchaincommunity.vectorstores import Chroma

from langchain.textsplitter import RecursiveCharacterTextSplitter

from langchain.chains import RetrievalQA

Simulated knowledge base

knowledgebase = """

AI Assistant Pro Product Guide

Key Features

AI Assistant Pro is an enterprise AI platform supporting natural language

processing, sentiment analysis, and information extraction. The platform

supports 50+ languages and can be integrated via REST API.

Pricing

Starter Plan: $50/month for 10,000 requests.

Business Plan: $200/month for 100,000 requests.

Enterprise Plan: Custom pricing for unlimited requests.

Integrations

The platform supports integration with Slack, Microsoft Teams, WhatsApp,

and Telegram. SDKs are available for Python, JavaScript, and Java.

Security

Data is encrypted using AES-256 at rest and TLS 1.3 in transit.

The platform is compliant with ISO 27001 and SOC 2 Type II.

Supports SSO via SAML 2.0 and OAuth 2.0.

Support

24/7 support is available for Business and Enterprise plans.

Starter plan gets email support with a 48-hour SLA.

Full documentation is available at docs.aiassistantpro.com.

"""

Save knowledge base

with open("/tmp/knowledgebase.txt", "w") as f:

f.write(knowledgebase)

Setup RAG pipeline v1 (baseline)

from langchaincommunity.documentloaders import TextLoader

loader = TextLoader("/tmp/knowledgebase.txt")

documents = loader.load()

Aggressive chunking (small chunks)

splitterv1 = RecursiveCharacterTextSplitter(

chunksize=200,

chunkoverlap=20

)

splitsv1 = splitterv1.splitdocuments(documents)

embeddings = OpenAIEmbeddings()

vectorstorev1 = Chroma.fromdocuments(splitsv1, embeddings,

collectionname="v1")

qav1 = RetrievalQA.fromchaintype(

llm=ChatOpenAI(model="gpt-4o-mini", temperature=0),

retriever=vectorstorev1.asretriever(searchkwargs={"k": 2}),

returnsourcedocuments=True

)

Step 2: Prepare Evaluation Dataset

from datasets import Dataset

evalquestions = [

"How much does the Business plan cost?",

"What security features are available?",

"What integrations are supported?",

"What is the support SLA for the Starter plan?",

"What programming languages are supported by the SDK?"

]

evalgroundtruths = [

"The Business plan costs $200/month for 100,000 requests.",

"Security features include AES-256 encryption at rest, TLS 1.3 in transit, ISO 27001 and SOC 2 Type II compliance, and SSO via SAML 2.0 and OAuth 2.0.",

"The platform supports integration with Slack, Microsoft Teams, WhatsApp, and Telegram.",

"The Starter plan gets email support with a 48-hour SLA.",

"SDKs are available for Python, JavaScript, and Java."

]

Step 3: Evaluate Version 1

from ragas import evaluate

from ragas.metrics import (

faithfulness, answerrelevancy,

contextprecision, contextrecall

)

Collect answers from pipeline v1

answersv1 = []

contextsv1 = []

for q in evalquestions:

resp = qav1.invoke({"query": q})

answersv1.append(resp["result"])

contextsv1.append([d.pagecontent for d in resp["sourcedocuments"]])

datasetv1 = Dataset.fromdict({

"question": evalquestions,

"answer": answersv1,

"contexts": contextsv1,

"groundtruth": evalgroundtruths

})

resultv1 = evaluate(

dataset=datasetv1,

metrics=[faithfulness, answerrelevancy, contextprecision, contextrecall]

)

print("=== V1 Evaluation Results (Baseline) ===")

print(f"Faithfulness: {resultv1['faithfulness']:.4f}")

print(f"Answer Relevancy: {resultv1['answerrelevancy']:.4f}")

print(f"Context Precision: {resultv1['contextprecision']:.4f}")

print(f"Context Recall: {resultv1['contextrecall']:.4f}")

Step 4: Identify Issues and Improve

Based on the evaluation results, if context recall is low, it means chunks that are too small are breaking apart important information. Let's fix this:

# Pipeline V2: Larger chunks + more overlap + more context retrieved

splitterv2 = RecursiveCharacterTextSplitter(

chunksize=500, # Larger chunks

chunkoverlap=100 # More overlap

)

splitsv2 = splitterv2.splitdocuments(documents)

vectorstorev2 = Chroma.fromdocuments(splitsv2, embeddings,

collectionname="v2")

qav2 = RetrievalQA.fromchaintype(

llm=ChatOpenAI(model="gpt-4o-mini", temperature=0),

retriever=vectorstorev2.asretriever(searchkwargs={"k": 4}), # Retrieve more

returnsourcedocuments=True

)

Evaluate V2

answersv2 = []

contextsv2 = []

for q in evalquestions:

resp = qav2.invoke({"query": q})

answersv2.append(resp["result"])

contextsv2.append([d.pagecontent for d in resp["sourcedocuments"]])

datasetv2 = Dataset.fromdict({

"question": evalquestions,

"answer": answersv2,

"contexts": contextsv2,

"groundtruth": evalgroundtruths

})

resultv2 = evaluate(

dataset=datasetv2,

metrics=[faithfulness, answerrelevancy, contextprecision, contextrecall]

)

print("\n=== V2 Evaluation Results (Improved) ===")

print(f"Faithfulness: {resultv2['faithfulness']:.4f}")

print(f"Answer Relevancy: {resultv2['answerrelevancy']:.4f}")

print(f"Context Precision: {resultv2['contextprecision']:.4f}")

print(f"Context Recall: {resultv2['contextrecall']:.4f}")

Step 5: Compare Results

import pandas as pd

comparison = pd.DataFrame({

"Metric": ["Faithfulness", "Answer Relevancy",

"Context Precision", "Context Recall"],

"V1 (Baseline)": [

resultv1['faithfulness'],

resultv1['answerrelevancy'],

resultv1['contextprecision'],

resultv1['contextrecall']

],

"V2 (Improved)": [

resultv2['faithfulness'],

resultv2['answerrelevancy'],

resultv2['contextprecision'],

resultv2['contextrecall']

]

})

comparison["Improvement"] = comparison["V2 (Improved)"] - comparison["V1 (Baseline)"]

print("\n=== V1 vs V2 Comparison ===")

print(comparison.tostring(index=False))

Tips and Best Practices

  • Evaluate regularly: Run RAGAS evaluations every time you change your RAG pipeline, such as modifying chunking strategy, embedding model, or prompt template.
  • Use diverse test sets: Ensure your evaluation dataset covers various question types, from simple factual queries to questions requiring complex reasoning.
  • Watch for metric trade-offs: Improving one metric can sometimes decrease another. For example, retrieving more context may improve recall but decrease precision.
  • Automate evaluation in CI/CD: Integrate RAGAS evaluation into your CI/CD pipeline to automatically detect quality regressions.
  • Combine with manual evaluation: Automated metrics are great for monitoring, but manual evaluation by domain experts remains important for capturing quality nuances.
  • Conclusion

    RAGAS provides a comprehensive and easy-to-use framework for evaluating RAG pipelines. With metrics like faithfulness, answer relevancy, context precision, and context recall, you can identify specific weaknesses in your pipeline and make measurable improvements.

    The key to effective evaluation is consistency. Run evaluations regularly, use representative datasets, and always compare results against your baseline. With this approach, you can build RAG pipelines that continuously improve over time.

    For more information, visit the official RAGAS documentation at docs.ragas.io.

    Related Articles

    DSPy: Stop Hand-Tuning Prompts, Let the Compiler Optimize Them

    DSPy: Berhenti Ngoprek Prompt Manual, Biarkan Compiler yang Optimasi Halo temen-temen, kali ini aku mau ngenalin satu li...

    Inspect AI: The LLM Evaluation Framework from the UK AI Safety Institute

    Inspect AI: Framework Evaluasi LLM dari UK AI Safety Institute yang Wajib Kamu Coba Temen-temen, kalau kamu udah mulai s...

    Complete Braintrust Tutorial: Evaluate, Test, and Improve Your LLM Applications

    Tutorial Lengkap Braintrust: Evaluasi, Testing, dan Improve Aplikasi LLM Halo temen-temen, di tutorial kali ini aku mau ...

    LangChain Tutorial: The Most Popular Framework for Building LLM Applications

    Tutorial LangChain: Framework Paling Populer untuk Membangun Aplikasi LLM LangChain adalah framework open-source yang di...