LangFuse: Open-Source Platform for LLM Application Observability
As more companies adopt Large Language Models (LLMs) in production applications, the need for observability becomes critical. How do you know the cost of each request? How do you track the performance of different prompts? How do you systematically evaluate the quality of LLM outputs?
LangFuse is an open-source solution that answers all these questions. It provides a comprehensive observability platform for tracing, monitoring, evaluating, and managing prompts in your LLM applications.What Is LangFuse?
LangFuse is an open-source LLM engineering platform that helps development teams with:
- Tracing: Track every step of LLM application execution, from input to output
- Monitoring: Monitor latency, cost, and error rates in real-time
- Evaluation: Score LLM outputs both manually and automatically
- Prompt Management: Centralized prompt versioning and deployment
- Analytics: Dashboards for analyzing usage and performance
LangFuse supports various LLM providers (OpenAI, Anthropic, Azure, etc.) and integrates with popular frameworks like LangChain, LlamaIndex, and Haystack.
Deployment Options
1. LangFuse Cloud (Managed)
The easiest way to get started is using LangFuse Cloud at cloud.langfuse.com.
# No installation needed - just sign up and get API keys
LANGFUSEPUBLICKEY=pk-lf-...
LANGFUSESECRETKEY=sk-lf-...
LANGFUSEHOST=https://cloud.langfuse.com
2. Self-Hosted with Docker
For full control over your data, you can deploy LangFuse yourself:
# docker-compose.yml
version: '3.8'
services:
langfuse:
image: langfuse/langfuse:2
ports:
- "3000:3000"
environment:
- DATABASEURL=postgresql://postgres:postgres@db:5432/langfuse
- NEXTAUTHSECRET=mysecret
- SALT=mysalt
- NEXTAUTHURL=http://localhost:3000
dependson:
- db
db:
image: postgres:15
environment:
- POSTGRES
USER=postgres
POSTGRESPASSWORD=postgres
POSTGRESDB=langfuse
volumes:
- pgdata:/var/lib/postgresql/data
volumes:
pgdata:
# Start LangFuse
docker-compose up -d
Access at http://localhost:3000
Create the first admin account through the UI
Python SDK Installation
pip install langfuse
Environment Configuration
import os
os.environ["LANGFUSEPUBLICKEY"] = "pk-lf-..."
os.environ["LANGFUSESECRETKEY"] = "sk-lf-..."
os.environ["LANGFUSEHOST"] = "https://cloud.langfuse.com" # or self-hosted URL
Client Initialization
from langfuse import Langfuse
langfuse = Langfuse(
publickey="pk-lf-...",
secretkey="sk-lf-...",
host="https://cloud.langfuse.com"
)
Verify connection
langfuse.authcheck()
print("LangFuse connected!")
Tracing: Tracking LLM Application Execution
Tracing is LangFuse's core feature. Every interaction with your LLM application is recorded as a trace consisting of several components:
Core Tracing Concepts
- Trace: The top-level unit representing one end-to-end execution
- Span: A sub-unit representing a single step in the trace (e.g., retrieval, preprocessing)
- Generation: A special span for LLM calls, automatically capturing token usage and cost
Creating a Simple Trace
from langfuse import Langfuse
langfuse = Langfuse()
Create a new trace
trace = langfuse.trace(
name="customer-support-chat",
userid="user-123",
metadata={"session": "abc-456"},
tags=["production", "customer-support"]
)
Add a span for preprocessing step
span = trace.span(
name="preprocess-input",
input={"rawmessage": "How do I reset my password?"}
)
Simulate preprocessing
processedinput = "How do I reset my password?"
span.end(output={"processed": processedinput})
Recording LLM Generations
# Record LLM call as a generation
generation = trace.generation(
name="llm-response",
model="gpt-4o",
modelparameters={
"temperature": 0.7,
"maxtokens": 500
},
input=[
{"role": "system", "content": "You are a customer support assistant."},
{"role": "user", "content": processedinput}
]
)
Simulate LLM response
responsetext = "To reset your password, please follow these steps..."
Update generation with output and usage
generation.end(
output=responsetext,
usage={
"input": 45,
"output": 120,
"unit": "TOKENS"
}
)
IMPORTANT: flush data to server
langfuse.flush()
Tracing with Decorators
LangFuse provides the @observe() decorator for cleaner automatic tracing:
from langfuse.decorators import observe, langfusecontext
@observe()
def processcustomerquery(query: str):
# Step 1: Preprocessing
cleaned = preprocess(query)
# Step 2: Retrieve context
context = retrievedocuments(cleaned)
# Step 3: Generate response
response = generateresponse(cleaned, context)
return response
@observe()
def preprocess(text: str) -> str:
return text.strip().lower()
@observe()
def retrievedocuments(query: str) -> list:
# Simulate vector search
return [
{"title": "Reset Password Guide", "content": "..."},
{"title": "Account Recovery", "content": "..."}
]
@observe(astype="generation")
def generateresponse(query: str, context: list) -> str:
import openai
client = openai.OpenAI()
messages = [
{"role": "system", "content": f"Context: {context}"},
{"role": "user", "content": query}
]
response = client.chat.completions.create(
model="gpt-4o",
messages=messages,
temperature=0.7
)
# Update observation with generation details
langfusecontext.updatecurrentobservation(
model="gpt-4o",
usage={
"input": response.usage.prompttokens,
"output": response.usage.completiontokens
}
)
return response.choices[0].message.content
Run - all steps are automatically traced
result = processcustomerquery("How do I reset my password?")
Prompt Management
LangFuse allows you to manage prompts centrally, with versioning and controlled deployment.
Creating and Managing Prompts
# Create a new prompt in LangFuse
langfuse.createprompt(
name="customer-support-v1",
prompt="You are a customer support assistant for {{companyname}}. "
"Answer the following question helpfully and informatively.\n\n"
"Question: {{question}}\n"
"Context: {{context}}",
config={
"model": "gpt-4o",
"temperature": 0.7,
"maxtokens": 500
},
labels=["production"] # Label for deployment
)
Using Prompts from LangFuse
# Fetch prompt from LangFuse
prompt = langfuse.getprompt("customer-support-v1")
Compile with variables
compiled = prompt.compile(
companyname="PT RUBYTHALIB DATA KONSULTA",
question="How do I reset my password?",
context="Password reset guide: ..."
)
print(compiled)
Output: "You are a customer support assistant for PT RUBYTHALIB DATA KONSULTA..."
Use config from prompt
print(prompt.config)
{'model': 'gpt-4o', 'temperature': 0.7, 'maxtokens': 500}
Chat Prompts (Multi-Message)
# Create a chat prompt
langfuse.createprompt(
name="chat-support-v1",
prompt=[
{"role": "system", "content": "You are an assistant for {{company}}."},
{"role": "user", "content": "{{usermessage}}"}
],
type="chat",
labels=["staging"]
)
Fetch and compile
chatprompt = langfuse.getprompt("chat-support-v1", type="chat")
messages = chatprompt.compile(
company="PT RUBYTHALIB DATA KONSULTA",
usermessage="Hello, I need help"
)
Evaluation and Scoring
LangFuse supports several evaluation methods for measuring LLM output quality.
Manual Scoring via SDK
# Score a trace
trace = langfuse.trace(name="eval-example")
... LLM processing ...
Manual scores
trace.score(
name="helpfulness",
value=0.9,
comment="Very helpful and complete answer"
)
trace.score(
name="accuracy",
value=0.85,
comment="Accurate information but lacking detail at the end"
)
trace.score(
name="user-feedback",
value=1, # thumbs up
datatype="BOOLEAN"
)
Automated Evaluation with LLM-as-Judge
from langfuse.decorators import observe, langfusecontext
import openai
@observe()
def evaluateresponse(question: str, response: str, reference: str) -> dict:
client = openai.OpenAI()
evalprompt = f"""Evaluate the following response based on these criteria:
Relevance (0-1): How relevant is the answer to the question
Accuracy (0-1): How accurate compared to the reference
Completeness (0-1): How complete is the answer
Question: {question}
Response: {response}
Reference: {reference}
Provide scores in JSON format."""
result = client.chat.completions.create(
model="gpt-4o",
messages=[{"role": "user", "content": evalprompt}],
responseformat={"type": "jsonobject"}
)
import json
scores = json.loads(result.choices[0].message.content)
# Record scores in LangFuse
langfusecontext.scorecurrenttrace(
name="relevance", value=scores.get("relevance", 0)
)
langfusecontext.scorecurrenttrace(
name="accuracy", value=scores.get("accuracy", 0)
)
langfusecontext.scorecurrenttrace(
name="completeness", value=scores.get("completeness", 0)
)
return scores
Integration with LangChain
LangFuse integrates seamlessly with LangChain through a callback handler.
pip install langfuse langchain langchain-openai
from langfuse.callback import CallbackHandler
from langchainopenai import ChatOpenAI
from langchain.prompts import ChatPromptTemplate
from langchain.schema.runnable import RunnablePassthrough
Initialize LangFuse handler
langfusehandler = CallbackHandler(
publickey="pk-lf-...",
secretkey="sk-lf-...",
host="https://cloud.langfuse.com"
)
Create LangChain chain
prompt = ChatPromptTemplate.frommessages([
("system", "You are a helpful assistant."),
("human", "{question}")
])
llm = ChatOpenAI(model="gpt-4o", temperature=0.7)
chain = prompt | llm
Run with LangFuse tracing
response = chain.invoke(
{"question": "What is machine learning?"},
config={"callbacks": [langfusehandler]}
)
print(response.content)
RAG Chain with LangFuse Tracing
from langchainopenai import OpenAIEmbeddings
from langchain
community.vectorstores import FAISS
from langchain.textsplitter import RecursiveCharacterTextSplitter
Setup vector store
embeddings = OpenAIEmbeddings()
textsplitter = RecursiveCharacterTextSplitter(
chunksize=1000,
chunkoverlap=200
)
Simulate documents
documents = textsplitter.createdocuments([
"Machine learning is a branch of AI...",
"Deep learning uses neural networks...",
])
vectorstore = FAISS.fromdocuments(documents, embeddings)
retriever = vectorstore.asretriever(searchkwargs={"k": 3})
RAG chain
from langchain.schema.outputparser import StrOutputParser
ragprompt = ChatPromptTemplate.frommessages([
("system", "Answer based on the following context:\n{context}"),
("human", "{question}")
])
ragchain = (
{"context": retriever, "question": RunnablePassthrough()}
| ragprompt
| llm
| StrOutputParser()
)
All steps (retrieval + generation) are automatically traced
result = ragchain.invoke(
"What is the difference between ML and DL?",
config={"callbacks": [langfusehandler]}
)
Integration with LlamaIndex
pip install langfuse llama-index
from langfuse.llamaindex import LlamaIndexCallbackHandler
Setup handler
langfuse
handler = LlamaIndexCallbackHandler(
publickey="pk-lf-...",
secretkey="sk-lf-...",
host="https://cloud.langfuse.com"
)
import llamaindex.core
llamaindex.core.globalhandler = langfusehandler
Now all LlamaIndex operations are automatically traced
from llamaindex.core import VectorStoreIndex, SimpleDirectoryReader
documents = SimpleDirectoryReader("data").loaddata()
index = VectorStoreIndex.fromdocuments(documents)
queryengine = index.asqueryengine()
Query - automatically traced in LangFuse
response = queryengine.query("Explain the concept of RAG")
print(response)
Cost Tracking
LangFuse automatically tracks costs based on model and token usage.
# Automatic cost tracking when using generation
generation = trace.generation(
name="cost-example",
model="gpt-4o",
input=[{"role": "user", "content": "Hello"}],
output="Hi there!",
usage={
"input": 10,
"output": 5,
"unit": "TOKENS"
}
# Cost is automatically calculated based on model pricing
)
Or set cost manually
generationcustom = trace.generation(
name="custom-model",
model="custom-llm-v1",
usage={
"input": 100,
"output": 50,
"unit": "TOKENS",
"totalcost": 0.003 # USD
}
)
Fetching Cost Data via API
# Using the API for analytics
import requests
response = requests.get(
"https://cloud.langfuse.com/api/public/metrics/daily",
headers={
"Authorization": "Basic key:secretkey)>"
},
params={
"traceName": "customer-support-chat",
"fromTimestamp": "2026-04-01T00:00:00Z",
"toTimestamp": "2026-04-30T23:59:59Z"
}
)
metrics = response.json()
for day in metrics["data"]:
print(f"Date: {day['date']}")
print(f" Total traces: {day['countTraces']}")
print(f" Total cost: ${day['totalCost']:.4f}")
print(f" Avg latency: {day['avgLatency']:.2f}s")
Dashboard Analytics
LangFuse provides a comprehensive web dashboard with metrics including:
- Trace Volume: Number of requests over time
- Latency Distribution: P50, P90, P99 latency
- Cost Breakdown: Cost per model, per feature, per user
- Token Usage: Input/output token consumption
- Error Rate: Error percentage and error types
- Score Distribution: Distribution of evaluation scores
- User Analytics: Usage per user/session
You can also create custom dashboards with filters based on tags, users, metadata, and time ranges.
Practical Example: Customer Support Bot with Full Observability
Let's build a complete customer support bot with full observability.
"""
Customer Support Bot with LangFuse Observability
"""
from langfuse.decorators import observe, langfusecontext
import openai
import json
client = openai.OpenAI()
Simple knowledge base
KNOWLEDGEBASE = {
"resetpassword": "To reset your password: 1) Go to the login page, "
"2) Click 'Forgot Password', 3) Enter your email, "
"4) Check your inbox for the reset link.",
"billing": "For billing questions: Contact the finance team at "
"finance@company.com or through the Billing menu in the dashboard.",
"technical": "For technical issues: 1) Check service status at status.company.com, "
"2) Restart the application, 3) Clear browser cache."
}
@observe()
def classifyintent(query: str) -> str:
"""Classify the intent of the user's question."""
response = client.chat.completions.create(
model="gpt-4o-mini",
messages=[
{"role": "system", "content": "Classify intent: resetpassword, billing, technical, other"},
{"role": "user", "content": query}
],
temperature=0
)
langfusecontext.updatecurrentobservation(
model="gpt-4o-mini",
usage={
"input": response.usage.prompttokens,
"output": response.usage.completiontokens
}
)
return response.choices[0].message.content.strip().lower()
@observe()
def retrieveknowledge(intent: str) -> str:
"""Retrieve knowledge base content based on intent."""
context = KNOWLEDGEBASE.get(intent, "No specific information available.")
langfusecontext.updatecurrentobservation(
metadata={"intent": intent, "found": intent in KNOWLEDGEBASE}
)
return context
@observe(astype="generation")
def generateanswer(query: str, context: str, intent: str) -> str:
"""Generate an answer using the LLM."""
response = client.chat.completions.create(
model="gpt-4o",
messages=[
{
"role": "system",
"content": f"You are a customer support assistant. "
f"Use the following context to answer: {context}"
},
{"role": "user", "content": query}
],
temperature=0.7,
maxtokens=500
)
langfusecontext.updatecurrentobservation(
model="gpt-4o",
usage={
"input": response.usage.prompttokens,
"output": response.usage.completiontokens
},
metadata={"intent": intent}
)
return response.choices[0].message.content
@observe()
def autoevaluate(query: str, response: str) -> dict:
"""Automatically evaluate response quality."""
evalresponse = client.chat.completions.create(
model="gpt-4o-mini",
messages=[
{
"role": "user",
"content": f"Rate this support response (0-1) for: "
f"helpfulness, clarity, accuracy.\n"
f"Question: {query}\nResponse: {response}\n"
f"Return JSON only."
}
],
responseformat={"type": "jsonobject"},
temperature=0
)
scores = json.loads(evalresponse.choices[0].message.content)
for metric, value in scores.items():
langfusecontext.scorecurrenttrace(
name=metric,
value=float(value)
)
return scores
@observe()
def handlesupportquery(userid: str, query: str) -> dict:
"""Main customer support pipeline."""
langfusecontext.updatecurrenttrace(
userid=userid,
tags=["customer-support", "production"],
metadata={"source": "web-chat"}
)
# Step 1: Classify intent
intent = classifyintent(query)
# Step 2: Retrieve knowledge
context = retrieveknowledge(intent)
# Step 3: Generate answer
answer = generateanswer(query, context, intent)
# Step 4: Auto-evaluate
scores = autoevaluate(query, answer)
return {
"intent": intent,
"answer": answer,
"scores": scores
}
Run
if name == "main":
result = handlesupportquery(
userid="user-456",
query="I forgot my password, how do I reset it?"
)
print(f"Intent: {result['intent']}")
print(f"Answer: {result['answer']}")
print(f"Scores: {result['scores']}")
Best Practices
langfuse.flush() at the end of requests or use the @observe() decorator which auto-flushes# Sampling example
import random
langfuse = Langfuse()
Only trace 10% of requests in production
if random.random() < 0.1:
trace = langfuse.trace(name="sampled-trace")
Conclusion
LangFuse is an invaluable tool for teams building LLM applications. With comprehensive tracing, monitoring, evaluation, and prompt management features, LangFuse helps you understand and optimize your LLM applications systematically.
Its open-source nature enables self-hosting for full data control, while integrations with LangChain and LlamaIndex make adoption very straightforward. Start with LangFuse Cloud for experimentation, then consider self-hosting when you're ready for production.
Happy building and we hope this tutorial is helpful!