LangFuse: Open-Source Platform for LLM Application Observability

# LangFuse: Platform Open-Source untuk Observability Aplikasi LLM Seiring semakin banyaknya perusahaan yang mengadopsi Large Language Model (LLM) dalam aplikasi produksi, kebutuhan akan **observabili...

By Ruby Abdullah · · tutorial
LangFuseLLMObservabilityMonitoringPython

LangFuse: Open-Source Platform for LLM Application Observability

As more companies adopt Large Language Models (LLMs) in production applications, the need for observability becomes critical. How do you know the cost of each request? How do you track the performance of different prompts? How do you systematically evaluate the quality of LLM outputs?

LangFuse is an open-source solution that answers all these questions. It provides a comprehensive observability platform for tracing, monitoring, evaluating, and managing prompts in your LLM applications.

What Is LangFuse?

LangFuse is an open-source LLM engineering platform that helps development teams with:

  • Tracing: Track every step of LLM application execution, from input to output
  • Monitoring: Monitor latency, cost, and error rates in real-time
  • Evaluation: Score LLM outputs both manually and automatically
  • Prompt Management: Centralized prompt versioning and deployment
  • Analytics: Dashboards for analyzing usage and performance

LangFuse supports various LLM providers (OpenAI, Anthropic, Azure, etc.) and integrates with popular frameworks like LangChain, LlamaIndex, and Haystack.

Deployment Options

1. LangFuse Cloud (Managed)

The easiest way to get started is using LangFuse Cloud at cloud.langfuse.com.

# No installation needed - just sign up and get API keys

LANGFUSEPUBLICKEY=pk-lf-...

LANGFUSESECRETKEY=sk-lf-...

LANGFUSEHOST=https://cloud.langfuse.com

2. Self-Hosted with Docker

For full control over your data, you can deploy LangFuse yourself:

# docker-compose.yml

version: '3.8'

services:

langfuse:

image: langfuse/langfuse:2

ports:

  • "3000:3000"
environment:

  • DATABASEURL=postgresql://postgres:postgres@db:5432/langfuse
  • NEXTAUTHSECRET=mysecret
  • SALT=mysalt
  • NEXTAUTHURL=http://localhost:3000
dependson:

  • db

db:

image: postgres:15

environment:

  • POSTGRESUSER=postgres
  • POSTGRESPASSWORD=postgres
  • POSTGRESDB=langfuse
volumes:

  • pgdata:/var/lib/postgresql/data

volumes:

pgdata:

# Start LangFuse

docker-compose up -d

Access at http://localhost:3000

Create the first admin account through the UI

Python SDK Installation

pip install langfuse

Environment Configuration

import os

os.environ["LANGFUSEPUBLICKEY"] = "pk-lf-..."

os.environ["LANGFUSESECRETKEY"] = "sk-lf-..."

os.environ["LANGFUSEHOST"] = "https://cloud.langfuse.com" # or self-hosted URL

Client Initialization

from langfuse import Langfuse

langfuse = Langfuse(

publickey="pk-lf-...",

secretkey="sk-lf-...",

host="https://cloud.langfuse.com"

)

Verify connection

langfuse.authcheck()

print("LangFuse connected!")

Tracing: Tracking LLM Application Execution

Tracing is LangFuse's core feature. Every interaction with your LLM application is recorded as a trace consisting of several components:

Core Tracing Concepts

  • Trace: The top-level unit representing one end-to-end execution
  • Span: A sub-unit representing a single step in the trace (e.g., retrieval, preprocessing)
  • Generation: A special span for LLM calls, automatically capturing token usage and cost

Creating a Simple Trace

from langfuse import Langfuse

langfuse = Langfuse()

Create a new trace

trace = langfuse.trace(

name="customer-support-chat",

userid="user-123",

metadata={"session": "abc-456"},

tags=["production", "customer-support"]

)

Add a span for preprocessing step

span = trace.span(

name="preprocess-input",

input={"rawmessage": "How do I reset my password?"}

)

Simulate preprocessing

processedinput = "How do I reset my password?"

span.end(output={"processed": processedinput})

Recording LLM Generations

# Record LLM call as a generation

generation = trace.generation(

name="llm-response",

model="gpt-4o",

modelparameters={

"temperature": 0.7,

"maxtokens": 500

},

input=[

{"role": "system", "content": "You are a customer support assistant."},

{"role": "user", "content": processedinput}

]

)

Simulate LLM response

responsetext = "To reset your password, please follow these steps..."

Update generation with output and usage

generation.end(

output=responsetext,

usage={

"input": 45,

"output": 120,

"unit": "TOKENS"

}

)

IMPORTANT: flush data to server

langfuse.flush()

Tracing with Decorators

LangFuse provides the @observe() decorator for cleaner automatic tracing:

from langfuse.decorators import observe, langfusecontext

@observe()

def processcustomerquery(query: str):

# Step 1: Preprocessing

cleaned = preprocess(query)

# Step 2: Retrieve context

context = retrievedocuments(cleaned)

# Step 3: Generate response

response = generateresponse(cleaned, context)

return response

@observe()

def preprocess(text: str) -> str:

return text.strip().lower()

@observe()

def retrievedocuments(query: str) -> list:

# Simulate vector search

return [

{"title": "Reset Password Guide", "content": "..."},

{"title": "Account Recovery", "content": "..."}

]

@observe(astype="generation")

def generateresponse(query: str, context: list) -> str:

import openai

client = openai.OpenAI()

messages = [

{"role": "system", "content": f"Context: {context}"},

{"role": "user", "content": query}

]

response = client.chat.completions.create(

model="gpt-4o",

messages=messages,

temperature=0.7

)

# Update observation with generation details

langfusecontext.updatecurrentobservation(

model="gpt-4o",

usage={

"input": response.usage.prompttokens,

"output": response.usage.completiontokens

}

)

return response.choices[0].message.content

Run - all steps are automatically traced

result = processcustomerquery("How do I reset my password?")

Prompt Management

LangFuse allows you to manage prompts centrally, with versioning and controlled deployment.

Creating and Managing Prompts

# Create a new prompt in LangFuse

langfuse.createprompt(

name="customer-support-v1",

prompt="You are a customer support assistant for {{companyname}}. "

"Answer the following question helpfully and informatively.\n\n"

"Question: {{question}}\n"

"Context: {{context}}",

config={

"model": "gpt-4o",

"temperature": 0.7,

"maxtokens": 500

},

labels=["production"] # Label for deployment

)

Using Prompts from LangFuse

# Fetch prompt from LangFuse

prompt = langfuse.getprompt("customer-support-v1")

Compile with variables

compiled = prompt.compile(

companyname="PT RUBYTHALIB DATA KONSULTA",

question="How do I reset my password?",

context="Password reset guide: ..."

)

print(compiled)

Output: "You are a customer support assistant for PT RUBYTHALIB DATA KONSULTA..."

Use config from prompt

print(prompt.config)

{'model': 'gpt-4o', 'temperature': 0.7, 'maxtokens': 500}

Chat Prompts (Multi-Message)

# Create a chat prompt

langfuse.createprompt(

name="chat-support-v1",

prompt=[

{"role": "system", "content": "You are an assistant for {{company}}."},

{"role": "user", "content": "{{usermessage}}"}

],

type="chat",

labels=["staging"]

)

Fetch and compile

chatprompt = langfuse.getprompt("chat-support-v1", type="chat")

messages = chatprompt.compile(

company="PT RUBYTHALIB DATA KONSULTA",

usermessage="Hello, I need help"

)

Evaluation and Scoring

LangFuse supports several evaluation methods for measuring LLM output quality.

Manual Scoring via SDK

# Score a trace

trace = langfuse.trace(name="eval-example")

... LLM processing ...

Manual scores

trace.score(

name="helpfulness",

value=0.9,

comment="Very helpful and complete answer"

)

trace.score(

name="accuracy",

value=0.85,

comment="Accurate information but lacking detail at the end"

)

trace.score(

name="user-feedback",

value=1, # thumbs up

datatype="BOOLEAN"

)

Automated Evaluation with LLM-as-Judge

from langfuse.decorators import observe, langfusecontext

import openai

@observe()

def evaluateresponse(question: str, response: str, reference: str) -> dict:

client = openai.OpenAI()

evalprompt = f"""Evaluate the following response based on these criteria:

  • Relevance (0-1): How relevant is the answer to the question
  • Accuracy (0-1): How accurate compared to the reference
  • Completeness (0-1): How complete is the answer
  • Question: {question}

    Response: {response}

    Reference: {reference}

    Provide scores in JSON format."""

    result = client.chat.completions.create(

    model="gpt-4o",

    messages=[{"role": "user", "content": evalprompt}],

    responseformat={"type": "jsonobject"}

    )

    import json

    scores = json.loads(result.choices[0].message.content)

    # Record scores in LangFuse

    langfusecontext.scorecurrenttrace(

    name="relevance", value=scores.get("relevance", 0)

    )

    langfusecontext.scorecurrenttrace(

    name="accuracy", value=scores.get("accuracy", 0)

    )

    langfusecontext.scorecurrenttrace(

    name="completeness", value=scores.get("completeness", 0)

    )

    return scores

    Integration with LangChain

    LangFuse integrates seamlessly with LangChain through a callback handler.

    pip install langfuse langchain langchain-openai
    

    from langfuse.callback import CallbackHandler
    

    from langchainopenai import ChatOpenAI

    from langchain.prompts import ChatPromptTemplate

    from langchain.schema.runnable import RunnablePassthrough

    Initialize LangFuse handler

    langfusehandler = CallbackHandler(

    publickey="pk-lf-...",

    secretkey="sk-lf-...",

    host="https://cloud.langfuse.com"

    )

    Create LangChain chain

    prompt = ChatPromptTemplate.frommessages([

    ("system", "You are a helpful assistant."),

    ("human", "{question}")

    ])

    llm = ChatOpenAI(model="gpt-4o", temperature=0.7)

    chain = prompt | llm

    Run with LangFuse tracing

    response = chain.invoke(

    {"question": "What is machine learning?"},

    config={"callbacks": [langfusehandler]}

    )

    print(response.content)

    RAG Chain with LangFuse Tracing

    from langchainopenai import OpenAIEmbeddings
    

    from langchaincommunity.vectorstores import FAISS

    from langchain.textsplitter import RecursiveCharacterTextSplitter

    Setup vector store

    embeddings = OpenAIEmbeddings()

    textsplitter = RecursiveCharacterTextSplitter(

    chunksize=1000,

    chunkoverlap=200

    )

    Simulate documents

    documents = textsplitter.createdocuments([

    "Machine learning is a branch of AI...",

    "Deep learning uses neural networks...",

    ])

    vectorstore = FAISS.fromdocuments(documents, embeddings)

    retriever = vectorstore.asretriever(searchkwargs={"k": 3})

    RAG chain

    from langchain.schema.outputparser import StrOutputParser

    ragprompt = ChatPromptTemplate.frommessages([

    ("system", "Answer based on the following context:\n{context}"),

    ("human", "{question}")

    ])

    ragchain = (

    {"context": retriever, "question": RunnablePassthrough()}

    | ragprompt

    | llm

    | StrOutputParser()

    )

    All steps (retrieval + generation) are automatically traced

    result = ragchain.invoke(

    "What is the difference between ML and DL?",

    config={"callbacks": [langfusehandler]}

    )

    Integration with LlamaIndex

    pip install langfuse llama-index
    

    from langfuse.llamaindex import LlamaIndexCallbackHandler
    
    

    Setup handler

    langfusehandler = LlamaIndexCallbackHandler(

    publickey="pk-lf-...",

    secretkey="sk-lf-...",

    host="https://cloud.langfuse.com"

    )

    import llamaindex.core

    llamaindex.core.globalhandler = langfusehandler

    Now all LlamaIndex operations are automatically traced

    from llamaindex.core import VectorStoreIndex, SimpleDirectoryReader

    documents = SimpleDirectoryReader("data").loaddata()

    index = VectorStoreIndex.fromdocuments(documents)

    queryengine = index.asqueryengine()

    Query - automatically traced in LangFuse

    response = queryengine.query("Explain the concept of RAG")

    print(response)

    Cost Tracking

    LangFuse automatically tracks costs based on model and token usage.

    # Automatic cost tracking when using generation
    

    generation = trace.generation(

    name="cost-example",

    model="gpt-4o",

    input=[{"role": "user", "content": "Hello"}],

    output="Hi there!",

    usage={

    "input": 10,

    "output": 5,

    "unit": "TOKENS"

    }

    # Cost is automatically calculated based on model pricing

    )

    Or set cost manually

    generationcustom = trace.generation(

    name="custom-model",

    model="custom-llm-v1",

    usage={

    "input": 100,

    "output": 50,

    "unit": "TOKENS",

    "totalcost": 0.003 # USD

    }

    )

    Fetching Cost Data via API

    # Using the API for analytics
    

    import requests

    response = requests.get(

    "https://cloud.langfuse.com/api/public/metrics/daily",

    headers={

    "Authorization": "Basic key:secretkey)>"

    },

    params={

    "traceName": "customer-support-chat",

    "fromTimestamp": "2026-04-01T00:00:00Z",

    "toTimestamp": "2026-04-30T23:59:59Z"

    }

    )

    metrics = response.json()

    for day in metrics["data"]:

    print(f"Date: {day['date']}")

    print(f" Total traces: {day['countTraces']}")

    print(f" Total cost: ${day['totalCost']:.4f}")

    print(f" Avg latency: {day['avgLatency']:.2f}s")

    Dashboard Analytics

    LangFuse provides a comprehensive web dashboard with metrics including:

    • Trace Volume: Number of requests over time
    • Latency Distribution: P50, P90, P99 latency
    • Cost Breakdown: Cost per model, per feature, per user
    • Token Usage: Input/output token consumption
    • Error Rate: Error percentage and error types
    • Score Distribution: Distribution of evaluation scores
    • User Analytics: Usage per user/session

    You can also create custom dashboards with filters based on tags, users, metadata, and time ranges.

    Practical Example: Customer Support Bot with Full Observability

    Let's build a complete customer support bot with full observability.

    """
    

    Customer Support Bot with LangFuse Observability

    """

    from langfuse.decorators import observe, langfusecontext

    import openai

    import json

    client = openai.OpenAI()

    Simple knowledge base

    KNOWLEDGEBASE = {

    "resetpassword": "To reset your password: 1) Go to the login page, "

    "2) Click 'Forgot Password', 3) Enter your email, "

    "4) Check your inbox for the reset link.",

    "billing": "For billing questions: Contact the finance team at "

    "finance@company.com or through the Billing menu in the dashboard.",

    "technical": "For technical issues: 1) Check service status at status.company.com, "

    "2) Restart the application, 3) Clear browser cache."

    }

    @observe()

    def classifyintent(query: str) -> str:

    """Classify the intent of the user's question."""

    response = client.chat.completions.create(

    model="gpt-4o-mini",

    messages=[

    {"role": "system", "content": "Classify intent: resetpassword, billing, technical, other"},

    {"role": "user", "content": query}

    ],

    temperature=0

    )

    langfusecontext.updatecurrentobservation(

    model="gpt-4o-mini",

    usage={

    "input": response.usage.prompttokens,

    "output": response.usage.completiontokens

    }

    )

    return response.choices[0].message.content.strip().lower()

    @observe()

    def retrieveknowledge(intent: str) -> str:

    """Retrieve knowledge base content based on intent."""

    context = KNOWLEDGEBASE.get(intent, "No specific information available.")

    langfusecontext.updatecurrentobservation(

    metadata={"intent": intent, "found": intent in KNOWLEDGEBASE}

    )

    return context

    @observe(astype="generation")

    def generateanswer(query: str, context: str, intent: str) -> str:

    """Generate an answer using the LLM."""

    response = client.chat.completions.create(

    model="gpt-4o",

    messages=[

    {

    "role": "system",

    "content": f"You are a customer support assistant. "

    f"Use the following context to answer: {context}"

    },

    {"role": "user", "content": query}

    ],

    temperature=0.7,

    maxtokens=500

    )

    langfusecontext.updatecurrentobservation(

    model="gpt-4o",

    usage={

    "input": response.usage.prompttokens,

    "output": response.usage.completiontokens

    },

    metadata={"intent": intent}

    )

    return response.choices[0].message.content

    @observe()

    def autoevaluate(query: str, response: str) -> dict:

    """Automatically evaluate response quality."""

    evalresponse = client.chat.completions.create(

    model="gpt-4o-mini",

    messages=[

    {

    "role": "user",

    "content": f"Rate this support response (0-1) for: "

    f"helpfulness, clarity, accuracy.\n"

    f"Question: {query}\nResponse: {response}\n"

    f"Return JSON only."

    }

    ],

    responseformat={"type": "jsonobject"},

    temperature=0

    )

    scores = json.loads(evalresponse.choices[0].message.content)

    for metric, value in scores.items():

    langfusecontext.scorecurrenttrace(

    name=metric,

    value=float(value)

    )

    return scores

    @observe()

    def handlesupportquery(userid: str, query: str) -> dict:

    """Main customer support pipeline."""

    langfusecontext.updatecurrenttrace(

    userid=userid,

    tags=["customer-support", "production"],

    metadata={"source": "web-chat"}

    )

    # Step 1: Classify intent

    intent = classifyintent(query)

    # Step 2: Retrieve knowledge

    context = retrieveknowledge(intent)

    # Step 3: Generate answer

    answer = generateanswer(query, context, intent)

    # Step 4: Auto-evaluate

    scores = autoevaluate(query, answer)

    return {

    "intent": intent,

    "answer": answer,

    "scores": scores

    }

    Run

    if name == "main":

    result = handlesupportquery(

    userid="user-456",

    query="I forgot my password, how do I reset it?"

    )

    print(f"Intent: {result['intent']}")

    print(f"Answer: {result['answer']}")

    print(f"Scores: {result['scores']}")

    Best Practices

  • Always flush: Call langfuse.flush() at the end of requests or use the @observe() decorator which auto-flushes
  • Use tags: Tag traces with environment (production, staging) and features for easy filtering
  • Track userid: Always include userid for per-user analytics
  • Prompt versioning: Use LangFuse prompt management for A/B testing prompts
  • Continuous evaluation: Set up automated evaluation for ongoing quality monitoring
  • Set budget alerts: Monitor costs and set alerts when thresholds are exceeded
  • Sampling in production: For high-traffic applications, use sampling rates to reduce overhead
  • # Sampling example
    

    import random

    langfuse = Langfuse()

    Only trace 10% of requests in production

    if random.random() < 0.1:

    trace = langfuse.trace(name="sampled-trace")

    Conclusion

    LangFuse is an invaluable tool for teams building LLM applications. With comprehensive tracing, monitoring, evaluation, and prompt management features, LangFuse helps you understand and optimize your LLM applications systematically.

    Its open-source nature enables self-hosting for full data control, while integrations with LangChain and LlamaIndex make adoption very straightforward. Start with LangFuse Cloud for experimentation, then consider self-hosting when you're ready for production.

    Happy building and we hope this tutorial is helpful!

    Related Articles

    LangSmith Tutorial: LLM Observability and Debugging

    Tutorial 8: LangSmith - Observabilitas dan Debugging LLM Daftar Isi Pendahuluan Prasyarat Menyiapkan LangSmith [Melacak ...

    DSPy: Stop Hand-Tuning Prompts, Let the Compiler Optimize Them

    DSPy: Berhenti Ngoprek Prompt Manual, Biarkan Compiler yang Optimasi Halo temen-temen, kali ini aku mau ngenalin satu li...

    Inspect AI: The LLM Evaluation Framework from the UK AI Safety Institute

    Inspect AI: Framework Evaluasi LLM dari UK AI Safety Institute yang Wajib Kamu Coba Temen-temen, kalau kamu udah mulai s...

    Complete Braintrust Tutorial: Evaluate, Test, and Improve Your LLM Applications

    Tutorial Lengkap Braintrust: Evaluasi, Testing, dan Improve Aplikasi LLM Halo temen-temen, di tutorial kali ini aku mau ...