LangSmith Tutorial: LLM Observability and Debugging

# Tutorial 8: LangSmith - Observabilitas dan Debugging LLM ## Daftar Isi 1. [Pendahuluan](#pendahuluan) 2. [Prasyarat](#prasyarat) 3. [Menyiapkan LangSmith](#menyiapkan-langsmith) 4. [Melacak Panggi...

By Ruby Abdullah · · tutorial
LangSmithLLMObservabilityDebuggingLangChainMonitoring

Tutorial 8: LangSmith - LLM Observability and Debugging

Table of Contents

  • Introduction
  • Prerequisites
  • Setting Up LangSmith
  • Tracing LLM Calls
  • Evaluation Datasets
  • Custom Evaluators
  • Prompt Versioning
  • Production Monitoring
  • Debugging Chains and Agents
  • Cost Tracking
  • Integration with LangChain
  • A/B Testing Prompts
  • Best Practices
  • Conclusion

  • Introduction

    Building LLM-powered applications is only half the battle. Understanding how your models behave in production, why they fail, how much they cost, and whether your prompts are actually improving over time is the other critical half. LangSmith is the observability and evaluation platform built by LangChain specifically to address these challenges.

    LangSmith provides end-to-end tracing of every LLM call, chain execution, and agent step. It enables you to build evaluation datasets, run automated evaluations, version your prompts, monitor production performance, and debug complex multi-step workflows. Whether you are prototyping a simple chatbot or running a production RAG pipeline, LangSmith gives you the visibility you need to ship reliable AI applications.

    This tutorial walks you through every major feature of LangSmith with practical, production-ready code examples.


    Prerequisites

    • Python 3.9 or higher
    • An active LangSmith account (sign up at smith.langchain.com)
    • An OpenAI API key (or any supported LLM provider)
    • Basic familiarity with LangChain concepts

    Install the required packages:

    pip install langsmith langchain langchain-openai openai
    


    Setting Up LangSmith

    Step 1: Configure Environment Variables

    LangSmith requires a few environment variables to connect your application to the platform.

    import os
    
    

    os.environ["LANGCHAINTRACINGV2"] = "true"

    os.environ["LANGCHAINAPIKEY"] = "ls_yourapikeyhere"

    os.environ["LANGCHAINPROJECT"] = "my-first-project"

    os.environ["LANGCHAINENDPOINT"] = "https://api.smith.langchain.com"

    Step 2: Verify the Connection

    from langsmith import Client
    
    

    client = Client()

    List existing projects to verify connectivity

    projects = list(client.listprojects())

    print(f"Connected to LangSmith. Found {len(projects)} project(s).")

    for project in projects:

    print(f" - {project.name} (created: {project.createdat})")

    Step 3: Create a Dedicated Project

    # Create a new project for organized tracking
    

    projectname = "tutorial-langsmith-demo"

    try:

    project = client.createproject(projectname)

    print(f"Project '{project.name}' created successfully.")

    except Exception as e:

    print(f"Project may already exist: {e}")

    Set the active project

    os.environ["LANGCHAINPROJECT"] = projectname


    Tracing LLM Calls

    Basic Tracing with LangChain

    When tracing is enabled, every LangChain call is automatically captured.

    from langchainopenai import ChatOpenAI
    

    from langchaincore.messages import HumanMessage, SystemMessage

    llm = ChatOpenAI(model="gpt-4o", temperature=0.7)

    messages = [

    SystemMessage(content="You are a helpful coding assistant."),

    HumanMessage(content="Explain Python decorators in 3 sentences.")

    ]

    response = llm.invoke(messages)

    print(response.content)

    The trace is automatically sent to LangSmith

    Manual Tracing with the @traceable Decorator

    For custom functions that are not LangChain components, use the @traceable decorator.

    from langsmith import traceable
    
    

    @traceable(name="processuserquery")

    def processuserquery(query: str, context: str = "") -> dict:

    """Process a user query with optional context."""

    llm = ChatOpenAI(model="gpt-4o", temperature=0)

    prompt = f"Context: {context}\n\nQuestion: {query}" if context else query

    response = llm.invoke([HumanMessage(content=prompt)])

    return {

    "query": query,

    "response": response.content,

    "model": "gpt-4o",

    "hascontext": bool(context)

    }

    Every call is traced with inputs, outputs, and metadata

    result = processuserquery(

    query="What is dependency injection?",

    context="We are discussing software design patterns."

    )

    print(result["response"])

    Nested Tracing

    Nested function calls create a trace tree, making it easy to see the full execution flow.

    @traceable(name="retrievedocuments")
    

    def retrievedocuments(query: str) -> list:

    """Simulate document retrieval."""

    return [

    {"id": 1, "text": "Python is a high-level programming language."},

    {"id": 2, "text": "Python supports multiple programming paradigms."},

    ]

    @traceable(name="generateanswer")

    def generateanswer(query: str, documents: list) -> str:

    """Generate answer from retrieved documents."""

    llm = ChatOpenAI(model="gpt-4o", temperature=0)

    context = "\n".join([doc["text"] for doc in documents])

    prompt = f"Based on the following context:\n{context}\n\nAnswer: {query}"

    response = llm.invoke([HumanMessage(content=prompt)])

    return response.content

    @traceable(name="ragpipeline")

    def ragpipeline(query: str) -> str:

    """Full RAG pipeline with nested tracing."""

    docs = retrievedocuments(query)

    answer = generateanswer(query, docs)

    return answer

    answer = ragpipeline("What programming paradigms does Python support?")

    print(answer)


    Evaluation Datasets

    Creating a Dataset

    from langsmith import Client
    
    

    client = Client()

    Create a dataset for evaluation

    datasetname = "qa-evaluation-set"

    dataset = client.createdataset(

    datasetname,

    description="Question-answer pairs for evaluating our RAG pipeline."

    )

    Add examples to the dataset

    examples = [

    {

    "inputs": {"query": "What is Python?"},

    "outputs": {"answer": "Python is a high-level, interpreted programming language."}

    },

    {

    "inputs": {"query": "What is machine learning?"},

    "outputs": {"answer": "Machine learning is a subset of AI that enables systems to learn from data."}

    },

    {

    "inputs": {"query": "What is a neural network?"},

    "outputs": {"answer": "A neural network is a computational model inspired by the human brain."}

    },

    {

    "inputs": {"query": "What is natural language processing?"},

    "outputs": {"answer": "NLP is a field of AI focused on interaction between computers and human language."}

    },

    ]

    for example in examples:

    client.createexample(

    inputs=example["inputs"],

    outputs=example["outputs"],

    datasetid=dataset.id

    )

    print(f"Dataset '{datasetname}' created with {len(examples)} examples.")

    Running Evaluations Against a Dataset

    from langsmith.evaluation import evaluate
    
    

    def predictanswer(inputs: dict) -> dict:

    """The function to evaluate."""

    llm = ChatOpenAI(model="gpt-4o", temperature=0)

    response = llm.invoke([HumanMessage(content=inputs["query"])])

    return {"answer": response.content}

    results = evaluate(

    predictanswer,

    data=datasetname,

    experimentprefix="gpt4o-baseline",

    maxconcurrency=2,

    )

    print(f"Evaluation complete. View results at: {results.experimenturl}")


    Custom Evaluators

    Building a Semantic Similarity Evaluator

    from langsmith.schemas import Example, Run
    
    

    def semanticsimilarityevaluator(run: Run, example: Example) -> dict:

    """Evaluate semantic similarity between predicted and expected answers."""

    from langchainopenai import OpenAIEmbeddings

    import numpy as np

    embeddings = OpenAIEmbeddings(model="text-embedding-3-small")

    predicted = run.outputs.get("answer", "")

    expected = example.outputs.get("answer", "")

    predembedding = embeddings.embedquery(predicted)

    expembedding = embeddings.embedquery(expected)

    similarity = np.dot(predembedding, expembedding) / (

    np.linalg.norm(predembedding) np.linalg.norm(expembedding)

    )

    return {

    "key": "semanticsimilarity",

    "score": float(similarity),

    "comment": f"Cosine similarity: {similarity:.4f}"

    }

    def answerrelevanceevaluator(run: Run, example: Example) -> dict:

    """Check if the answer is relevant to the question."""

    llm = ChatOpenAI(model="gpt-4o", temperature=0)

    query = example.inputs.get("query", "")

    predicted = run.outputs.get("answer", "")

    evalprompt = f"""Rate how relevant the following answer is to the question.

    Question: {query}

    Answer: {predicted}

    Return only a number between 0 and 1, where 1 means perfectly relevant."""

    response = llm.invoke([HumanMessage(content=evalprompt)])

    try:

    score = float(response.content.strip())

    except ValueError:

    score = 0.0

    return {

    "key": "answerrelevance",

    "score": score,

    }

    Run evaluation with custom evaluators

    results = evaluate(

    predictanswer,

    data=datasetname,

    evaluators=[semanticsimilarityevaluator, answerrelevanceevaluator],

    experimentprefix="gpt4o-custom-eval",

    maxconcurrency=2,

    )


    Prompt Versioning

    Managing Prompts in LangSmith Hub

    from langchain import hub
    
    

    Push a prompt to LangSmith Hub

    from langchaincore.prompts import ChatPromptTemplate

    qaprompt = ChatPromptTemplate.frommessages([

    ("system", "You are an expert assistant. Answer questions accurately and concisely. "

    "If you don't know the answer, say so."),

    ("human", "Context: {context}\n\nQuestion: {question}")

    ])

    Push the prompt with a commit message

    hub.push("my-org/qa-prompt", qaprompt, newrepoispublic=False)

    print("Prompt pushed to LangSmith Hub.")

    Pull a specific version

    promptv1 = hub.pull("my-org/qa-prompt")

    print(f"Pulled prompt with {len(promptv1.messages)} messages.")

    Update the prompt (creates a new version automatically)

    qapromptv2 = ChatPromptTemplate.frommessages([

    ("system", "You are an expert assistant specializing in technology topics. "

    "Provide detailed, well-structured answers with examples when possible. "

    "If uncertain, express your confidence level."),

    ("human", "Context: {context}\n\nQuestion: {question}")

    ])

    hub.push("my-org/qa-prompt", qapromptv2)

    print("Prompt v2 pushed. Previous version is preserved in history.")


    Production Monitoring

    Setting Up Monitoring Rules

    from langsmith import Client
    
    

    client = Client()

    Query recent runs for monitoring

    runs = client.listruns(

    projectname="tutorial-langsmith-demo",

    executionorder=1, # Top-level runs only

    error=False,

    limit=100,

    )

    Analyze latency and token usage

    latencies = []

    totaltokens = []

    for run in runs:

    if run.endtime and run.starttime:

    latency = (run.endtime - run.starttime).totalseconds()

    latencies.append(latency)

    if run.totaltokens:

    totaltokens.append(run.totaltokens)

    if latencies:

    avglatency = sum(latencies) / len(latencies)

    maxlatency = max(latencies)

    p95latency = sorted(latencies)[int(len(latencies) 0.95)]

    print(f"Average latency: {avglatency:.2f}s")

    print(f"P95 latency: {p95latency:.2f}s")

    print(f"Max latency: {maxlatency:.2f}s")

    if totaltokens:

    avgtokens = sum(totaltokens) / len(totaltokens)

    print(f"Average tokens per call: {avgtokens:.0f}")

    Automated Alerting

    import smtplib
    

    from email.mime.text import MIMEText

    @traceable(name="monitoredpipeline")

    def monitoredpipeline(query: str) -> dict:

    """Pipeline with built-in monitoring."""

    import time

    start = time.time()

    llm = ChatOpenAI(model="gpt-4o", temperature=0)

    response = llm.invoke([HumanMessage(content=query)])

    elapsed = time.time() - start

    result = {

    "answer": response.content,

    "latencyseconds": elapsed,

    "tokencount": response.usagemetadata.get("totaltokens", 0)

    if hasattr(response, "usagemetadata") and response.usagemetadata

    else 0,

    }

    # Alert if latency exceeds threshold

    if elapsed > 10.0:

    print(f"[ALERT] High latency detected: {elapsed:.2f}s for query: {query[:50]}...")

    return result


    Debugging Chains and Agents

    Inspecting Failed Runs

    # Fetch runs that resulted in errors
    

    errorruns = client.listruns(

    projectname="tutorial-langsmith-demo",

    error=True,

    limit=20,

    )

    for run in errorruns:

    print(f"Run ID: {run.id}")

    print(f" Name: {run.name}")

    print(f" Error: {run.error}")

    print(f" Inputs: {run.inputs}")

    print(f" Start: {run.starttime}")

    print("---")

    Debugging Agent Execution Step-by-Step

    from langchain.agents import AgentExecutor, createopenaitoolsagent
    

    from langchaincore.tools import tool

    from langchaincore.prompts import ChatPromptTemplate, MessagesPlaceholder

    @tool

    def calculate(expression: str) -> str:

    """Evaluate a mathematical expression."""

    try:

    result = eval(expression, {"builtins": {}}, {})

    return str(result)

    except Exception as e:

    return f"Error: {e}"

    @tool

    def searchknowledge(query: str) -> str:

    """Search the knowledge base for information."""

    knowledge = {

    "python": "Python is a versatile programming language created by Guido van Rossum.",

    "javascript": "JavaScript is the language of the web, running in browsers and Node.js.",

    }

    for key, value in knowledge.items():

    if key in query.lower():

    return value

    return "No relevant information found."

    prompt = ChatPromptTemplate.frommessages([

    ("system", "You are a helpful assistant with access to tools."),

    MessagesPlaceholder(variablename="chathistory", optional=True),

    ("human", "{input}"),

    MessagesPlaceholder(variablename="agentscratchpad"),

    ])

    llm = ChatOpenAI(model="gpt-4o", temperature=0)

    agent = createopenaitoolsagent(llm, [calculate, searchknowledge], prompt)

    agentexecutor = AgentExecutor(agent=agent, tools=[calculate, searchknowledge], verbose=True)

    This entire execution is traced - every tool call, LLM call, and decision

    result = agentexecutor.invoke({"input": "What is 15 23 + 7? Also, tell me about Python."})

    print(result["output"])

    After execution, inspect the trace in LangSmith UI to see:

    1. The agent's reasoning at each step

    2. Which tools were called and their inputs/outputs

    3. Token usage per step

    4. Total latency breakdown


    Cost Tracking

    Calculating Costs from Traced Runs

    # Token pricing (example rates, adjust to current pricing)
    

    PRICING = {

    "gpt-4o": {"input": 2.50 / 1000000, "output": 10.00 / 1000000},

    "gpt-4o-mini": {"input": 0.15 / 1000000, "output": 0.60 / 1000000},

    "gpt-3.5-turbo": {"input": 0.50 / 1000000, "output": 1.50 / 1000000},

    }

    def calculateprojectcosts(projectname: str, days: int = 7) -> dict:

    """Calculate costs for a project over the specified number of days."""

    from datetime import datetime, timedelta

    client = Client()

    startdate = datetime.now() - timedelta(days=days)

    runs = client.listruns(

    projectname=projectname,

    runtype="llm",

    starttime=startdate,

    )

    totalcost = 0.0

    modelcosts = {}

    for run in runs:

    model = run.extra.get("metadata", {}).get("lsmodelname", "unknown") if run.extra else "unknown"

    inputtokens = run.prompttokens or 0

    outputtokens = run.completiontokens or 0

    if model in PRICING:

    cost = (inputtokens PRICING[model]["input"] +

    outputtokens * PRICING[model]["output"])

    else:

    cost = 0.0

    totalcost += cost

    if model not in modelcosts:

    modelcosts[model] = {"cost": 0.0, "calls": 0, "tokens": 0}

    modelcosts[model]["cost"] += cost

    modelcosts[model]["calls"] += 1

    modelcosts[model]["tokens"] += inputtokens + outputtokens

    print(f"Cost report for '{projectname}' (last {days} days):")

    print(f" Total cost: ${totalcost:.4f}")

    for model, data in modelcosts.items():

    print(f" {model}: ${data['cost']:.4f} ({data['calls']} calls, {data['tokens']} tokens)")

    return {"totalcost": totalcost, "modelcosts": modelcosts}

    costs = calculateprojectcosts("tutorial-langsmith-demo", days=30)


    Integration with LangChain

    Full RAG Pipeline with LangSmith Tracing

    from langchainopenai import ChatOpenAI, OpenAIEmbeddings
    

    from langchaincore.prompts import ChatPromptTemplate

    from langchaincore.outputparsers import StrOutputParser

    from langchaincore.runnables import RunnablePassthrough

    All LangChain components are automatically traced

    embeddings = OpenAIEmbeddings(model="text-embedding-3-small")

    llm = ChatOpenAI(model="gpt-4o", temperature=0)

    Simulated vector store retriever

    @traceable(name="vectorretriever")

    def retrieve(query: str) -> str:

    """Simulate vector store retrieval."""

    docs = {

    "ai": "Artificial intelligence encompasses machine learning, deep learning, and NLP.",

    "ml": "Machine learning algorithms learn patterns from data without explicit programming.",

    }

    results = []

    for key, value in docs.items():

    if key in query.lower():

    results.append(value)

    return "\n".join(results) if results else "No relevant documents found."

    prompt = ChatPromptTemplate.fromtemplate(

    "Answer the question based on the context.\n\n"

    "Context: {context}\n\n"

    "Question: {question}\n\n"

    "Answer:"

    )

    Build the chain - every step is traced

    chain = (

    {"context": lambda x: retrieve(x["question"]), "question": lambda x: x["question"]}

    | prompt

    | llm

    | StrOutputParser()

    )

    Execute and trace

    answer = chain.invoke({"question": "What is machine learning?"})

    print(answer)


    A/B Testing Prompts

    Comparing Prompt Variants

    from langsmith.evaluation import evaluate
    
    

    Define prompt variants

    PROMPTA = "Answer the following question concisely: {query}"

    PROMPTB = """You are an expert assistant. Provide a clear, well-structured answer

    to the following question. Include relevant examples when helpful.

    Question: {query}

    Answer:"""

    def makepredictor(prompttemplate: str):

    """Create a predictor function for a given prompt template."""

    def predict(inputs: dict) -> dict:

    llm = ChatOpenAI(model="gpt-4o", temperature=0)

    prompt = prompttemplate.format(query=inputs["query"])

    response = llm.invoke([HumanMessage(content=prompt)])

    return {"answer": response.content}

    return predict

    def qualityevaluator(run: Run, example: Example) -> dict:

    """Evaluate answer quality using an LLM judge."""

    llm = ChatOpenAI(model="gpt-4o", temperature=0)

    predicted = run.outputs.get("answer", "")

    expected = example.outputs.get("answer", "")

    evalprompt = f"""Compare the predicted answer against the reference answer.

    Rate the quality from 0 to 1.

    Reference: {expected}

    Predicted: {predicted}

    Return only a decimal number between 0 and 1."""

    response = llm.invoke([HumanMessage(content=evalprompt)])

    try:

    score = float(response.content.strip())

    except ValueError:

    score = 0.0

    return {"key": "qualityscore", "score": score}

    Run A/B test

    resultsa = evaluate(

    makepredictor(PROMPTA),

    data=datasetname,

    evaluators=[qualityevaluator],

    experimentprefix="prompt-variant-A",

    )

    resultsb = evaluate(

    makepredictor(PROMPTB),

    data=datasetname,

    evaluators=[qualityevaluator],

    experimentprefix="prompt-variant-B",

    )

    print("A/B test complete. Compare results in the LangSmith UI.")

    print(f"Variant A: {resultsa.experimenturl}")

    print(f"Variant B: {resultsb.experimenturl}")


    Best Practices

  • Always set a project name. Use LANGCHAINPROJECT to organize traces by application, environment, or team. Avoid dumping everything into the default project.
  • Use @traceable for custom logic. Any function that performs meaningful work (API calls, data transformations, business logic) should be traced to provide full visibility.
  • Build evaluation datasets incrementally. Start with 20-30 high-quality examples and grow your dataset over time. Include edge cases and failure modes.
  • Combine multiple evaluators. Use semantic similarity, relevance, faithfulness, and custom business-logic evaluators together for comprehensive assessment.
  • Version your prompts in LangSmith Hub. Never change prompts in code without tracking the change. Hub provides full version history and easy rollback.
  • Set up cost tracking early. Monitor token usage and costs from day one. Small inefficiencies compound quickly at scale.
  • Review error traces weekly. Regularly inspect failed runs to identify patterns. Common issues include context window overflow, malformed tool inputs, and hallucination.
  • Use tags and metadata. Add tags to runs (e.g., usertier:premium, feature:search) to enable filtered analysis.
  • Separate development and production projects. Use different LANGCHAINPROJECT values for dev, staging, and production to keep data clean.
  • Automate evaluations in CI/CD. Run evaluations as part of your deployment pipeline to catch regressions before they reach production.

  • Conclusion

    LangSmith transforms LLM application development from guesswork into a data-driven process. By tracing every call, evaluating against curated datasets, versioning prompts, and monitoring production performance, you gain the confidence needed to ship and maintain reliable AI applications.

    The key takeaway is to treat observability as a first-class concern from the start of your project, not an afterthought. Set up tracing on day one, build evaluation datasets as you develop, and establish cost monitoring before you scale. LangSmith provides all the tools you need to make this workflow seamless.

    Start by integrating LangSmith into your existing LangChain project, build a small evaluation dataset from real user queries, and iterate on your prompts with data backing every decision.

    Related Articles

    LangFuse: Open-Source Platform for LLM Application Observability

    LangFuse: Platform Open-Source untuk Observability Aplikasi LLM Seiring semakin banyaknya perusahaan yang mengadopsi Lar...

    LangChain Tutorial: The Most Popular Framework for Building LLM Applications

    Tutorial LangChain: Framework Paling Populer untuk Membangun Aplikasi LLM LangChain adalah framework open-source yang di...

    Complete LangGraph Tutorial: Building Complex AI Agents

    Tutorial Lengkap LangGraph: Membangun AI Agents yang Kompleks LangGraph adalah library dari LangChain untuk membangun st...

    DSPy: Stop Hand-Tuning Prompts, Let the Compiler Optimize Them

    DSPy: Berhenti Ngoprek Prompt Manual, Biarkan Compiler yang Optimasi Halo temen-temen, kali ini aku mau ngenalin satu li...