Tutorial 8: LangSmith - LLM Observability and Debugging
Table of Contents
Introduction
Building LLM-powered applications is only half the battle. Understanding how your models behave in production, why they fail, how much they cost, and whether your prompts are actually improving over time is the other critical half. LangSmith is the observability and evaluation platform built by LangChain specifically to address these challenges.
LangSmith provides end-to-end tracing of every LLM call, chain execution, and agent step. It enables you to build evaluation datasets, run automated evaluations, version your prompts, monitor production performance, and debug complex multi-step workflows. Whether you are prototyping a simple chatbot or running a production RAG pipeline, LangSmith gives you the visibility you need to ship reliable AI applications.
This tutorial walks you through every major feature of LangSmith with practical, production-ready code examples.
Prerequisites
- Python 3.9 or higher
- An active LangSmith account (sign up at smith.langchain.com)
- An OpenAI API key (or any supported LLM provider)
- Basic familiarity with LangChain concepts
Install the required packages:
pip install langsmith langchain langchain-openai openai
Setting Up LangSmith
Step 1: Configure Environment Variables
LangSmith requires a few environment variables to connect your application to the platform.
import os
os.environ["LANGCHAINTRACINGV2"] = "true"
os.environ["LANGCHAINAPIKEY"] = "ls_yourapikeyhere"
os.environ["LANGCHAINPROJECT"] = "my-first-project"
os.environ["LANGCHAINENDPOINT"] = "https://api.smith.langchain.com"
Step 2: Verify the Connection
from langsmith import Client
client = Client()
List existing projects to verify connectivity
projects = list(client.listprojects())
print(f"Connected to LangSmith. Found {len(projects)} project(s).")
for project in projects:
print(f" - {project.name} (created: {project.createdat})")
Step 3: Create a Dedicated Project
# Create a new project for organized tracking
projectname = "tutorial-langsmith-demo"
try:
project = client.createproject(projectname)
print(f"Project '{project.name}' created successfully.")
except Exception as e:
print(f"Project may already exist: {e}")
Set the active project
os.environ["LANGCHAINPROJECT"] = projectname
Tracing LLM Calls
Basic Tracing with LangChain
When tracing is enabled, every LangChain call is automatically captured.
from langchainopenai import ChatOpenAI
from langchaincore.messages import HumanMessage, SystemMessage
llm = ChatOpenAI(model="gpt-4o", temperature=0.7)
messages = [
SystemMessage(content="You are a helpful coding assistant."),
HumanMessage(content="Explain Python decorators in 3 sentences.")
]
response = llm.invoke(messages)
print(response.content)
The trace is automatically sent to LangSmith
Manual Tracing with the @traceable Decorator
For custom functions that are not LangChain components, use the @traceable decorator.
from langsmith import traceable
@traceable(name="processuserquery")
def processuserquery(query: str, context: str = "") -> dict:
"""Process a user query with optional context."""
llm = ChatOpenAI(model="gpt-4o", temperature=0)
prompt = f"Context: {context}\n\nQuestion: {query}" if context else query
response = llm.invoke([HumanMessage(content=prompt)])
return {
"query": query,
"response": response.content,
"model": "gpt-4o",
"hascontext": bool(context)
}
Every call is traced with inputs, outputs, and metadata
result = processuserquery(
query="What is dependency injection?",
context="We are discussing software design patterns."
)
print(result["response"])
Nested Tracing
Nested function calls create a trace tree, making it easy to see the full execution flow.
@traceable(name="retrievedocuments")
def retrieve
documents(query: str) -> list:
"""Simulate document retrieval."""
return [
{"id": 1, "text": "Python is a high-level programming language."},
{"id": 2, "text": "Python supports multiple programming paradigms."},
]
@traceable(name="generateanswer")
def generateanswer(query: str, documents: list) -> str:
"""Generate answer from retrieved documents."""
llm = ChatOpenAI(model="gpt-4o", temperature=0)
context = "\n".join([doc["text"] for doc in documents])
prompt = f"Based on the following context:\n{context}\n\nAnswer: {query}"
response = llm.invoke([HumanMessage(content=prompt)])
return response.content
@traceable(name="ragpipeline")
def ragpipeline(query: str) -> str:
"""Full RAG pipeline with nested tracing."""
docs = retrievedocuments(query)
answer = generateanswer(query, docs)
return answer
answer = ragpipeline("What programming paradigms does Python support?")
print(answer)
Evaluation Datasets
Creating a Dataset
from langsmith import Client
client = Client()
Create a dataset for evaluation
datasetname = "qa-evaluation-set"
dataset = client.createdataset(
datasetname,
description="Question-answer pairs for evaluating our RAG pipeline."
)
Add examples to the dataset
examples = [
{
"inputs": {"query": "What is Python?"},
"outputs": {"answer": "Python is a high-level, interpreted programming language."}
},
{
"inputs": {"query": "What is machine learning?"},
"outputs": {"answer": "Machine learning is a subset of AI that enables systems to learn from data."}
},
{
"inputs": {"query": "What is a neural network?"},
"outputs": {"answer": "A neural network is a computational model inspired by the human brain."}
},
{
"inputs": {"query": "What is natural language processing?"},
"outputs": {"answer": "NLP is a field of AI focused on interaction between computers and human language."}
},
]
for example in examples:
client.createexample(
inputs=example["inputs"],
outputs=example["outputs"],
datasetid=dataset.id
)
print(f"Dataset '{datasetname}' created with {len(examples)} examples.")
Running Evaluations Against a Dataset
from langsmith.evaluation import evaluate
def predictanswer(inputs: dict) -> dict:
"""The function to evaluate."""
llm = ChatOpenAI(model="gpt-4o", temperature=0)
response = llm.invoke([HumanMessage(content=inputs["query"])])
return {"answer": response.content}
results = evaluate(
predictanswer,
data=datasetname,
experimentprefix="gpt4o-baseline",
maxconcurrency=2,
)
print(f"Evaluation complete. View results at: {results.experimenturl}")
Custom Evaluators
Building a Semantic Similarity Evaluator
from langsmith.schemas import Example, Run
def semanticsimilarityevaluator(run: Run, example: Example) -> dict:
"""Evaluate semantic similarity between predicted and expected answers."""
from langchainopenai import OpenAIEmbeddings
import numpy as np
embeddings = OpenAIEmbeddings(model="text-embedding-3-small")
predicted = run.outputs.get("answer", "")
expected = example.outputs.get("answer", "")
predembedding = embeddings.embedquery(predicted)
expembedding = embeddings.embedquery(expected)
similarity = np.dot(predembedding, expembedding) / (
np.linalg.norm(predembedding) np.linalg.norm(expembedding)
)
return {
"key": "semanticsimilarity",
"score": float(similarity),
"comment": f"Cosine similarity: {similarity:.4f}"
}
def answerrelevanceevaluator(run: Run, example: Example) -> dict:
"""Check if the answer is relevant to the question."""
llm = ChatOpenAI(model="gpt-4o", temperature=0)
query = example.inputs.get("query", "")
predicted = run.outputs.get("answer", "")
evalprompt = f"""Rate how relevant the following answer is to the question.
Question: {query}
Answer: {predicted}
Return only a number between 0 and 1, where 1 means perfectly relevant."""
response = llm.invoke([HumanMessage(content=evalprompt)])
try:
score = float(response.content.strip())
except ValueError:
score = 0.0
return {
"key": "answerrelevance",
"score": score,
}
Run evaluation with custom evaluators
results = evaluate(
predictanswer,
data=datasetname,
evaluators=[semanticsimilarityevaluator, answerrelevanceevaluator],
experimentprefix="gpt4o-custom-eval",
maxconcurrency=2,
)
Prompt Versioning
Managing Prompts in LangSmith Hub
from langchain import hub
Push a prompt to LangSmith Hub
from langchaincore.prompts import ChatPromptTemplate
qaprompt = ChatPromptTemplate.frommessages([
("system", "You are an expert assistant. Answer questions accurately and concisely. "
"If you don't know the answer, say so."),
("human", "Context: {context}\n\nQuestion: {question}")
])
Push the prompt with a commit message
hub.push("my-org/qa-prompt", qaprompt, newrepoispublic=False)
print("Prompt pushed to LangSmith Hub.")
Pull a specific version
promptv1 = hub.pull("my-org/qa-prompt")
print(f"Pulled prompt with {len(promptv1.messages)} messages.")
Update the prompt (creates a new version automatically)
qapromptv2 = ChatPromptTemplate.frommessages([
("system", "You are an expert assistant specializing in technology topics. "
"Provide detailed, well-structured answers with examples when possible. "
"If uncertain, express your confidence level."),
("human", "Context: {context}\n\nQuestion: {question}")
])
hub.push("my-org/qa-prompt", qapromptv2)
print("Prompt v2 pushed. Previous version is preserved in history.")
Production Monitoring
Setting Up Monitoring Rules
from langsmith import Client
client = Client()
Query recent runs for monitoring
runs = client.listruns(
projectname="tutorial-langsmith-demo",
executionorder=1, # Top-level runs only
error=False,
limit=100,
)
Analyze latency and token usage
latencies = []
totaltokens = []
for run in runs:
if run.endtime and run.starttime:
latency = (run.endtime - run.starttime).totalseconds()
latencies.append(latency)
if run.totaltokens:
totaltokens.append(run.totaltokens)
if latencies:
avglatency = sum(latencies) / len(latencies)
maxlatency = max(latencies)
p95latency = sorted(latencies)[int(len(latencies) 0.95)]
print(f"Average latency: {avglatency:.2f}s")
print(f"P95 latency: {p95latency:.2f}s")
print(f"Max latency: {maxlatency:.2f}s")
if totaltokens:
avgtokens = sum(totaltokens) / len(totaltokens)
print(f"Average tokens per call: {avgtokens:.0f}")
Automated Alerting
import smtplib
from email.mime.text import MIMEText
@traceable(name="monitoredpipeline")
def monitoredpipeline(query: str) -> dict:
"""Pipeline with built-in monitoring."""
import time
start = time.time()
llm = ChatOpenAI(model="gpt-4o", temperature=0)
response = llm.invoke([HumanMessage(content=query)])
elapsed = time.time() - start
result = {
"answer": response.content,
"latencyseconds": elapsed,
"tokencount": response.usagemetadata.get("totaltokens", 0)
if hasattr(response, "usagemetadata") and response.usagemetadata
else 0,
}
# Alert if latency exceeds threshold
if elapsed > 10.0:
print(f"[ALERT] High latency detected: {elapsed:.2f}s for query: {query[:50]}...")
return result
Debugging Chains and Agents
Inspecting Failed Runs
# Fetch runs that resulted in errors
errorruns = client.listruns(
projectname="tutorial-langsmith-demo",
error=True,
limit=20,
)
for run in errorruns:
print(f"Run ID: {run.id}")
print(f" Name: {run.name}")
print(f" Error: {run.error}")
print(f" Inputs: {run.inputs}")
print(f" Start: {run.starttime}")
print("---")
Debugging Agent Execution Step-by-Step
from langchain.agents import AgentExecutor, createopenaitoolsagent
from langchain
core.tools import tool
from langchaincore.prompts import ChatPromptTemplate, MessagesPlaceholder
@tool
def calculate(expression: str) -> str:
"""Evaluate a mathematical expression."""
try:
result = eval(expression, {"builtins": {}}, {})
return str(result)
except Exception as e:
return f"Error: {e}"
@tool
def searchknowledge(query: str) -> str:
"""Search the knowledge base for information."""
knowledge = {
"python": "Python is a versatile programming language created by Guido van Rossum.",
"javascript": "JavaScript is the language of the web, running in browsers and Node.js.",
}
for key, value in knowledge.items():
if key in query.lower():
return value
return "No relevant information found."
prompt = ChatPromptTemplate.frommessages([
("system", "You are a helpful assistant with access to tools."),
MessagesPlaceholder(variablename="chathistory", optional=True),
("human", "{input}"),
MessagesPlaceholder(variablename="agentscratchpad"),
])
llm = ChatOpenAI(model="gpt-4o", temperature=0)
agent = createopenaitoolsagent(llm, [calculate, searchknowledge], prompt)
agentexecutor = AgentExecutor(agent=agent, tools=[calculate, searchknowledge], verbose=True)
This entire execution is traced - every tool call, LLM call, and decision
result = agentexecutor.invoke({"input": "What is 15 23 + 7? Also, tell me about Python."})
print(result["output"])
After execution, inspect the trace in LangSmith UI to see:
1. The agent's reasoning at each step
2. Which tools were called and their inputs/outputs
3. Token usage per step
4. Total latency breakdown
Cost Tracking
Calculating Costs from Traced Runs
# Token pricing (example rates, adjust to current pricing)
PRICING = {
"gpt-4o": {"input": 2.50 / 1000000, "output": 10.00 / 1000000},
"gpt-4o-mini": {"input": 0.15 / 1000000, "output": 0.60 / 1000000},
"gpt-3.5-turbo": {"input": 0.50 / 1000000, "output": 1.50 / 1000000},
}
def calculateprojectcosts(projectname: str, days: int = 7) -> dict:
"""Calculate costs for a project over the specified number of days."""
from datetime import datetime, timedelta
client = Client()
startdate = datetime.now() - timedelta(days=days)
runs = client.listruns(
projectname=projectname,
runtype="llm",
starttime=startdate,
)
totalcost = 0.0
modelcosts = {}
for run in runs:
model = run.extra.get("metadata", {}).get("lsmodelname", "unknown") if run.extra else "unknown"
inputtokens = run.prompttokens or 0
outputtokens = run.completiontokens or 0
if model in PRICING:
cost = (inputtokens PRICING[model]["input"] +
outputtokens * PRICING[model]["output"])
else:
cost = 0.0
totalcost += cost
if model not in modelcosts:
modelcosts[model] = {"cost": 0.0, "calls": 0, "tokens": 0}
modelcosts[model]["cost"] += cost
modelcosts[model]["calls"] += 1
modelcosts[model]["tokens"] += inputtokens + outputtokens
print(f"Cost report for '{projectname}' (last {days} days):")
print(f" Total cost: ${totalcost:.4f}")
for model, data in modelcosts.items():
print(f" {model}: ${data['cost']:.4f} ({data['calls']} calls, {data['tokens']} tokens)")
return {"totalcost": totalcost, "modelcosts": modelcosts}
costs = calculateprojectcosts("tutorial-langsmith-demo", days=30)
Integration with LangChain
Full RAG Pipeline with LangSmith Tracing
from langchainopenai import ChatOpenAI, OpenAIEmbeddings
from langchaincore.prompts import ChatPromptTemplate
from langchaincore.outputparsers import StrOutputParser
from langchaincore.runnables import RunnablePassthrough
All LangChain components are automatically traced
embeddings = OpenAIEmbeddings(model="text-embedding-3-small")
llm = ChatOpenAI(model="gpt-4o", temperature=0)
Simulated vector store retriever
@traceable(name="vectorretriever")
def retrieve(query: str) -> str:
"""Simulate vector store retrieval."""
docs = {
"ai": "Artificial intelligence encompasses machine learning, deep learning, and NLP.",
"ml": "Machine learning algorithms learn patterns from data without explicit programming.",
}
results = []
for key, value in docs.items():
if key in query.lower():
results.append(value)
return "\n".join(results) if results else "No relevant documents found."
prompt = ChatPromptTemplate.fromtemplate(
"Answer the question based on the context.\n\n"
"Context: {context}\n\n"
"Question: {question}\n\n"
"Answer:"
)
Build the chain - every step is traced
chain = (
{"context": lambda x: retrieve(x["question"]), "question": lambda x: x["question"]}
| prompt
| llm
| StrOutputParser()
)
Execute and trace
answer = chain.invoke({"question": "What is machine learning?"})
print(answer)
A/B Testing Prompts
Comparing Prompt Variants
from langsmith.evaluation import evaluate
Define prompt variants
PROMPTA = "Answer the following question concisely: {query}"
PROMPTB = """You are an expert assistant. Provide a clear, well-structured answer
to the following question. Include relevant examples when helpful.
Question: {query}
Answer:"""
def makepredictor(prompttemplate: str):
"""Create a predictor function for a given prompt template."""
def predict(inputs: dict) -> dict:
llm = ChatOpenAI(model="gpt-4o", temperature=0)
prompt = prompttemplate.format(query=inputs["query"])
response = llm.invoke([HumanMessage(content=prompt)])
return {"answer": response.content}
return predict
def qualityevaluator(run: Run, example: Example) -> dict:
"""Evaluate answer quality using an LLM judge."""
llm = ChatOpenAI(model="gpt-4o", temperature=0)
predicted = run.outputs.get("answer", "")
expected = example.outputs.get("answer", "")
evalprompt = f"""Compare the predicted answer against the reference answer.
Rate the quality from 0 to 1.
Reference: {expected}
Predicted: {predicted}
Return only a decimal number between 0 and 1."""
response = llm.invoke([HumanMessage(content=evalprompt)])
try:
score = float(response.content.strip())
except ValueError:
score = 0.0
return {"key": "qualityscore", "score": score}
Run A/B test
resultsa = evaluate(
makepredictor(PROMPTA),
data=datasetname,
evaluators=[qualityevaluator],
experimentprefix="prompt-variant-A",
)
resultsb = evaluate(
makepredictor(PROMPTB),
data=datasetname,
evaluators=[qualityevaluator],
experimentprefix="prompt-variant-B",
)
print("A/B test complete. Compare results in the LangSmith UI.")
print(f"Variant A: {resultsa.experimenturl}")
print(f"Variant B: {resultsb.experimenturl}")
Best Practices
LANGCHAINPROJECT to organize traces by application, environment, or team. Avoid dumping everything into the default project.@traceable for custom logic. Any function that performs meaningful work (API calls, data transformations, business logic) should be traced to provide full visibility.usertier:premium, feature:search) to enable filtered analysis.LANGCHAINPROJECT values for dev, staging, and production to keep data clean.Conclusion
LangSmith transforms LLM application development from guesswork into a data-driven process. By tracing every call, evaluating against curated datasets, versioning prompts, and monitoring production performance, you gain the confidence needed to ship and maintain reliable AI applications.
The key takeaway is to treat observability as a first-class concern from the start of your project, not an afterthought. Set up tracing on day one, build evaluation datasets as you develop, and establish cost monitoring before you scale. LangSmith provides all the tools you need to make this workflow seamless.
Start by integrating LangSmith into your existing LangChain project, build a small evaluation dataset from real user queries, and iterate on your prompts with data backing every decision.