DSPy: A Framework for Programmatic LLM Optimization
Manual prompt engineering is a tedious and hard-to-maintain process. Every time a model changes or data shifts, prompts need to be rewritten. DSPy (Declarative Self-improving Python) provides a revolutionary solution: a framework that lets you programmatically optimize LLM prompts, much like training a machine learning model.
In this tutorial, we will learn how to use DSPy to build LLM pipelines that can be automatically optimized, from basic concepts to a complete RAG (Retrieval-Augmented Generation) implementation.
What Is DSPy?
DSPy is a Python framework that fundamentally changes how we work with LLMs. Instead of writing prompts manually, DSPy allows you to define what you want to achieve (through signatures and modules), then automatically optimizes how to achieve it (through optimizers/teleprompters).
A simple analogy: if manual prompt engineering is like writing assembly code, DSPy is like using a compiler that automatically optimizes your code.
Core concepts of DSPy:
- Signatures: Define the input and output of a task
- Modules: Building blocks that implement prompting strategies
- Optimizers (Teleprompters): Algorithms that automatically optimize prompts
- Metrics: Evaluation functions to measure output quality
- Compilation: The process of combining all components and optimizing
Installation
Install DSPy and its required dependencies:
pip install dspy-ai openai
For additional features:
# For retrieval with ChromaDB
pip install dspy-ai chromadb
For using local models
pip install dspy-ai transformers torch
Configure your API key:
export OPENAIAPIKEY="sk-your-api-key-here"
Basic Setup and Configuration
The first step is configuring the LLM you will use.
import dspy
Configure with OpenAI
lm = dspy.LM("openai/gpt-4o-mini")
dspy.configure(lm=lm)
Or with other models
lm = dspy.LM("anthropic/claude-sonnet-4-20250514")
lm = dspy.LM("openai/gpt-4o")
Signatures: Defining Tasks
Signatures are how DSPy defines the input and output of a task. They are the most fundamental abstraction in DSPy.
Inline Signatures (simple)# Format: "input -> output"
Classify sentiment
classify = dspy.Predict("text -> sentiment")
result = classify(text="This product is amazing and high quality!")
print(result.sentiment) # positive
Question answering
qa = dspy.Predict("question -> answer")
result = qa(question="What is the capital of France?")
print(result.answer) # Paris
Summarization
summarize = dspy.Predict("document -> summary")
result = summarize(
document="A long article about AI and its impact..."
)
print(result.summary)
Class-based Signatures (more expressive)
class SentimentAnalysis(dspy.Signature):
"""Analyze sentiment from product review text."""
text: str = dspy.InputField(desc="Product review text")
sentiment: str = dspy.OutputField(
desc="Sentiment: positive, negative, or neutral"
)
confidence: float = dspy.OutputField(
desc="Confidence level between 0.0 and 1.0"
)
reasoning: str = dspy.OutputField(
desc="Reasoning behind the sentiment classification"
)
Use the signature
predictor = dspy.Predict(SentimentAnalysis)
result = predictor(
text="Item arrived on time, quality matches description. "
"But the packaging was sloppy."
)
print(f"Sentiment: {result.sentiment}")
print(f"Confidence: {result.confidence}")
print(f"Reasoning: {result.reasoning}")
Modules: Building Blocks
DSPy provides several built-in modules that implement different prompting strategies.
dspy.Predict
The most basic module that performs a single LLM call.
predictor = dspy.Predict("question -> answer")
result = predictor(question="Explain the concept of machine learning")
print(result.answer)
dspy.ChainOfThought
Adds a reasoning step before producing output, automatically implementing Chain-of-Thought prompting.
class MathProblem(dspy.Signature):
"""Solve the math problem step by step."""
problem: str = dspy.InputField(desc="Math problem")
answer: float = dspy.OutputField(desc="Numerical answer")
With ChainOfThought, the model will think step-by-step
cot = dspy.ChainOfThought(MathProblem)
result = cot(
problem="A store sells 150 items at $25 each. "
"There is a 15% discount. "
"What is the total revenue?"
)
print(f"Reasoning: {result.reasoning}")
print(f"Answer: {result.answer}")
dspy.ReAct
Implements the Reasoning + Acting pattern, where the model can use tools to solve tasks.
def searchdatabase(query: str) -> str:
"""Simulate database search."""
data = {
"python": "Python is a high-level programming language",
"javascript": "JavaScript is a programming language for the web",
"rust": "Rust is a safe systems programming language",
}
for key, value in data.items():
if key in query.lower():
return value
return "Data not found"
def calculate(expression: str) -> str:
"""Calculate a mathematical expression."""
try:
return str(eval(expression))
except Exception:
return "Calculation error"
ReAct with tools
react = dspy.ReAct(
"question -> answer",
tools=[search
database, calculate],
)
result = react(
question="What is Python and what is 15 23?"
)
print(result.answer)
dspy.ChainOfThoughtWithHint
A version of ChainOfThought that accepts additional hints to guide reasoning.
cothint = dspy.ChainOfThoughtWithHint("question -> answer")
result = cot
hint(
question="How to optimize a slow SQL query?",
hint="Consider indexing, query plans, and data normalization"
)
print(result.answer)
Building Custom Modules
You can create custom modules by extending dspy.Module.
class ArticleGenerator(dspy.Module):
def init(self):
super().init()
self.outline = dspy.ChainOfThought(
"topic, audience -> outline"
)
self.draft = dspy.ChainOfThought(
"topic, outline -> article"
)
self.review = dspy.Predict(
"article -> improvedarticle, feedback"
)
def forward(self, topic: str, audience: str):
# Step 1: Create outline
outlineresult = self.outline(
topic=topic, audience=audience
)
# Step 2: Write draft based on outline
draftresult = self.draft(
topic=topic,
outline=outlineresult.outline
)
# Step 3: Review and improve
finalresult = self.review(
article=draftresult.article
)
return dspy.Prediction(
outline=outlineresult.outline,
article=finalresult.improvedarticle,
feedback=finalresult.feedback,
)
Use the module
generator = ArticleGenerator()
result = generator(
topic="Applications of AI in Agriculture",
audience="Farmers and agricultural professionals"
)
print(f"Outline: {result.outline}")
print(f"Article: {result.article}")
Metrics: Measuring Quality
Metrics are functions that evaluate the quality of outputs. They are essential for the optimization process.
# Simple metric: exact match
def exactmatch(example, prediction, trace=None):
return example.answer.lower() == prediction.answer.lower()
Metric with partial scoring
def qualitymetric(example, prediction, trace=None):
score = 0.0
# Check if the answer contains important keywords
keywords = example.get("keywords", [])
if keywords:
matches = sum(
1 for kw in keywords
if kw.lower() in prediction.answer.lower()
)
score += matches / len(keywords) 0.5
# Check answer length (not too short/long)
wordcount = len(prediction.answer.split())
if 50 <= wordcount <= 300:
score += 0.3
elif 20 <= wordcount < 50:
score += 0.15
# Check if reasoning is present
if hasattr(prediction, "reasoning") and prediction.reasoning:
score += 0.2
return score
Metric using LLM as judge
class AnswerQuality(dspy.Signature):
"""Evaluate answer quality based on the question and reference."""
question: str = dspy.InputField()
referenceanswer: str = dspy.InputField()
predictedanswer: str = dspy.InputField()
qualityscore: float = dspy.OutputField(
desc="Quality score between 0.0 and 1.0"
)
def llmjudgemetric(example, prediction, trace=None):
judge = dspy.Predict(AnswerQuality)
result = judge(
question=example.question,
referenceanswer=example.answer,
predictedanswer=prediction.answer,
)
return float(result.qualityscore)
Optimizers (Teleprompters): Automatic Optimization
Optimizers are the component that sets DSPy apart from other frameworks. They automatically optimize prompts based on training data and metrics.
Preparing Training Data
# Create training dataset
trainset = [
dspy.Example(
question="What is machine learning?",
answer="Machine learning is a branch of artificial intelligence "
"that enables computers to learn from data "
"without being explicitly programmed."
).withinputs("question"),
dspy.Example(
question="Explain the difference between supervised "
"and unsupervised learning",
answer="Supervised learning uses labeled data for training, "
"while unsupervised learning discovers patterns "
"from unlabeled data."
).withinputs("question"),
dspy.Example(
question="What is the purpose of activation functions "
"in neural networks?",
answer="Activation functions determine a neuron's output "
"based on its input, adding non-linearity "
"to the model."
).withinputs("question"),
# Add more examples...
]
devset = [
dspy.Example(
question="What is transfer learning?",
answer="Transfer learning is a technique of using a model "
"trained on one task as a starting point "
"for a different but related task."
).withinputs("question"),
]
BootstrapFewShot
An optimizer that automatically selects the best few-shot examples.
from dspy.teleprompt import BootstrapFewShot
Define module
qamodule = dspy.ChainOfThought("question -> answer")
Define metric
def answerquality(example, prediction, trace=None):
# Simple: check if prediction is long enough and relevant
if len(prediction.answer.split()) < 10:
return False
return True
Compile with BootstrapFewShot
optimizer = BootstrapFewShot(
metric=answerquality,
maxbootstrappeddemos=4,
maxlabeleddemos=4,
)
compiledqa = optimizer.compile(
qamodule,
trainset=trainset,
)
Use the optimized module
result = compiledqa(question="What is deep learning?")
print(result.answer)
BootstrapFewShotWithRandomSearch
A variant that tries various combinations of examples and selects the best one.
from dspy.teleprompt import BootstrapFewShotWithRandomSearch
optimizer = BootstrapFewShotWithRandomSearch(
metric=answerquality,
maxbootstrappeddemos=4,
maxlabeleddemos=4,
numcandidateprograms=10,
numthreads=4,
)
compiledqa = optimizer.compile(
qamodule,
trainset=trainset,
valset=devset,
)
MIPROv2
An advanced optimizer that simultaneously optimizes instructions and few-shot examples.
from dspy.teleprompt import MIPROv2
optimizer = MIPROv2(
metric=answerquality,
numcandidates=7,
inittemperature=1.0,
)
compiledqa = optimizer.compile(
qamodule,
trainset=trainset,
maxbootstrappeddemos=4,
maxlabeleddemos=4,
)
Saving and Loading Compiled Programs
After optimization, you can save and reload compiled programs.
# Save the optimized program
compiledqa.save("optimizedqaprogram.json")
Load it back
loadedqa = dspy.ChainOfThought("question -> answer")
loadedqa.load("optimizedqaprogram.json")
Use as usual
result = loadedqa(question="What is a neural network?")
print(result.answer)
Practical Example: RAG Pipeline with DSPy
Here is a complete RAG (Retrieval-Augmented Generation) implementation using DSPy.
import dspy
from dspy.retrieve.chromadbrm import ChromadbRM
import chromadb
Setup ChromaDB
chromaclient = chromadb.PersistentClient(path="./chromadb")
Create collection and add documents
collection = chromaclient.getorcreatecollection("knowledgebase")
Add documents
documents = [
"Python is a programming language created by Guido van Rossum "
"in 1991. Python is known for its readable and easy-to-learn "
"syntax.",
"Machine learning has three main paradigms: supervised learning, "
"unsupervised learning, and reinforcement learning. Each has "
"different use cases.",
"Deep learning is a subset of machine learning that uses "
"neural networks with many layers. Popular frameworks include "
"TensorFlow, PyTorch, and JAX.",
"Natural Language Processing (NLP) is a branch of AI that focuses "
"on the interaction between computers and human language. NLP tasks "
"include sentiment analysis, named entity recognition, "
"and machine translation.",
"MLOps is a practice that combines Machine Learning and DevOps "
"to automate and standardize the ML model lifecycle "
"from development to production.",
]
collection.add(
documents=documents,
ids=[f"doc{i}" for i in range(len(documents))],
)
Setup retriever
retriever = ChromadbRM(
collectionname="knowledgebase",
persistdirectory="./chromadb",
k=3,
)
dspy.configure(
lm=dspy.LM("openai/gpt-4o-mini"),
rm=retriever,
)
Define RAG Signature
class GenerateAnswer(dspy.Signature):
"""Answer the question based on the provided context.
Provide an informative and accurate answer."""
context: list[str] = dspy.InputField(
desc="Relevant context from the knowledge base"
)
question: str = dspy.InputField(
desc="The question to answer"
)
answer: str = dspy.OutputField(
desc="A detailed and accurate answer based on the context"
)
Build RAG Module
class RAGModule(dspy.Module):
def init(self, numpassages=3):
super().init()
self.retrieve = dspy.Retrieve(k=numpassages)
self.generate = dspy.ChainOfThought(GenerateAnswer)
def forward(self, question):
# Retrieve relevant passages
context = self.retrieve(question).passages
# Generate answer
result = self.generate(
context=context,
question=question,
)
return dspy.Prediction(
context=context,
answer=result.answer,
reasoning=result.reasoning,
)
Create instance and use
rag = RAGModule()
result = rag(question="What are the main paradigms in machine learning?")
print(f"Answer: {result.answer}")
print(f"Reasoning: {result.reasoning}")
print(f"Context used: {result.context}")
Optimize RAG with training data
trainsetrag = [
dspy.Example(
question="Who created Python?",
answer="Python was created by Guido van Rossum in 1991."
).withinputs("question"),
dspy.Example(
question="What is deep learning?",
answer="Deep learning is a subset of machine learning "
"that uses neural networks with many layers."
).withinputs("question"),
dspy.Example(
question="Name popular deep learning frameworks",
answer="Popular deep learning frameworks include TensorFlow, "
"PyTorch, and JAX."
).withinputs("question"),
]
Compile RAG module
from dspy.teleprompt import BootstrapFewShot
def ragmetric(example, prediction, trace=None):
# Check if the answer contains information from reference
refwords = set(example.answer.lower().split())
predwords = set(prediction.answer.lower().split())
overlap = len(refwords & predwords) / len(refwords)
return overlap > 0.5
optimizer = BootstrapFewShot(
metric=ragmetric,
maxbootstrappeddemos=2,
)
compiledrag = optimizer.compile(rag, trainset=trainsetrag)
Use the optimized RAG
result = compiledrag(
question="What are the tasks in NLP?"
)
print(f"Optimized Answer: {result.answer}")
Evaluating Pipelines
DSPy provides utilities for systematically evaluating pipelines.
from dspy.evaluate import Evaluate
Create evaluator
evaluator = Evaluate(
devset=devset,
metric=ragmetric,
numthreads=4,
displayprogress=True,
displaytable=5,
)
Evaluate module before optimization
scorebefore = evaluator(rag)
print(f"Score before optimization: {scorebefore}")
Evaluate module after optimization
scoreafter = evaluator(compiledrag)
print(f"Score after optimization: {scoreafter}")
Tips and Best Practices
1. Start with clear SignaturesGood descriptions in signatures help both the LLM and optimizer understand the task better.
# Good
class GoodSignature(dspy.Signature):
"""Extract business entities from financial reports."""
report: str = dspy.InputField(desc="Financial report text")
entities: list[str] = dspy.OutputField(
desc="List of company names, financial figures, and key metrics"
)
2. Use high-quality training data
Training data quality significantly affects optimization results. Make sure your examples are representative and accurate.
3. Choose the right optimizerBootstrapFewShot: Fast, great for getting startedBootstrapFewShotWithRandomSearch: Better results but slowerMIPROv2: Best for production, but slowest
Always evaluate before and after optimization. If results are unsatisfactory, try adding more training data or using a different optimizer.
5. Save compilation resultsAfter finding the optimal configuration, always save the results so you do not need to repeat the optimization process.
Conclusion
DSPy changes the paradigm of working with LLMs. Instead of spending time writing and tweaking prompts manually, DSPy allows you to focus on defining tasks and letting the framework optimize the implementation.
Key takeaways from this tutorial:
- Signatures for defining task input/output
- Modules like ChainOfThought and ReAct for prompting strategies
- Optimizers for automatically tuning prompts
- Metrics for measuring output quality
- A complete RAG pipeline implementation
- Evaluation and saving/loading compiled programs
With DSPy, you can build LLM pipelines that are more robust, maintainable, and performant. Start with simple tasks, optimize incrementally, and scale as your project demands.