DSPy: A Framework for Programmatic LLM Optimization

# DSPy: Framework untuk Optimasi LLM Secara Programatik Prompt engineering secara manual adalah proses yang melelahkan dan sulit di-maintain. Setiap kali model berubah atau data berubah, prompt harus...

By Ruby Abdullah · · tutorial
DSPyLLMPrompt OptimizationMachine LearningPython

DSPy: A Framework for Programmatic LLM Optimization

Manual prompt engineering is a tedious and hard-to-maintain process. Every time a model changes or data shifts, prompts need to be rewritten. DSPy (Declarative Self-improving Python) provides a revolutionary solution: a framework that lets you programmatically optimize LLM prompts, much like training a machine learning model.

In this tutorial, we will learn how to use DSPy to build LLM pipelines that can be automatically optimized, from basic concepts to a complete RAG (Retrieval-Augmented Generation) implementation.

What Is DSPy?

DSPy is a Python framework that fundamentally changes how we work with LLMs. Instead of writing prompts manually, DSPy allows you to define what you want to achieve (through signatures and modules), then automatically optimizes how to achieve it (through optimizers/teleprompters).

A simple analogy: if manual prompt engineering is like writing assembly code, DSPy is like using a compiler that automatically optimizes your code.

Core concepts of DSPy:

  • Signatures: Define the input and output of a task
  • Modules: Building blocks that implement prompting strategies
  • Optimizers (Teleprompters): Algorithms that automatically optimize prompts
  • Metrics: Evaluation functions to measure output quality
  • Compilation: The process of combining all components and optimizing

Installation

Install DSPy and its required dependencies:

pip install dspy-ai openai

For additional features:

# For retrieval with ChromaDB

pip install dspy-ai chromadb

For using local models

pip install dspy-ai transformers torch

Configure your API key:

export OPENAIAPIKEY="sk-your-api-key-here"

Basic Setup and Configuration

The first step is configuring the LLM you will use.

import dspy

Configure with OpenAI

lm = dspy.LM("openai/gpt-4o-mini")

dspy.configure(lm=lm)

Or with other models

lm = dspy.LM("anthropic/claude-sonnet-4-20250514")

lm = dspy.LM("openai/gpt-4o")

Signatures: Defining Tasks

Signatures are how DSPy defines the input and output of a task. They are the most fundamental abstraction in DSPy.

Inline Signatures (simple)
# Format: "input -> output"

Classify sentiment

classify = dspy.Predict("text -> sentiment")

result = classify(text="This product is amazing and high quality!")

print(result.sentiment) # positive

Question answering

qa = dspy.Predict("question -> answer")

result = qa(question="What is the capital of France?")

print(result.answer) # Paris

Summarization

summarize = dspy.Predict("document -> summary")

result = summarize(

document="A long article about AI and its impact..."

)

print(result.summary)

Class-based Signatures (more expressive)
class SentimentAnalysis(dspy.Signature):

"""Analyze sentiment from product review text."""

text: str = dspy.InputField(desc="Product review text")

sentiment: str = dspy.OutputField(

desc="Sentiment: positive, negative, or neutral"

)

confidence: float = dspy.OutputField(

desc="Confidence level between 0.0 and 1.0"

)

reasoning: str = dspy.OutputField(

desc="Reasoning behind the sentiment classification"

)

Use the signature

predictor = dspy.Predict(SentimentAnalysis)

result = predictor(

text="Item arrived on time, quality matches description. "

"But the packaging was sloppy."

)

print(f"Sentiment: {result.sentiment}")

print(f"Confidence: {result.confidence}")

print(f"Reasoning: {result.reasoning}")

Modules: Building Blocks

DSPy provides several built-in modules that implement different prompting strategies.

dspy.Predict

The most basic module that performs a single LLM call.

predictor = dspy.Predict("question -> answer")

result = predictor(question="Explain the concept of machine learning")

print(result.answer)

dspy.ChainOfThought

Adds a reasoning step before producing output, automatically implementing Chain-of-Thought prompting.

class MathProblem(dspy.Signature):

"""Solve the math problem step by step."""

problem: str = dspy.InputField(desc="Math problem")

answer: float = dspy.OutputField(desc="Numerical answer")

With ChainOfThought, the model will think step-by-step

cot = dspy.ChainOfThought(MathProblem)

result = cot(

problem="A store sells 150 items at $25 each. "

"There is a 15% discount. "

"What is the total revenue?"

)

print(f"Reasoning: {result.reasoning}")

print(f"Answer: {result.answer}")

dspy.ReAct

Implements the Reasoning + Acting pattern, where the model can use tools to solve tasks.

def searchdatabase(query: str) -> str:

"""Simulate database search."""

data = {

"python": "Python is a high-level programming language",

"javascript": "JavaScript is a programming language for the web",

"rust": "Rust is a safe systems programming language",

}

for key, value in data.items():

if key in query.lower():

return value

return "Data not found"

def calculate(expression: str) -> str:

"""Calculate a mathematical expression."""

try:

return str(eval(expression))

except Exception:

return "Calculation error"

ReAct with tools

react = dspy.ReAct(

"question -> answer",

tools=[searchdatabase, calculate],

)

result = react(

question="What is Python and what is 15 23?"

)

print(result.answer)

dspy.ChainOfThoughtWithHint

A version of ChainOfThought that accepts additional hints to guide reasoning.

cothint = dspy.ChainOfThoughtWithHint("question -> answer")

result = cothint(

question="How to optimize a slow SQL query?",

hint="Consider indexing, query plans, and data normalization"

)

print(result.answer)

Building Custom Modules

You can create custom modules by extending dspy.Module.

class ArticleGenerator(dspy.Module):

def init(self):

super().init()

self.outline = dspy.ChainOfThought(

"topic, audience -> outline"

)

self.draft = dspy.ChainOfThought(

"topic, outline -> article"

)

self.review = dspy.Predict(

"article -> improvedarticle, feedback"

)

def forward(self, topic: str, audience: str):

# Step 1: Create outline

outlineresult = self.outline(

topic=topic, audience=audience

)

# Step 2: Write draft based on outline

draftresult = self.draft(

topic=topic,

outline=outlineresult.outline

)

# Step 3: Review and improve

finalresult = self.review(

article=draftresult.article

)

return dspy.Prediction(

outline=outlineresult.outline,

article=finalresult.improvedarticle,

feedback=finalresult.feedback,

)

Use the module

generator = ArticleGenerator()

result = generator(

topic="Applications of AI in Agriculture",

audience="Farmers and agricultural professionals"

)

print(f"Outline: {result.outline}")

print(f"Article: {result.article}")

Metrics: Measuring Quality

Metrics are functions that evaluate the quality of outputs. They are essential for the optimization process.

# Simple metric: exact match

def exactmatch(example, prediction, trace=None):

return example.answer.lower() == prediction.answer.lower()

Metric with partial scoring

def qualitymetric(example, prediction, trace=None):

score = 0.0

# Check if the answer contains important keywords

keywords = example.get("keywords", [])

if keywords:

matches = sum(

1 for kw in keywords

if kw.lower() in prediction.answer.lower()

)

score += matches / len(keywords) 0.5

# Check answer length (not too short/long)

wordcount = len(prediction.answer.split())

if 50 <= wordcount <= 300:

score += 0.3

elif 20 <= wordcount < 50:

score += 0.15

# Check if reasoning is present

if hasattr(prediction, "reasoning") and prediction.reasoning:

score += 0.2

return score

Metric using LLM as judge

class AnswerQuality(dspy.Signature):

"""Evaluate answer quality based on the question and reference."""

question: str = dspy.InputField()

referenceanswer: str = dspy.InputField()

predictedanswer: str = dspy.InputField()

qualityscore: float = dspy.OutputField(

desc="Quality score between 0.0 and 1.0"

)

def llmjudgemetric(example, prediction, trace=None):

judge = dspy.Predict(AnswerQuality)

result = judge(

question=example.question,

referenceanswer=example.answer,

predictedanswer=prediction.answer,

)

return float(result.qualityscore)

Optimizers (Teleprompters): Automatic Optimization

Optimizers are the component that sets DSPy apart from other frameworks. They automatically optimize prompts based on training data and metrics.

Preparing Training Data

# Create training dataset

trainset = [

dspy.Example(

question="What is machine learning?",

answer="Machine learning is a branch of artificial intelligence "

"that enables computers to learn from data "

"without being explicitly programmed."

).withinputs("question"),

dspy.Example(

question="Explain the difference between supervised "

"and unsupervised learning",

answer="Supervised learning uses labeled data for training, "

"while unsupervised learning discovers patterns "

"from unlabeled data."

).withinputs("question"),

dspy.Example(

question="What is the purpose of activation functions "

"in neural networks?",

answer="Activation functions determine a neuron's output "

"based on its input, adding non-linearity "

"to the model."

).withinputs("question"),

# Add more examples...

]

devset = [

dspy.Example(

question="What is transfer learning?",

answer="Transfer learning is a technique of using a model "

"trained on one task as a starting point "

"for a different but related task."

).withinputs("question"),

]

BootstrapFewShot

An optimizer that automatically selects the best few-shot examples.

from dspy.teleprompt import BootstrapFewShot

Define module

qamodule = dspy.ChainOfThought("question -> answer")

Define metric

def answerquality(example, prediction, trace=None):

# Simple: check if prediction is long enough and relevant

if len(prediction.answer.split()) < 10:

return False

return True

Compile with BootstrapFewShot

optimizer = BootstrapFewShot(

metric=answerquality,

maxbootstrappeddemos=4,

maxlabeleddemos=4,

)

compiledqa = optimizer.compile(

qamodule,

trainset=trainset,

)

Use the optimized module

result = compiledqa(question="What is deep learning?")

print(result.answer)

BootstrapFewShotWithRandomSearch

A variant that tries various combinations of examples and selects the best one.

from dspy.teleprompt import BootstrapFewShotWithRandomSearch

optimizer = BootstrapFewShotWithRandomSearch(

metric=answerquality,

maxbootstrappeddemos=4,

maxlabeleddemos=4,

numcandidateprograms=10,

numthreads=4,

)

compiledqa = optimizer.compile(

qamodule,

trainset=trainset,

valset=devset,

)

MIPROv2

An advanced optimizer that simultaneously optimizes instructions and few-shot examples.

from dspy.teleprompt import MIPROv2

optimizer = MIPROv2(

metric=answerquality,

numcandidates=7,

inittemperature=1.0,

)

compiledqa = optimizer.compile(

qamodule,

trainset=trainset,

maxbootstrappeddemos=4,

maxlabeleddemos=4,

)

Saving and Loading Compiled Programs

After optimization, you can save and reload compiled programs.

# Save the optimized program

compiledqa.save("optimizedqaprogram.json")

Load it back

loadedqa = dspy.ChainOfThought("question -> answer")

loadedqa.load("optimizedqaprogram.json")

Use as usual

result = loadedqa(question="What is a neural network?")

print(result.answer)

Practical Example: RAG Pipeline with DSPy

Here is a complete RAG (Retrieval-Augmented Generation) implementation using DSPy.

import dspy

from dspy.retrieve.chromadbrm import ChromadbRM

import chromadb

Setup ChromaDB

chromaclient = chromadb.PersistentClient(path="./chromadb")

Create collection and add documents

collection = chromaclient.getorcreatecollection("knowledgebase")

Add documents

documents = [

"Python is a programming language created by Guido van Rossum "

"in 1991. Python is known for its readable and easy-to-learn "

"syntax.",

"Machine learning has three main paradigms: supervised learning, "

"unsupervised learning, and reinforcement learning. Each has "

"different use cases.",

"Deep learning is a subset of machine learning that uses "

"neural networks with many layers. Popular frameworks include "

"TensorFlow, PyTorch, and JAX.",

"Natural Language Processing (NLP) is a branch of AI that focuses "

"on the interaction between computers and human language. NLP tasks "

"include sentiment analysis, named entity recognition, "

"and machine translation.",

"MLOps is a practice that combines Machine Learning and DevOps "

"to automate and standardize the ML model lifecycle "

"from development to production.",

]

collection.add(

documents=documents,

ids=[f"doc{i}" for i in range(len(documents))],

)

Setup retriever

retriever = ChromadbRM(

collectionname="knowledgebase",

persistdirectory="./chromadb",

k=3,

)

dspy.configure(

lm=dspy.LM("openai/gpt-4o-mini"),

rm=retriever,

)

Define RAG Signature

class GenerateAnswer(dspy.Signature):

"""Answer the question based on the provided context.

Provide an informative and accurate answer."""

context: list[str] = dspy.InputField(

desc="Relevant context from the knowledge base"

)

question: str = dspy.InputField(

desc="The question to answer"

)

answer: str = dspy.OutputField(

desc="A detailed and accurate answer based on the context"

)

Build RAG Module

class RAGModule(dspy.Module):

def init(self, numpassages=3):

super().init()

self.retrieve = dspy.Retrieve(k=numpassages)

self.generate = dspy.ChainOfThought(GenerateAnswer)

def forward(self, question):

# Retrieve relevant passages

context = self.retrieve(question).passages

# Generate answer

result = self.generate(

context=context,

question=question,

)

return dspy.Prediction(

context=context,

answer=result.answer,

reasoning=result.reasoning,

)

Create instance and use

rag = RAGModule()

result = rag(question="What are the main paradigms in machine learning?")

print(f"Answer: {result.answer}")

print(f"Reasoning: {result.reasoning}")

print(f"Context used: {result.context}")

Optimize RAG with training data

trainsetrag = [

dspy.Example(

question="Who created Python?",

answer="Python was created by Guido van Rossum in 1991."

).withinputs("question"),

dspy.Example(

question="What is deep learning?",

answer="Deep learning is a subset of machine learning "

"that uses neural networks with many layers."

).withinputs("question"),

dspy.Example(

question="Name popular deep learning frameworks",

answer="Popular deep learning frameworks include TensorFlow, "

"PyTorch, and JAX."

).withinputs("question"),

]

Compile RAG module

from dspy.teleprompt import BootstrapFewShot

def ragmetric(example, prediction, trace=None):

# Check if the answer contains information from reference

refwords = set(example.answer.lower().split())

predwords = set(prediction.answer.lower().split())

overlap = len(refwords & predwords) / len(refwords)

return overlap > 0.5

optimizer = BootstrapFewShot(

metric=ragmetric,

maxbootstrappeddemos=2,

)

compiledrag = optimizer.compile(rag, trainset=trainsetrag)

Use the optimized RAG

result = compiledrag(

question="What are the tasks in NLP?"

)

print(f"Optimized Answer: {result.answer}")

Evaluating Pipelines

DSPy provides utilities for systematically evaluating pipelines.

from dspy.evaluate import Evaluate

Create evaluator

evaluator = Evaluate(

devset=devset,

metric=ragmetric,

numthreads=4,

displayprogress=True,

displaytable=5,

)

Evaluate module before optimization

scorebefore = evaluator(rag)

print(f"Score before optimization: {scorebefore}")

Evaluate module after optimization

scoreafter = evaluator(compiledrag)

print(f"Score after optimization: {scoreafter}")

Tips and Best Practices

1. Start with clear Signatures

Good descriptions in signatures help both the LLM and optimizer understand the task better.

# Good

class GoodSignature(dspy.Signature):

"""Extract business entities from financial reports."""

report: str = dspy.InputField(desc="Financial report text")

entities: list[str] = dspy.OutputField(

desc="List of company names, financial figures, and key metrics"

)

2. Use high-quality training data

Training data quality significantly affects optimization results. Make sure your examples are representative and accurate.

3. Choose the right optimizer
  • BootstrapFewShot: Fast, great for getting started
  • BootstrapFewShotWithRandomSearch: Better results but slower
  • MIPROv2: Best for production, but slowest

4. Iterate and evaluate

Always evaluate before and after optimization. If results are unsatisfactory, try adding more training data or using a different optimizer.

5. Save compilation results

After finding the optimal configuration, always save the results so you do not need to repeat the optimization process.

Conclusion

DSPy changes the paradigm of working with LLMs. Instead of spending time writing and tweaking prompts manually, DSPy allows you to focus on defining tasks and letting the framework optimize the implementation.

Key takeaways from this tutorial:

  • Signatures for defining task input/output
  • Modules like ChainOfThought and ReAct for prompting strategies
  • Optimizers for automatically tuning prompts
  • Metrics for measuring output quality
  • A complete RAG pipeline implementation
  • Evaluation and saving/loading compiled programs

With DSPy, you can build LLM pipelines that are more robust, maintainable, and performant. Start with simple tasks, optimize incrementally, and scale as your project demands.

Related Articles

DSPy: Stop Hand-Tuning Prompts, Let the Compiler Optimize Them

DSPy: Berhenti Ngoprek Prompt Manual, Biarkan Compiler yang Optimasi Halo temen-temen, kali ini aku mau ngenalin satu li...

MLX Tutorial: Apple's Machine Learning Framework for Apple Silicon

Tutorial MLX: Framework Machine Learning Apple untuk Apple Silicon MLX adalah framework machine learning open-source dar...

Complete vLLM Tutorial: High-Performance LLM Serving

Tutorial Lengkap vLLM: High-Performance LLM Serving vLLM adalah library Python untuk inference dan serving LLM dengan pe...

Complete Ollama Tutorial: Deploy LLMs Locally

Tutorial Lengkap Ollama: Deploy LLM Secara Lokal Ollama adalah tool open-source yang memudahkan Anda menjalankan Large L...