LiteLLM: Universal API Gateway for 100+ LLM Models
In the rapidly evolving AI landscape, we face dozens of LLM (Large Language Model) providers such as OpenAI, Anthropic, Google Gemini, Cohere, Ollama, and many more. Each provider has different API formats, authentication methods, and parameters. Imagine if you could call all these models with a single, unified interface. That is exactly what LiteLLM offers.
LiteLLM is an open-source Python library that provides a unified interface to call 100+ LLM models using the OpenAI API format. With LiteLLM, you write your code once and can switch between providers without changing a single line of code.
Why LiteLLM?
Here are several reasons why LiteLLM has become a popular choice:
- Unified API: One calling format for all providers
- 100+ Models: Supports OpenAI, Anthropic, Google, Cohere, Ollama, HuggingFace, and more
- OpenAI-Compatible Proxy: Run a proxy server compatible with the OpenAI API
- Fallback & Retry: Automatically switch to another model if one fails
- Cost Tracking: Track usage costs for each model
- Load Balancing: Distribute requests across multiple models/providers
- Streaming: Built-in streaming response support
- Production Ready: Router for large-scale production deployments
Installation and Setup
Installing LiteLLM
Install LiteLLM using pip:
pip install litellm
For the proxy server feature, install with additional dependencies:
pip install 'litellm[proxy]'
Configuring API Keys
Before getting started, prepare your API keys from the providers you want to use. Store them as environment variables:
# OpenAI
export OPENAIAPIKEY="sk-your-openai-key"
Anthropic
export ANTHROPICAPIKEY="sk-ant-your-anthropic-key"
Google Gemini
export GEMINIAPIKEY="your-gemini-key"
Or use a .env file
You can also use a .env file and load it with python-dotenv:
from dotenv import loaddotenv
load
dotenv()
Basic Completion Calls
Calling OpenAI GPT
import litellm
Call OpenAI GPT-4o
response = litellm.completion(
model="gpt-4o",
messages=[
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": "Explain what machine learning is in 2 sentences."}
]
)
print(response.choices[0].message.content)
Calling Anthropic Claude
# Call Anthropic Claude - just change the model name!
response = litellm.completion(
model="anthropic/claude-sonnet-4-20250514",
messages=[
{"role": "user", "content": "Explain what a neural network is in 2 sentences."}
]
)
print(response.choices[0].message.content)
Calling Google Gemini
# Call Google Gemini
response = litellm.completion(
model="gemini/gemini-2.0-flash",
messages=[
{"role": "user", "content": "What is the difference between supervised and unsupervised learning?"}
]
)
print(response.choices[0].message.content)
Calling Ollama (Local Models)
# Call a local model via Ollama
response = litellm.completion(
model="ollama/llama3",
messages=[
{"role": "user", "content": "Write a Python function to calculate fibonacci numbers."}
],
apibase="http://localhost:11434"
)
print(response.choices[0].message.content)
Unified Interface - Same Code, Different Providers
The key advantage of LiteLLM is that you can create a generic function that works with any provider:
import litellm
def askai(question: str, model: str = "gpt-4o") -> str:
"""Universal function to query any AI model."""
response = litellm.completion(
model=model,
messages=[
{"role": "user", "content": question}
],
temperature=0.7,
maxtokens=1000
)
return response.choices[0].message.content
Use with different providers - the same code!
question = "What is an API Gateway?"
OpenAI
openaianswer = askai(question, model="gpt-4o")
print(f"OpenAI: {openaianswer}\n")
Anthropic
claudeanswer = askai(question, model="anthropic/claude-sonnet-4-20250514")
print(f"Claude: {claudeanswer}\n")
Gemini
geminianswer = askai(question, model="gemini/gemini-2.0-flash")
print(f"Gemini: {geminianswer}\n")
Ollama (local)
ollamaanswer = askai(question, model="ollama/llama3")
print(f"Ollama: {ollamaanswer}\n")
Notice that the messages, temperature, and maxtokens format is completely uniform. LiteLLM automatically translates these parameters to the format understood by each provider.
LiteLLM Proxy Server
The LiteLLM Proxy Server allows you to run a server that is compatible with the OpenAI API. This means any application that already supports the OpenAI API can connect to 100+ models through this proxy.
Proxy Configuration
Create a litellmconfig.yaml file:
modellist:
- modelname: gpt-4o
litellmparams:
model: gpt-4o
apikey: sk-your-openai-key
- modelname: claude-sonnet
litellmparams:
model: anthropic/claude-sonnet-4-20250514
api
key: sk-ant-your-anthropic-key
- modelname: gemini-flash
litellmparams:
model: gemini/gemini-2.0-flash
apikey: your-gemini-key
- modelname: llama-local
litellmparams:
model: ollama/llama3
api
base: http://localhost:11434
litellmsettings:
dropparams: true
setverbose: false
Running the Proxy
litellm --config litellmconfig.yaml --port 4000
Using the Proxy from a Client
Now you can use the standard OpenAI SDK pointing to the proxy:
from openai import OpenAI
Point to LiteLLM Proxy
client = OpenAI(
apikey="sk-anything", # proxy doesn't validate keys by default
baseurl="http://localhost:4000"
)
Call Claude through an OpenAI-compatible endpoint!
response = client.chat.completions.create(
model="claude-sonnet",
messages=[
{"role": "user", "content": "Hello, who are you?"}
]
)
print(response.choices[0].message.content)
This is extremely useful for:
- Teams that want to use a single endpoint for all models
- Applications already integrated with the OpenAI API
- Centralized access control and cost management
Streaming Responses
Streaming allows you to receive responses incrementally (token by token), providing a more responsive user experience:
import litellm
Streaming with OpenAI
response = litellm.completion(
model="gpt-4o",
messages=[
{"role": "user", "content": "Write a short story about a robot learning to cook."}
],
stream=True
)
Receive the response incrementally
for chunk in response:
content = chunk.choices[0].delta.content
if content:
print(content, end="", flush=True)
print() # Newline at the end
Streaming works with all providers:
# Streaming with Anthropic Claude
response = litellm.completion(
model="anthropic/claude-sonnet-4-20250514",
messages=[
{"role": "user", "content": "Explain the concept of recursion using an analogy."}
],
stream=True
)
for chunk in response:
content = chunk.choices[0].delta.content
if content:
print(content, end="", flush=True)
Fallbacks and Retries
In production, you need to handle cases when a provider experiences issues. LiteLLM provides an elegant fallback mechanism:
Basic Fallback
import litellm
from litellm import completion
Configure fallback
litellm.setverbose = False
def completionwithfallback(messages, modelchain=None):
"""Try models sequentially until one succeeds."""
if modelchain is None:
modelchain = [
"gpt-4o",
"anthropic/claude-sonnet-4-20250514",
"gemini/gemini-2.0-flash"
]
for model in modelchain:
try:
response = completion(
model=model,
messages=messages,
timeout=30
)
print(f"Successfully used model: {model}")
return response
except Exception as e:
print(f"Model {model} failed: {e}")
continue
raise Exception("All models in the fallback chain failed!")
Usage
response = completionwithfallback(
messages=[{"role": "user", "content": "Hello!"}]
)
print(response.choices[0].message.content)
Retry with Backoff
import litellm
Configure global retries
litellm.numretries = 3 # Retry 3 times
litellm.retryafter = 5 # Wait 5 seconds between retries
response = litellm.completion(
model="gpt-4o",
messages=[{"role": "user", "content": "Hello!"}],
numretries=3,
timeout=30
)
Load Balancing
Load balancing allows you to distribute requests across multiple deployments of the same model to increase throughput:
from litellm import Router
Configure router with multiple deployments
modellist = [
{
"modelname": "gpt-4o",
"litellmparams": {
"model": "gpt-4o",
"apikey": "sk-key-1",
}
},
{
"modelname": "gpt-4o",
"litellmparams": {
"model": "gpt-4o",
"apikey": "sk-key-2",
}
},
{
"modelname": "gpt-4o",
"litellmparams": {
"model": "azure/gpt-4o",
"apikey": "azure-key",
"apibase": "https://your-resource.openai.azure.com",
"apiversion": "2024-02-15-preview"
}
}
]
router = Router(
modellist=modellist,
routingstrategy="least-busy", # Options: simple-shuffle, least-busy, latency-based-routing
numretries=2
)
Router automatically distributes requests
response = router.completion(
model="gpt-4o",
messages=[{"role": "user", "content": "Explain load balancing."}]
)
print(response.choices[0].message.content)
Available routing strategies:
simple-shuffle: Random distributionleast-busy: Send to the deployment with the fewest queued requestslatency-based-routing: Send to the deployment with the lowest latencyusage-based-routing: Distribute based on usage (TPM/RPM)
Cost Tracking and Budgeting
LiteLLM provides built-in cost tracking features that are extremely useful:
Per-Request Cost Tracking
import litellm
from litellm import completioncost
response = litellm.completion(
model="gpt-4o",
messages=[
{"role": "user", "content": "Write a poem about Python programming."}
]
)
Calculate cost
cost = completioncost(completionresponse=response)
print(f"Cost of this request: ${cost:.6f}")
print(f"Input tokens: {response.usage.prompttokens}")
print(f"Output tokens: {response.usage.completiontokens}")
print(f"Total tokens: {response.usage.totaltokens}")
Cumulative Budget Tracking
import litellm
class BudgetTracker:
"""Cumulative cost tracker to control spending."""
def init(self, budgetlimit: float = 10.0):
self.totalcost = 0.0
self.budgetlimit = budgetlimit
self.requesthistory = []
def trackedcompletion(self, model: str, messages: list, kwargs):
if self.totalcost >= self.budgetlimit:
raise Exception(
f"Budget exhausted! Total: ${self.totalcost:.4f}, "
f"Limit: ${self.budgetlimit:.2f}"
)
response = litellm.completion(
model=model,
messages=messages,
kwargs
)
cost = litellm.completioncost(completionresponse=response)
self.totalcost += cost
self.requesthistory.append({
"model": model,
"cost": cost,
"tokens": response.usage.totaltokens
})
print(f"Request cost: ${cost:.6f} | Total: ${self.totalcost:.6f} | "
f"Remaining: ${self.budgetlimit - self.totalcost:.6f}")
return response
def getsummary(self):
print(f"\n{'='50}")
print(f"Budget Summary")
print(f"{'='50}")
print(f"Total requests: {len(self.requesthistory)}")
print(f"Total cost: ${self.totalcost:.6f}")
print(f"Budget limit: ${self.budgetlimit:.2f}")
print(f"Remaining: ${self.budgetlimit - self.totalcost:.6f}")
Usage
tracker = BudgetTracker(budgetlimit=5.0)
response = tracker.trackedcompletion(
model="gpt-4o",
messages=[{"role": "user", "content": "Hello!"}]
)
tracker.getsummary()
Embedding Calls
LiteLLM also supports embedding model calls with a uniform interface:
import litellm
OpenAI Embedding
response = litellm.embedding(
model="text-embedding-3-small",
input=["Machine learning is a branch of AI",
"Deep learning uses neural networks"]
)
embedding1 = response.data[0].embedding
embedding2 = response.data[1].embedding
print(f"Embedding dimensions: {len(embedding1)}")
Calculate cosine similarity
import numpy as np
def cosinesimilarity(a, b):
return np.dot(a, b) / (np.linalg.norm(a) np.linalg.norm(b))
similarity = cosinesimilarity(embedding1, embedding2)
print(f"Cosine similarity: {similarity:.4f}")
Embeddings with other providers:
# Cohere Embedding
response = litellm.embedding(
model="cohere/embed-english-v3.0",
input=["Hello world", "Machine learning is great"]
)
Azure OpenAI Embedding
response = litellm.embedding(
model="azure/text-embedding-3-small",
input=["Hello world"],
apikey="your-azure-key",
apibase="https://your-resource.openai.azure.com",
apiversion="2024-02-15-preview"
)
Router for Production Deployments
The LiteLLM Router is the core component for production deployments. It combines fallback, load balancing, and retry in a single configuration:
from litellm import Router
import asyncio
Production router configuration
modellist = [
# Primary: GPT-4o
{
"modelname": "primary-model",
"litellmparams": {
"model": "gpt-4o",
"apikey": "sk-openai-key",
"rpm": 100, # Rate limit: 100 requests per minute
"tpm": 100000 # Token limit: 100k tokens per minute
}
},
# Fallback 1: Claude
{
"modelname": "fallback-model",
"litellmparams": {
"model": "anthropic/claude-sonnet-4-20250514",
"apikey": "sk-ant-key",
"rpm": 50
}
},
# Fallback 2: Gemini
{
"modelname": "fallback-model-2",
"litellmparams": {
"model": "gemini/gemini-2.0-flash",
"apikey": "gemini-key",
"rpm": 60
}
}
]
router = Router(
modellist=modellist,
fallbacks=[
{"primary-model": ["fallback-model", "fallback-model-2"]}
],
routingstrategy="latency-based-routing",
numretries=3,
retryafter=5,
timeout=30,
allowedfails=2, # Mark deployment as failed after 2 errors
cooldowntime=60 # 60-second cooldown for failed deployments
)
Synchronous call
response = router.completion(
model="primary-model",
messages=[{"role": "user", "content": "Hello!"}]
)
Async call for high throughput
async def asynccompletion():
response = await router.acompletion(
model="primary-model",
messages=[{"role": "user", "content": "Hello from async!"}]
)
return response
Run async
result = asyncio.run(asynccompletion())
print(result.choices[0].message.content)
Practical Example: Building a Multi-Provider Chatbot with Automatic Fallback
Let us build a complete chatbot that leverages various LiteLLM features:
import litellm
from litellm import Router
from datetime import datetime
class MultiProviderChatbot:
"""Production chatbot with multi-provider support and automatic fallback."""
def init(self):
modellist = [
{
"modelname": "smart-model",
"litellmparams": {
"model": "gpt-4o",
"rpm": 100
}
},
{
"modelname": "smart-model",
"litellmparams": {
"model": "anthropic/claude-sonnet-4-20250514",
"rpm": 50
}
},
{
"modelname": "fast-model",
"litellmparams": {
"model": "gpt-4o-mini",
"rpm": 200
}
},
{
"modelname": "fast-model",
"litellmparams": {
"model": "gemini/gemini-2.0-flash",
"rpm": 100
}
}
]
self.router = Router(
modellist=modellist,
routingstrategy="latency-based-routing",
numretries=2,
timeout=30,
allowedfails=2,
cooldowntime=30
)
self.conversationhistory = []
self.totalcost = 0.0
self.totaltokens = 0
def selectmodel(self, message: str) -> str:
"""Select model based on message complexity."""
complexkeywords = [
"analyze", "explain in detail", "compare",
"write code", "debug", "architecture"
]
if any(kw in message.lower() for kw in complexkeywords):
return "smart-model"
return "fast-model"
def chat(self, usermessage: str, usestreaming: bool = False) -> str:
"""Send a message and get a response."""
self.conversationhistory.append({
"role": "user",
"content": usermessage
})
model = self.selectmodel(usermessage)
systemmessage = {
"role": "system",
"content": (
"You are a friendly and helpful AI assistant. "
"Respond in the same language as the user's question."
)
}
messages = [systemmessage] + self.conversationhistory[-10:]
try:
if usestreaming:
return self.streamresponse(model, messages)
else:
return self.normalresponse(model, messages)
except Exception as e:
errormsg = f"Sorry, an error occurred: {str(e)}"
return errormsg
def normalresponse(self, model: str, messages: list) -> str:
"""Normal (non-streaming) response."""
response = self.router.completion(
model=model,
messages=messages
)
assistantmessage = response.choices[0].message.content
self.conversationhistory.append({
"role": "assistant",
"content": assistantmessage
})
cost = litellm.completioncost(completionresponse=response)
self.totalcost += cost
self.totaltokens += response.usage.totaltokens
return assistantmessage
def streamresponse(self, model: str, messages: list) -> str:
"""Streaming response."""
response = self.router.completion(
model=model,
messages=messages,
stream=True
)
fullresponse = ""
for chunk in response:
content = chunk.choices[0].delta.content
if content:
print(content, end="", flush=True)
fullresponse += content
print()
self.conversationhistory.append({
"role": "assistant",
"content": fullresponse
})
return fullresponse
def getstats(self):
"""Display usage statistics."""
print(f"\n{'='50}")
print(f"Chatbot Statistics")
print(f"{'='50}")
print(f"Total messages: {len(self.conversationhistory)}")
print(f"Total tokens: {self.totaltokens:,}")
print(f"Total cost: ${self.totalcost:.6f}")
print(f"Average cost/message: "
f"${self.totalcost / max(len(self.conversationhistory), 1):.6f}")
def reset(self):
"""Reset conversation history."""
self.conversationhistory = []
print("Conversation history reset.")
Run the chatbot
def main():
chatbot = MultiProviderChatbot()
print("Multi-Provider Chatbot (type 'quit' to exit)")
print("Type 'stats' to view statistics")
print("Type 'stream' at the beginning of a message for streaming")
print(f"{'='50}\n")
while True:
userinput = input("You: ").strip()
if userinput.lower() == 'quit':
chatbot.getstats()
break
elif userinput.lower() == 'stats':
chatbot.getstats()
continue
elif userinput.lower() == 'reset':
chatbot.reset()
continue
usestreaming = userinput.lower().startswith('stream ')
if usestreaming:
userinput = userinput[7:]
print(f"\nAssistant: ", end="")
if not usestreaming:
response = chatbot.chat(userinput)
print(response)
else:
chatbot.chat(userinput, use_streaming=True)
print()
if name == "main":
main()
Tips and Best Practices
acompletion() and router.acompletion() for applications that require high throughput.import litellm
litellm.cache = litellm.Cache(type="redis", host="localhost", port=6379)
Identical requests will use the cache
response = litellm.completion(
model="gpt-4o",
messages=[{"role": "user", "content": "What is Python?"}],
caching=True
)
Conclusion
LiteLLM is an extremely powerful tool for working with multiple LLM providers. With its unified interface, proxy server, fallback mechanisms, cost tracking, and load balancing, LiteLLM simplifies the complexity of managing multi-provider LLMs into just a few lines of code.
Start with a simple installation, try a few providers, then implement production features like the Router and Proxy Server as your needs grow. With LiteLLM, you get complete flexibility to switch between providers without vendor lock-in.
References and resources:
- Official LiteLLM documentation: https://docs.litellm.ai/
- GitHub repository: https://github.com/BerriAI/litellm
- Supported model list: https://docs.litellm.ai/docs/providers