LiteLLM: Universal API Gateway for 100+ LLM Models

# LiteLLM: Universal API Gateway untuk 100+ Model LLM Dalam dunia AI yang berkembang pesat, kita dihadapkan dengan puluhan penyedia LLM (Large Language Model) seperti OpenAI, Anthropic, Google Gemini...

By Ruby Abdullah · · tutorial
LiteLLMLLMAPI GatewayOpenAIPython

LiteLLM: Universal API Gateway for 100+ LLM Models

In the rapidly evolving AI landscape, we face dozens of LLM (Large Language Model) providers such as OpenAI, Anthropic, Google Gemini, Cohere, Ollama, and many more. Each provider has different API formats, authentication methods, and parameters. Imagine if you could call all these models with a single, unified interface. That is exactly what LiteLLM offers.

LiteLLM is an open-source Python library that provides a unified interface to call 100+ LLM models using the OpenAI API format. With LiteLLM, you write your code once and can switch between providers without changing a single line of code.

Why LiteLLM?

Here are several reasons why LiteLLM has become a popular choice:

  • Unified API: One calling format for all providers
  • 100+ Models: Supports OpenAI, Anthropic, Google, Cohere, Ollama, HuggingFace, and more
  • OpenAI-Compatible Proxy: Run a proxy server compatible with the OpenAI API
  • Fallback & Retry: Automatically switch to another model if one fails
  • Cost Tracking: Track usage costs for each model
  • Load Balancing: Distribute requests across multiple models/providers
  • Streaming: Built-in streaming response support
  • Production Ready: Router for large-scale production deployments

Installation and Setup

Installing LiteLLM

Install LiteLLM using pip:

pip install litellm

For the proxy server feature, install with additional dependencies:

pip install 'litellm[proxy]'

Configuring API Keys

Before getting started, prepare your API keys from the providers you want to use. Store them as environment variables:

# OpenAI

export OPENAIAPIKEY="sk-your-openai-key"

Anthropic

export ANTHROPICAPIKEY="sk-ant-your-anthropic-key"

Google Gemini

export GEMINIAPIKEY="your-gemini-key"

Or use a .env file

You can also use a .env file and load it with python-dotenv:

from dotenv import loaddotenv

loaddotenv()

Basic Completion Calls

Calling OpenAI GPT

import litellm

Call OpenAI GPT-4o

response = litellm.completion(

model="gpt-4o",

messages=[

{"role": "system", "content": "You are a helpful assistant."},

{"role": "user", "content": "Explain what machine learning is in 2 sentences."}

]

)

print(response.choices[0].message.content)

Calling Anthropic Claude

# Call Anthropic Claude - just change the model name!

response = litellm.completion(

model="anthropic/claude-sonnet-4-20250514",

messages=[

{"role": "user", "content": "Explain what a neural network is in 2 sentences."}

]

)

print(response.choices[0].message.content)

Calling Google Gemini

# Call Google Gemini

response = litellm.completion(

model="gemini/gemini-2.0-flash",

messages=[

{"role": "user", "content": "What is the difference between supervised and unsupervised learning?"}

]

)

print(response.choices[0].message.content)

Calling Ollama (Local Models)

# Call a local model via Ollama

response = litellm.completion(

model="ollama/llama3",

messages=[

{"role": "user", "content": "Write a Python function to calculate fibonacci numbers."}

],

apibase="http://localhost:11434"

)

print(response.choices[0].message.content)

Unified Interface - Same Code, Different Providers

The key advantage of LiteLLM is that you can create a generic function that works with any provider:

import litellm

def askai(question: str, model: str = "gpt-4o") -> str:

"""Universal function to query any AI model."""

response = litellm.completion(

model=model,

messages=[

{"role": "user", "content": question}

],

temperature=0.7,

maxtokens=1000

)

return response.choices[0].message.content

Use with different providers - the same code!

question = "What is an API Gateway?"

OpenAI

openaianswer = askai(question, model="gpt-4o")

print(f"OpenAI: {openaianswer}\n")

Anthropic

claudeanswer = askai(question, model="anthropic/claude-sonnet-4-20250514")

print(f"Claude: {claudeanswer}\n")

Gemini

geminianswer = askai(question, model="gemini/gemini-2.0-flash")

print(f"Gemini: {geminianswer}\n")

Ollama (local)

ollamaanswer = askai(question, model="ollama/llama3")

print(f"Ollama: {ollamaanswer}\n")

Notice that the messages, temperature, and maxtokens format is completely uniform. LiteLLM automatically translates these parameters to the format understood by each provider.

LiteLLM Proxy Server

The LiteLLM Proxy Server allows you to run a server that is compatible with the OpenAI API. This means any application that already supports the OpenAI API can connect to 100+ models through this proxy.

Proxy Configuration

Create a litellmconfig.yaml file:

modellist:
  • modelname: gpt-4o
litellmparams:

model: gpt-4o

apikey: sk-your-openai-key

  • modelname: claude-sonnet
litellmparams:

model: anthropic/claude-sonnet-4-20250514

apikey: sk-ant-your-anthropic-key

  • modelname: gemini-flash
litellmparams:

model: gemini/gemini-2.0-flash

apikey: your-gemini-key

  • modelname: llama-local
litellmparams:

model: ollama/llama3

apibase: http://localhost:11434

litellmsettings:

dropparams: true

setverbose: false

Running the Proxy

litellm --config litellmconfig.yaml --port 4000

Using the Proxy from a Client

Now you can use the standard OpenAI SDK pointing to the proxy:

from openai import OpenAI

Point to LiteLLM Proxy

client = OpenAI(

apikey="sk-anything", # proxy doesn't validate keys by default

baseurl="http://localhost:4000"

)

Call Claude through an OpenAI-compatible endpoint!

response = client.chat.completions.create(

model="claude-sonnet",

messages=[

{"role": "user", "content": "Hello, who are you?"}

]

)

print(response.choices[0].message.content)

This is extremely useful for:

  • Teams that want to use a single endpoint for all models
  • Applications already integrated with the OpenAI API
  • Centralized access control and cost management

Streaming Responses

Streaming allows you to receive responses incrementally (token by token), providing a more responsive user experience:

import litellm

Streaming with OpenAI

response = litellm.completion(

model="gpt-4o",

messages=[

{"role": "user", "content": "Write a short story about a robot learning to cook."}

],

stream=True

)

Receive the response incrementally

for chunk in response:

content = chunk.choices[0].delta.content

if content:

print(content, end="", flush=True)

print() # Newline at the end

Streaming works with all providers:

# Streaming with Anthropic Claude

response = litellm.completion(

model="anthropic/claude-sonnet-4-20250514",

messages=[

{"role": "user", "content": "Explain the concept of recursion using an analogy."}

],

stream=True

)

for chunk in response:

content = chunk.choices[0].delta.content

if content:

print(content, end="", flush=True)

Fallbacks and Retries

In production, you need to handle cases when a provider experiences issues. LiteLLM provides an elegant fallback mechanism:

Basic Fallback

import litellm

from litellm import completion

Configure fallback

litellm.setverbose = False

def completionwithfallback(messages, modelchain=None):

"""Try models sequentially until one succeeds."""

if modelchain is None:

modelchain = [

"gpt-4o",

"anthropic/claude-sonnet-4-20250514",

"gemini/gemini-2.0-flash"

]

for model in modelchain:

try:

response = completion(

model=model,

messages=messages,

timeout=30

)

print(f"Successfully used model: {model}")

return response

except Exception as e:

print(f"Model {model} failed: {e}")

continue

raise Exception("All models in the fallback chain failed!")

Usage

response = completionwithfallback(

messages=[{"role": "user", "content": "Hello!"}]

)

print(response.choices[0].message.content)

Retry with Backoff

import litellm

Configure global retries

litellm.numretries = 3 # Retry 3 times

litellm.retryafter = 5 # Wait 5 seconds between retries

response = litellm.completion(

model="gpt-4o",

messages=[{"role": "user", "content": "Hello!"}],

numretries=3,

timeout=30

)

Load Balancing

Load balancing allows you to distribute requests across multiple deployments of the same model to increase throughput:

from litellm import Router

Configure router with multiple deployments

modellist = [

{

"modelname": "gpt-4o",

"litellmparams": {

"model": "gpt-4o",

"apikey": "sk-key-1",

}

},

{

"modelname": "gpt-4o",

"litellmparams": {

"model": "gpt-4o",

"apikey": "sk-key-2",

}

},

{

"modelname": "gpt-4o",

"litellmparams": {

"model": "azure/gpt-4o",

"apikey": "azure-key",

"apibase": "https://your-resource.openai.azure.com",

"apiversion": "2024-02-15-preview"

}

}

]

router = Router(

modellist=modellist,

routingstrategy="least-busy", # Options: simple-shuffle, least-busy, latency-based-routing

numretries=2

)

Router automatically distributes requests

response = router.completion(

model="gpt-4o",

messages=[{"role": "user", "content": "Explain load balancing."}]

)

print(response.choices[0].message.content)

Available routing strategies:

  • simple-shuffle: Random distribution
  • least-busy: Send to the deployment with the fewest queued requests
  • latency-based-routing: Send to the deployment with the lowest latency
  • usage-based-routing: Distribute based on usage (TPM/RPM)

Cost Tracking and Budgeting

LiteLLM provides built-in cost tracking features that are extremely useful:

Per-Request Cost Tracking

import litellm

from litellm import completioncost

response = litellm.completion(

model="gpt-4o",

messages=[

{"role": "user", "content": "Write a poem about Python programming."}

]

)

Calculate cost

cost = completioncost(completionresponse=response)

print(f"Cost of this request: ${cost:.6f}")

print(f"Input tokens: {response.usage.prompttokens}")

print(f"Output tokens: {response.usage.completiontokens}")

print(f"Total tokens: {response.usage.totaltokens}")

Cumulative Budget Tracking

import litellm

class BudgetTracker:

"""Cumulative cost tracker to control spending."""

def init(self, budgetlimit: float = 10.0):

self.totalcost = 0.0

self.budgetlimit = budgetlimit

self.requesthistory = []

def trackedcompletion(self, model: str, messages: list, kwargs):

if self.totalcost >= self.budgetlimit:

raise Exception(

f"Budget exhausted! Total: ${self.totalcost:.4f}, "

f"Limit: ${self.budgetlimit:.2f}"

)

response = litellm.completion(

model=model,

messages=messages,

kwargs

)

cost = litellm.completioncost(completionresponse=response)

self.totalcost += cost

self.requesthistory.append({

"model": model,

"cost": cost,

"tokens": response.usage.totaltokens

})

print(f"Request cost: ${cost:.6f} | Total: ${self.totalcost:.6f} | "

f"Remaining: ${self.budgetlimit - self.totalcost:.6f}")

return response

def getsummary(self):

print(f"\n{'='50}")

print(f"Budget Summary")

print(f"{'='50}")

print(f"Total requests: {len(self.requesthistory)}")

print(f"Total cost: ${self.totalcost:.6f}")

print(f"Budget limit: ${self.budgetlimit:.2f}")

print(f"Remaining: ${self.budgetlimit - self.totalcost:.6f}")

Usage

tracker = BudgetTracker(budgetlimit=5.0)

response = tracker.trackedcompletion(

model="gpt-4o",

messages=[{"role": "user", "content": "Hello!"}]

)

tracker.getsummary()

Embedding Calls

LiteLLM also supports embedding model calls with a uniform interface:

import litellm

OpenAI Embedding

response = litellm.embedding(

model="text-embedding-3-small",

input=["Machine learning is a branch of AI",

"Deep learning uses neural networks"]

)

embedding1 = response.data[0].embedding

embedding2 = response.data[1].embedding

print(f"Embedding dimensions: {len(embedding1)}")

Calculate cosine similarity

import numpy as np

def cosinesimilarity(a, b):

return np.dot(a, b) / (np.linalg.norm(a) np.linalg.norm(b))

similarity = cosinesimilarity(embedding1, embedding2)

print(f"Cosine similarity: {similarity:.4f}")

Embeddings with other providers:

# Cohere Embedding

response = litellm.embedding(

model="cohere/embed-english-v3.0",

input=["Hello world", "Machine learning is great"]

)

Azure OpenAI Embedding

response = litellm.embedding(

model="azure/text-embedding-3-small",

input=["Hello world"],

apikey="your-azure-key",

apibase="https://your-resource.openai.azure.com",

apiversion="2024-02-15-preview"

)

Router for Production Deployments

The LiteLLM Router is the core component for production deployments. It combines fallback, load balancing, and retry in a single configuration:

from litellm import Router

import asyncio

Production router configuration

modellist = [

# Primary: GPT-4o

{

"modelname": "primary-model",

"litellmparams": {

"model": "gpt-4o",

"apikey": "sk-openai-key",

"rpm": 100, # Rate limit: 100 requests per minute

"tpm": 100000 # Token limit: 100k tokens per minute

}

},

# Fallback 1: Claude

{

"modelname": "fallback-model",

"litellmparams": {

"model": "anthropic/claude-sonnet-4-20250514",

"apikey": "sk-ant-key",

"rpm": 50

}

},

# Fallback 2: Gemini

{

"modelname": "fallback-model-2",

"litellmparams": {

"model": "gemini/gemini-2.0-flash",

"apikey": "gemini-key",

"rpm": 60

}

}

]

router = Router(

modellist=modellist,

fallbacks=[

{"primary-model": ["fallback-model", "fallback-model-2"]}

],

routingstrategy="latency-based-routing",

numretries=3,

retryafter=5,

timeout=30,

allowedfails=2, # Mark deployment as failed after 2 errors

cooldowntime=60 # 60-second cooldown for failed deployments

)

Synchronous call

response = router.completion(

model="primary-model",

messages=[{"role": "user", "content": "Hello!"}]

)

Async call for high throughput

async def asynccompletion():

response = await router.acompletion(

model="primary-model",

messages=[{"role": "user", "content": "Hello from async!"}]

)

return response

Run async

result = asyncio.run(asynccompletion())

print(result.choices[0].message.content)

Practical Example: Building a Multi-Provider Chatbot with Automatic Fallback

Let us build a complete chatbot that leverages various LiteLLM features:

import litellm

from litellm import Router

from datetime import datetime

class MultiProviderChatbot:

"""Production chatbot with multi-provider support and automatic fallback."""

def init(self):

modellist = [

{

"modelname": "smart-model",

"litellmparams": {

"model": "gpt-4o",

"rpm": 100

}

},

{

"modelname": "smart-model",

"litellmparams": {

"model": "anthropic/claude-sonnet-4-20250514",

"rpm": 50

}

},

{

"modelname": "fast-model",

"litellmparams": {

"model": "gpt-4o-mini",

"rpm": 200

}

},

{

"modelname": "fast-model",

"litellmparams": {

"model": "gemini/gemini-2.0-flash",

"rpm": 100

}

}

]

self.router = Router(

modellist=modellist,

routingstrategy="latency-based-routing",

numretries=2,

timeout=30,

allowedfails=2,

cooldowntime=30

)

self.conversationhistory = []

self.totalcost = 0.0

self.totaltokens = 0

def selectmodel(self, message: str) -> str:

"""Select model based on message complexity."""

complexkeywords = [

"analyze", "explain in detail", "compare",

"write code", "debug", "architecture"

]

if any(kw in message.lower() for kw in complexkeywords):

return "smart-model"

return "fast-model"

def chat(self, usermessage: str, usestreaming: bool = False) -> str:

"""Send a message and get a response."""

self.conversationhistory.append({

"role": "user",

"content": usermessage

})

model = self.selectmodel(usermessage)

systemmessage = {

"role": "system",

"content": (

"You are a friendly and helpful AI assistant. "

"Respond in the same language as the user's question."

)

}

messages = [systemmessage] + self.conversationhistory[-10:]

try:

if usestreaming:

return self.streamresponse(model, messages)

else:

return self.normalresponse(model, messages)

except Exception as e:

errormsg = f"Sorry, an error occurred: {str(e)}"

return errormsg

def normalresponse(self, model: str, messages: list) -> str:

"""Normal (non-streaming) response."""

response = self.router.completion(

model=model,

messages=messages

)

assistantmessage = response.choices[0].message.content

self.conversationhistory.append({

"role": "assistant",

"content": assistantmessage

})

cost = litellm.completioncost(completionresponse=response)

self.totalcost += cost

self.totaltokens += response.usage.totaltokens

return assistantmessage

def streamresponse(self, model: str, messages: list) -> str:

"""Streaming response."""

response = self.router.completion(

model=model,

messages=messages,

stream=True

)

fullresponse = ""

for chunk in response:

content = chunk.choices[0].delta.content

if content:

print(content, end="", flush=True)

fullresponse += content

print()

self.conversationhistory.append({

"role": "assistant",

"content": fullresponse

})

return fullresponse

def getstats(self):

"""Display usage statistics."""

print(f"\n{'='50}")

print(f"Chatbot Statistics")

print(f"{'='50}")

print(f"Total messages: {len(self.conversationhistory)}")

print(f"Total tokens: {self.totaltokens:,}")

print(f"Total cost: ${self.totalcost:.6f}")

print(f"Average cost/message: "

f"${self.totalcost / max(len(self.conversationhistory), 1):.6f}")

def reset(self):

"""Reset conversation history."""

self.conversationhistory = []

print("Conversation history reset.")

Run the chatbot

def main():

chatbot = MultiProviderChatbot()

print("Multi-Provider Chatbot (type 'quit' to exit)")

print("Type 'stats' to view statistics")

print("Type 'stream' at the beginning of a message for streaming")

print(f"{'='50}\n")

while True:

userinput = input("You: ").strip()

if userinput.lower() == 'quit':

chatbot.getstats()

break

elif userinput.lower() == 'stats':

chatbot.getstats()

continue

elif userinput.lower() == 'reset':

chatbot.reset()

continue

usestreaming = userinput.lower().startswith('stream ')

if usestreaming:

userinput = userinput[7:]

print(f"\nAssistant: ", end="")

if not usestreaming:

response = chatbot.chat(userinput)

print(response)

else:

chatbot.chat(userinput, use_streaming=True)

print()

if name == "main":

main()

Tips and Best Practices

  • Use Environment Variables: Never hardcode API keys in your code. Always use environment variables or a secret manager.
  • Implement Fallbacks: Always prepare at least 2 providers as fallbacks to avoid downtime.
  • Monitor Costs: Use cost tracking to monitor spending and prevent unexpectedly high bills.
  • Choose the Right Model: Use lighter models (like GPT-4o-mini or Gemini Flash) for simple tasks, and more powerful models for complex tasks.
  • Rate Limiting: Configure RPM and TPM in the router to avoid rate limits from providers.
  • Async for High Throughput: Use acompletion() and router.acompletion() for applications that require high throughput.
  • Caching: Enable LiteLLM caching to reduce costs and latency for repeated requests.
  • import litellm
    
    

    litellm.cache = litellm.Cache(type="redis", host="localhost", port=6379)

    Identical requests will use the cache

    response = litellm.completion(

    model="gpt-4o",

    messages=[{"role": "user", "content": "What is Python?"}],

    caching=True

    )

    Conclusion

    LiteLLM is an extremely powerful tool for working with multiple LLM providers. With its unified interface, proxy server, fallback mechanisms, cost tracking, and load balancing, LiteLLM simplifies the complexity of managing multi-provider LLMs into just a few lines of code.

    Start with a simple installation, try a few providers, then implement production features like the Router and Proxy Server as your needs grow. With LiteLLM, you get complete flexibility to switch between providers without vendor lock-in.

    References and resources:

    • Official LiteLLM documentation: https://docs.litellm.ai/
    • GitHub repository: https://github.com/BerriAI/litellm
    • Supported model list: https://docs.litellm.ai/docs/providers

    Related Articles

    DSPy: Stop Hand-Tuning Prompts, Let the Compiler Optimize Them

    DSPy: Berhenti Ngoprek Prompt Manual, Biarkan Compiler yang Optimasi Halo temen-temen, kali ini aku mau ngenalin satu li...

    Inspect AI: The LLM Evaluation Framework from the UK AI Safety Institute

    Inspect AI: Framework Evaluasi LLM dari UK AI Safety Institute yang Wajib Kamu Coba Temen-temen, kalau kamu udah mulai s...

    Complete Braintrust Tutorial: Evaluate, Test, and Improve Your LLM Applications

    Tutorial Lengkap Braintrust: Evaluasi, Testing, dan Improve Aplikasi LLM Halo temen-temen, di tutorial kali ini aku mau ...

    Helicone: How to Monitor and Control Every LLM Call in Production

    Helicone: Cara Aku Memonitor dan Mengontrol Semua LLM Call di Produksi Temen-temen, kalau kalian sudah mulai serius bang...