Complete vLLM Tutorial: High-Performance LLM Serving
vLLM is a Python library for high-performance LLM inference and serving. With its innovative PagedAttention technology, vLLM can achieve 10-24x higher throughput compared to standard implementations, making it the ideal choice for production deployments.
Why vLLM?
vLLM Advantages:- High throughput: PagedAttention for memory efficiency
- Continuous batching: Maximize GPU utilization
- OpenAI-compatible API: Easy integration
- Tensor parallelism: Multi-GPU support
- Quantization support: AWQ, GPTQ, FP8
- Production LLM serving
- High-traffic API endpoints
- Batch inference
- Multi-tenant deployments
Installation
# Install vLLM
pip install vllm
With CUDA support
pip install vllm --extra-index-url https://download.pytorch.org/whl/cu118
Verify installation
python -c "import vllm; print(vllm.version)"
Quick Start
1. Offline Inference
from vllm import LLM, SamplingParams
Load model
llm = LLM(model="meta-llama/Llama-2-7b-hf")
Sampling parameters
samplingparams = SamplingParams(
temperature=0.8,
topp=0.95,
maxtokens=256
)
Generate
prompts = [
"Explain machine learning in simple terms:",
"Write a Python function to calculate factorial:",
"What is the capital of France?"
]
outputs = llm.generate(prompts, samplingparams)
for output in outputs:
prompt = output.prompt
generatedtext = output.outputs[0].text
print(f"Prompt: {prompt[:50]}...")
print(f"Response: {generatedtext}\n")
2. OpenAI-Compatible Server
# Start vLLM server
python -m vllm.entrypoints.openai.apiserver \
--model meta-llama/Llama-2-7b-hf \
--host 0.0.0.0 \
--port 8000
from openai import OpenAI
client = OpenAI(
baseurl="http://localhost:8000/v1",
apikey="token-abc123" # Required but not validated
)
Chat completion
response = client.chat.completions.create(
model="meta-llama/Llama-2-7b-hf",
messages=[
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": "What is Python?"}
],
temperature=0.7,
maxtokens=256
)
print(response.choices[0].message.content)
Streaming
stream = client.chat.completions.create(
model="meta-llama/Llama-2-7b-hf",
messages=[{"role": "user", "content": "Tell me a story"}],
stream=True
)
for chunk in stream:
if chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="")
Sampling Parameters
from vllm import SamplingParams
Basic parameters
params = SamplingParams(
temperature=0.8, # Randomness (0.0 = deterministic)
topp=0.95, # Nucleus sampling
topk=50, # Top-k sampling
maxtokens=512, # Max output tokens
stop=["", "\n\n"], # Stop sequences
)
Advanced parameters
params = SamplingParams(
n=3, # Number of outputs per prompt
bestof=5, # Generate 5, return best 3
presencepenalty=0.5, # Penalize repeated tokens
frequencypenalty=0.5, # Penalize frequent tokens
repetitionpenalty=1.1, # Repetition penalty
lengthpenalty=1.0, # Length penalty for beam search
usebeamsearch=False, # Enable beam search
earlystopping=True,
skipspecialtokens=True,
ignoreeos=False,
)
Server Configuration
1. Basic Server Options
python -m vllm.entrypoints.openai.apiserver \
--model meta-llama/Llama-2-7b-hf \
--host 0.0.0.0 \
--port 8000 \
--dtype auto \
--max-model-len 4096 \
--gpu-memory-utilization 0.9
2. Multi-GPU (Tensor Parallelism)
# 2 GPUs
python -m vllm.entrypoints.openai.apiserver \
--model meta-llama/Llama-2-70b-hf \
--tensor-parallel-size 2
4 GPUs
python -m vllm.entrypoints.openai.apiserver \
--model meta-llama/Llama-2-70b-hf \
--tensor-parallel-size 4
3. Quantization
# AWQ quantization
python -m vllm.entrypoints.openai.apiserver \
--model TheBloke/Llama-2-7B-AWQ \
--quantization awq
GPTQ quantization
python -m vllm.entrypoints.openai.apiserver \
--model TheBloke/Llama-2-7B-GPTQ \
--quantization gptq
4. Full Configuration
python -m vllm.entrypoints.openai.apiserver \
--model meta-llama/Llama-2-7b-hf \
--host 0.0.0.0 \
--port 8000 \
--dtype float16 \
--max-model-len 4096 \
--gpu-memory-utilization 0.95 \
--tensor-parallel-size 1 \
--max-num-seqs 256 \
--max-num-batched-tokens 32768 \
--trust-remote-code \
--enforce-eager \
--disable-log-requests
Python API
1. LLM Class Options
from vllm import LLM
llm = LLM(
model="meta-llama/Llama-2-7b-hf",
tokenizer=None, # Use model's tokenizer
dtype="auto", # auto, float16, bfloat16, float32
trustremotecode=True,
tensorparallelsize=1,
gpumemoryutilization=0.9,
maxmodellen=4096,
quantization=None, # awq, gptq, squeezellm
enforceeager=False,
maxnumseqs=256,
maxnumbatchedtokens=32768,
seed=42,
)
2. Batch Processing
from vllm import LLM, SamplingParams
llm = LLM(model="meta-llama/Llama-2-7b-hf")
samplingparams = SamplingParams(temperature=0.7, maxtokens=256)
Large batch processing
prompts = [f"Question {i}: What is {i} + {i}?" for i in range(100)]
vLLM handles batching automatically
outputs = llm.generate(prompts, samplingparams)
for output in outputs:
print(f"{output.prompt[:30]}... -> {output.outputs[0].text[:50]}...")
3. Async Engine
import asyncio
from vllm import AsyncLLMEngine
from vllm.engine.argutils import AsyncEngineArgs
from vllm.samplingparams import SamplingParams
async def main():
# Create engine
engineargs = AsyncEngineArgs(
model="meta-llama/Llama-2-7b-hf",
tensorparallelsize=1,
)
engine = AsyncLLMEngine.fromengineargs(engineargs)
# Generate
samplingparams = SamplingParams(temperature=0.8, maxtokens=128)
async def generate(prompt, requestid):
results = []
async for output in engine.generate(prompt, samplingparams, requestid):
results.append(output)
return results[-1]
# Multiple concurrent requests
tasks = [
generate("What is AI?", "req-1"),
generate("Explain Python", "req-2"),
generate("What is ML?", "req-3"),
]
results = await asyncio.gather(tasks)
for result in results:
print(result.outputs[0].text[:100])
asyncio.run(main())
Chat Templates
1. Using Chat Templates
from vllm import LLM, SamplingParams
from transformers import AutoTokenizer
modelname = "meta-llama/Llama-2-7b-chat-hf"
llm = LLM(model=modelname)
tokenizer = AutoTokenizer.frompretrained(modelname)
Format messages
messages = [
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": "What is machine learning?"}
]
Apply chat template
prompt = tokenizer.applychattemplate(
messages,
tokenize=False,
addgenerationprompt=True
)
Generate
samplingparams = SamplingParams(temperature=0.7, maxtokens=256)
outputs = llm.generate([prompt], samplingparams)
print(outputs[0].outputs[0].text)
2. Multi-turn Conversation
from vllm import LLM, SamplingParams
from transformers import AutoTokenizer
modelname = "meta-llama/Llama-2-7b-chat-hf"
llm = LLM(model=modelname)
tokenizer = AutoTokenizer.frompretrained(modelname)
samplingparams = SamplingParams(temperature=0.7, maxtokens=256)
class ChatSession:
def init(self):
self.messages = [
{"role": "system", "content": "You are a helpful assistant."}
]
def chat(self, usermessage):
self.messages.append({"role": "user", "content": usermessage})
prompt = tokenizer.applychattemplate(
self.messages,
tokenize=False,
addgenerationprompt=True
)
outputs = llm.generate([prompt], samplingparams)
response = outputs[0].outputs[0].text
self.messages.append({"role": "assistant", "content": response})
return response
Usage
session = ChatSession()
print(session.chat("Hello!"))
print(session.chat("What did I just say?"))
Deployment
1. Docker Deployment
# Dockerfile
FROM nvidia/cuda:12.1.0-runtime-ubuntu22.04
RUN apt-get update && apt-get install -y python3 python3-pip
RUN pip3 install vllm
EXPOSE 8000
CMD ["python3", "-m", "vllm.entrypoints.openai.apiserver", \
"--model", "meta-llama/Llama-2-7b-hf", \
"--host", "0.0.0.0", \
"--port", "8000"]
docker build -t vllm-server .
docker run --gpus all -p 8000:8000 vllm-server
2. Docker Compose
# docker-compose.yml
version: '3.8'
services:
vllm:
image: vllm/vllm-openai:latest
runtime: nvidia
ports:
- "8000:8000"
volumes:
- ~/.cache/huggingface:/root/.cache/huggingface
environment:
- HUGGINGFACEHUBTOKEN=${HFTOKEN}
command: >
--model meta-llama/Llama-2-7b-hf
--host 0.0.0.0
--port 8000
deploy:
resources:
reservations:
devices:
- driver: nvidia
count: 1
capabilities: [gpu]
3. Kubernetes Deployment
# vllm-deployment.yaml
apiVersion: apps/v1
kind: Deployment
metadata:
name: vllm-server
spec:
replicas: 1
selector:
matchLabels:
app: vllm
template:
metadata:
labels:
app: vllm
spec:
containers:
- name: vllm
image: vllm/vllm-openai:latest
ports:
- containerPort: 8000
resources:
limits:
nvidia.com/gpu: 1
args:
- "--model"
- "meta-llama/Llama-2-7b-hf"
- "--host"
- "0.0.0.0"
- "--port"
- "8000"
volumeMounts:
- name: cache
mountPath: /root/.cache
volumes:
- name: cache
persistentVolumeClaim:
claimName: model-cache-pvc
apiVersion: v1
kind: Service
metadata:
name: vllm-service
spec:
selector:
app: vllm
ports:
- port: 80
targetPort: 8000
type: LoadBalancer
Performance Optimization
1. Memory Optimization
# Increase GPU memory utilization
--gpu-memory-utilization 0.95
Reduce max model length
--max-model-len 2048
Use quantization
--quantization awq
2. Throughput Optimization
# Increase batch size
--max-num-seqs 512
--max-num-batched-tokens 65536
Disable logging for production
--disable-log-requests
--disable-log-stats
3. Latency Optimization
# Enable eager mode (faster first token)
--enforce-eager
Use speculative decoding (if supported)
--speculative-model small-model
Monitoring
1. Built-in Metrics
# Enable metrics endpoint
python -m vllm.entrypoints.openai.apiserver \
--model meta-llama/Llama-2-7b-hf \
--served-model-name llama2
import requests
Get metrics (Prometheus format)
response = requests.get("http://localhost:8000/metrics")
print(response.text)
2. Custom Monitoring
import time
from openai import OpenAI
client = OpenAI(baseurl="http://localhost:8000/v1", apikey="none")
def benchmark(numrequests=100):
latencies = []
for i in range(numrequests):
start = time.time()
response = client.chat.completions.create(
model="meta-llama/Llama-2-7b-hf",
messages=[{"role": "user", "content": f"Count to {i}"}],
maxtokens=50
)
latencies.append(time.time() - start)
print(f"Average latency: {sum(latencies)/len(latencies):.3f}s")
print(f"P50 latency: {sorted(latencies)[len(latencies)//2]:.3f}s")
print(f"P99 latency: {sorted(latencies)[int(len(latencies)0.99)]:.3f}s")
print(f"Throughput: {numrequests/sum(latencies):.2f} req/s")
benchmark()
Conclusion
vLLM is the best choice for production LLM serving with:
Key takeaways:
- Use vLLM for production workloads
- Leverage quantization for memory efficiency
- Multi-GPU for large models
- Monitor performance with metrics endpoint