Complete vLLM Tutorial: High-Performance LLM Serving

# Tutorial Lengkap vLLM: High-Performance LLM Serving vLLM adalah library Python untuk inference dan serving LLM dengan performa tinggi. Dengan teknologi PagedAttention yang inovatif, vLLM dapat menc...

By Ruby Abdullah · · tutorial
vLLMLLMInferenceGPUPythonMachine Learning

Complete vLLM Tutorial: High-Performance LLM Serving

vLLM is a Python library for high-performance LLM inference and serving. With its innovative PagedAttention technology, vLLM can achieve 10-24x higher throughput compared to standard implementations, making it the ideal choice for production deployments.

Why vLLM?

vLLM Advantages:
  • High throughput: PagedAttention for memory efficiency
  • Continuous batching: Maximize GPU utilization
  • OpenAI-compatible API: Easy integration
  • Tensor parallelism: Multi-GPU support
  • Quantization support: AWQ, GPTQ, FP8

Use Cases:
  • Production LLM serving
  • High-traffic API endpoints
  • Batch inference
  • Multi-tenant deployments

Installation

# Install vLLM

pip install vllm

With CUDA support

pip install vllm --extra-index-url https://download.pytorch.org/whl/cu118

Verify installation

python -c "import vllm; print(vllm.version)"

Quick Start

1. Offline Inference

from vllm import LLM, SamplingParams

Load model

llm = LLM(model="meta-llama/Llama-2-7b-hf")

Sampling parameters

samplingparams = SamplingParams(

temperature=0.8,

topp=0.95,

maxtokens=256

)

Generate

prompts = [

"Explain machine learning in simple terms:",

"Write a Python function to calculate factorial:",

"What is the capital of France?"

]

outputs = llm.generate(prompts, samplingparams)

for output in outputs:

prompt = output.prompt

generatedtext = output.outputs[0].text

print(f"Prompt: {prompt[:50]}...")

print(f"Response: {generatedtext}\n")

2. OpenAI-Compatible Server

# Start vLLM server

python -m vllm.entrypoints.openai.apiserver \

--model meta-llama/Llama-2-7b-hf \

--host 0.0.0.0 \

--port 8000

from openai import OpenAI

client = OpenAI(

baseurl="http://localhost:8000/v1",

apikey="token-abc123" # Required but not validated

)

Chat completion

response = client.chat.completions.create(

model="meta-llama/Llama-2-7b-hf",

messages=[

{"role": "system", "content": "You are a helpful assistant."},

{"role": "user", "content": "What is Python?"}

],

temperature=0.7,

maxtokens=256

)

print(response.choices[0].message.content)

Streaming

stream = client.chat.completions.create(

model="meta-llama/Llama-2-7b-hf",

messages=[{"role": "user", "content": "Tell me a story"}],

stream=True

)

for chunk in stream:

if chunk.choices[0].delta.content:

print(chunk.choices[0].delta.content, end="")

Sampling Parameters

from vllm import SamplingParams

Basic parameters

params = SamplingParams(

temperature=0.8, # Randomness (0.0 = deterministic)

topp=0.95, # Nucleus sampling

topk=50, # Top-k sampling

maxtokens=512, # Max output tokens

stop=["", "\n\n"], # Stop sequences

)

Advanced parameters

params = SamplingParams(

n=3, # Number of outputs per prompt

bestof=5, # Generate 5, return best 3

presencepenalty=0.5, # Penalize repeated tokens

frequencypenalty=0.5, # Penalize frequent tokens

repetitionpenalty=1.1, # Repetition penalty

lengthpenalty=1.0, # Length penalty for beam search

usebeamsearch=False, # Enable beam search

earlystopping=True,

skipspecialtokens=True,

ignoreeos=False,

)

Server Configuration

1. Basic Server Options

python -m vllm.entrypoints.openai.apiserver \

--model meta-llama/Llama-2-7b-hf \

--host 0.0.0.0 \

--port 8000 \

--dtype auto \

--max-model-len 4096 \

--gpu-memory-utilization 0.9

2. Multi-GPU (Tensor Parallelism)

# 2 GPUs

python -m vllm.entrypoints.openai.apiserver \

--model meta-llama/Llama-2-70b-hf \

--tensor-parallel-size 2

4 GPUs

python -m vllm.entrypoints.openai.apiserver \

--model meta-llama/Llama-2-70b-hf \

--tensor-parallel-size 4

3. Quantization

# AWQ quantization

python -m vllm.entrypoints.openai.apiserver \

--model TheBloke/Llama-2-7B-AWQ \

--quantization awq

GPTQ quantization

python -m vllm.entrypoints.openai.apiserver \

--model TheBloke/Llama-2-7B-GPTQ \

--quantization gptq

4. Full Configuration

python -m vllm.entrypoints.openai.apiserver \

--model meta-llama/Llama-2-7b-hf \

--host 0.0.0.0 \

--port 8000 \

--dtype float16 \

--max-model-len 4096 \

--gpu-memory-utilization 0.95 \

--tensor-parallel-size 1 \

--max-num-seqs 256 \

--max-num-batched-tokens 32768 \

--trust-remote-code \

--enforce-eager \

--disable-log-requests

Python API

1. LLM Class Options

from vllm import LLM

llm = LLM(

model="meta-llama/Llama-2-7b-hf",

tokenizer=None, # Use model's tokenizer

dtype="auto", # auto, float16, bfloat16, float32

trustremotecode=True,

tensorparallelsize=1,

gpumemoryutilization=0.9,

maxmodellen=4096,

quantization=None, # awq, gptq, squeezellm

enforceeager=False,

maxnumseqs=256,

maxnumbatchedtokens=32768,

seed=42,

)

2. Batch Processing

from vllm import LLM, SamplingParams

llm = LLM(model="meta-llama/Llama-2-7b-hf")

samplingparams = SamplingParams(temperature=0.7, maxtokens=256)

Large batch processing

prompts = [f"Question {i}: What is {i} + {i}?" for i in range(100)]

vLLM handles batching automatically

outputs = llm.generate(prompts, samplingparams)

for output in outputs:

print(f"{output.prompt[:30]}... -> {output.outputs[0].text[:50]}...")

3. Async Engine

import asyncio

from vllm import AsyncLLMEngine

from vllm.engine.argutils import AsyncEngineArgs

from vllm.samplingparams import SamplingParams

async def main():

# Create engine

engineargs = AsyncEngineArgs(

model="meta-llama/Llama-2-7b-hf",

tensorparallelsize=1,

)

engine = AsyncLLMEngine.fromengineargs(engineargs)

# Generate

samplingparams = SamplingParams(temperature=0.8, maxtokens=128)

async def generate(prompt, requestid):

results = []

async for output in engine.generate(prompt, samplingparams, requestid):

results.append(output)

return results[-1]

# Multiple concurrent requests

tasks = [

generate("What is AI?", "req-1"),

generate("Explain Python", "req-2"),

generate("What is ML?", "req-3"),

]

results = await asyncio.gather(tasks)

for result in results:

print(result.outputs[0].text[:100])

asyncio.run(main())

Chat Templates

1. Using Chat Templates

from vllm import LLM, SamplingParams

from transformers import AutoTokenizer

modelname = "meta-llama/Llama-2-7b-chat-hf"

llm = LLM(model=modelname)

tokenizer = AutoTokenizer.frompretrained(modelname)

Format messages

messages = [

{"role": "system", "content": "You are a helpful assistant."},

{"role": "user", "content": "What is machine learning?"}

]

Apply chat template

prompt = tokenizer.applychattemplate(

messages,

tokenize=False,

addgenerationprompt=True

)

Generate

samplingparams = SamplingParams(temperature=0.7, maxtokens=256)

outputs = llm.generate([prompt], samplingparams)

print(outputs[0].outputs[0].text)

2. Multi-turn Conversation

from vllm import LLM, SamplingParams

from transformers import AutoTokenizer

modelname = "meta-llama/Llama-2-7b-chat-hf"

llm = LLM(model=modelname)

tokenizer = AutoTokenizer.frompretrained(modelname)

samplingparams = SamplingParams(temperature=0.7, maxtokens=256)

class ChatSession:

def init(self):

self.messages = [

{"role": "system", "content": "You are a helpful assistant."}

]

def chat(self, usermessage):

self.messages.append({"role": "user", "content": usermessage})

prompt = tokenizer.applychattemplate(

self.messages,

tokenize=False,

addgenerationprompt=True

)

outputs = llm.generate([prompt], samplingparams)

response = outputs[0].outputs[0].text

self.messages.append({"role": "assistant", "content": response})

return response

Usage

session = ChatSession()

print(session.chat("Hello!"))

print(session.chat("What did I just say?"))

Deployment

1. Docker Deployment

# Dockerfile

FROM nvidia/cuda:12.1.0-runtime-ubuntu22.04

RUN apt-get update && apt-get install -y python3 python3-pip

RUN pip3 install vllm

EXPOSE 8000

CMD ["python3", "-m", "vllm.entrypoints.openai.apiserver", \

"--model", "meta-llama/Llama-2-7b-hf", \

"--host", "0.0.0.0", \

"--port", "8000"]

docker build -t vllm-server .

docker run --gpus all -p 8000:8000 vllm-server

2. Docker Compose

# docker-compose.yml

version: '3.8'

services:

vllm:

image: vllm/vllm-openai:latest

runtime: nvidia

ports:

  • "8000:8000"
volumes:

  • ~/.cache/huggingface:/root/.cache/huggingface
environment:

  • HUGGINGFACEHUBTOKEN=${HFTOKEN}
command: >

--model meta-llama/Llama-2-7b-hf

--host 0.0.0.0

--port 8000

deploy:

resources:

reservations:

devices:

  • driver: nvidia
count: 1

capabilities: [gpu]

3. Kubernetes Deployment

# vllm-deployment.yaml

apiVersion: apps/v1

kind: Deployment

metadata:

name: vllm-server

spec:

replicas: 1

selector:

matchLabels:

app: vllm

template:

metadata:

labels:

app: vllm

spec:

containers:

  • name: vllm
image: vllm/vllm-openai:latest

ports:

  • containerPort: 8000
resources:

limits:

nvidia.com/gpu: 1

args:

  • "--model"
  • "meta-llama/Llama-2-7b-hf"
  • "--host"
  • "0.0.0.0"
  • "--port"
  • "8000"
volumeMounts:

  • name: cache
mountPath: /root/.cache

volumes:

  • name: cache
persistentVolumeClaim:

claimName: model-cache-pvc


apiVersion: v1

kind: Service

metadata:

name: vllm-service

spec:

selector:

app: vllm

ports:

  • port: 80
targetPort: 8000

type: LoadBalancer

Performance Optimization

1. Memory Optimization

# Increase GPU memory utilization

--gpu-memory-utilization 0.95

Reduce max model length

--max-model-len 2048

Use quantization

--quantization awq

2. Throughput Optimization

# Increase batch size

--max-num-seqs 512

--max-num-batched-tokens 65536

Disable logging for production

--disable-log-requests

--disable-log-stats

3. Latency Optimization

# Enable eager mode (faster first token)

--enforce-eager

Use speculative decoding (if supported)

--speculative-model small-model

Monitoring

1. Built-in Metrics

# Enable metrics endpoint

python -m vllm.entrypoints.openai.apiserver \

--model meta-llama/Llama-2-7b-hf \

--served-model-name llama2

import requests

Get metrics (Prometheus format)

response = requests.get("http://localhost:8000/metrics")

print(response.text)

2. Custom Monitoring

import time

from openai import OpenAI

client = OpenAI(baseurl="http://localhost:8000/v1", apikey="none")

def benchmark(numrequests=100):

latencies = []

for i in range(numrequests):

start = time.time()

response = client.chat.completions.create(

model="meta-llama/Llama-2-7b-hf",

messages=[{"role": "user", "content": f"Count to {i}"}],

maxtokens=50

)

latencies.append(time.time() - start)

print(f"Average latency: {sum(latencies)/len(latencies):.3f}s")

print(f"P50 latency: {sorted(latencies)[len(latencies)//2]:.3f}s")

print(f"P99 latency: {sorted(latencies)[int(len(latencies)0.99)]:.3f}s")

print(f"Throughput: {numrequests/sum(latencies):.2f} req/s")

benchmark()

Conclusion

vLLM is the best choice for production LLM serving with:

  • PagedAttention: Maximum memory efficiency
  • Continuous Batching: High throughput
  • OpenAI Compatible: Easy integration
  • Multi-GPU: Tensor parallelism
  • Quantization: AWQ, GPTQ support
  • Key takeaways:

    • Use vLLM for production workloads
    • Leverage quantization for memory efficiency
    • Multi-GPU for large models
    • Monitor performance with metrics endpoint

    Related Articles

    Fireworks AI: A Super Fast Inference Platform for Open LLMs and Multimodal Models

    Fireworks AI: Platform Inference Super Cepat buat Open LLM dan Model Multimodal Halo temen-temen, di tutorial kali ini a...

    MLX Tutorial: Apple's Machine Learning Framework for Apple Silicon

    Tutorial MLX: Framework Machine Learning Apple untuk Apple Silicon MLX adalah framework machine learning open-source dar...

    DSPy: A Framework for Programmatic LLM Optimization

    DSPy: Framework untuk Optimasi LLM Secara Programatik Prompt engineering secara manual adalah proses yang melelahkan dan...

    Complete Ollama Tutorial: Deploy LLMs Locally

    Tutorial Lengkap Ollama: Deploy LLM Secara Lokal Ollama adalah tool open-source yang memudahkan Anda menjalankan Large L...