Complete Ollama Tutorial: Deploy LLMs Locally
Ollama is an open-source tool that makes it easy to run Large Language Models (LLMs) locally on your computer. With Ollama, you can use models like Llama 3, Mistral, Gemma, and many more without requiring internet connection or paid APIs.
Why Ollama?
Benefits of using Ollama:- Privacy: Data never leaves your computer
- No API costs: Free after downloading models
- Offline capable: Works without internet
- Easy setup: One command to run models
- OpenAI-compatible API: Drop-in replacement for OpenAI
- Development and testing AI applications
- Private/sensitive data processing
- Offline AI applications
- Learning and experimenting with LLMs
- Cost-effective inference
Installation
1. Install on Linux
# Install with script
curl -fsSL https://ollama.com/install.sh | sh
Verify installation
ollama --version
2. Install on macOS
# Download from website or use Homebrew
brew install ollama
Or download .dmg from https://ollama.com/download
3. Install on Windows
Download installer from ollama.com/download and run it.
4. Install via Docker
# Pull image
docker pull ollama/ollama
Run container
docker run -d -v ollama:/root/.ollama -p 11434:11434 --name ollama ollama/ollama
With GPU (NVIDIA)
docker run -d --gpus all -v ollama:/root/.ollama -p 11434:11434 --name ollama ollama/ollama
Quick Start
1. Pull and Run Model
# Download and run Llama 3
ollama run llama3
Or other models
ollama run mistral
ollama run gemma:7b
ollama run phi3
ollama run codellama
Chat directly in terminal
>> Hello, who are you?
I am an AI assistant...
>> /bye # Exit chat
2. Available Models
| Model | Size | Use Case |
|-------|------|----------|
| llama3:8b | 4.7GB | General purpose, balanced |
| llama3:70b | 40GB | High quality responses |
| mistral | 4.1GB | Fast, efficient |
| gemma:7b | 5GB | Google's open model |
| phi3 | 2.2GB | Small, efficient |
| codellama | 3.8GB | Code generation |
| llava | 4.5GB | Vision + Language |
| mixtral | 26GB | Mixture of experts |
# List available models
ollama list
Pull specific version
ollama pull llama3:8b
ollama pull llama3:70b
Remove model
ollama rm llama3:8b
3. Model Commands
# Show model info
ollama show llama3
Copy model (for custom)
ollama cp llama3 my-llama3
Push to registry (if you have account)
ollama push username/my-model
REST API
Ollama provides an OpenAI-compatible REST API.
1. Generate Completion
# Simple generation
curl http://localhost:11434/api/generate -d '{
"model": "llama3",
"prompt": "Explain machine learning in 2 sentences"
}'
With streaming disabled
curl http://localhost:11434/api/generate -d '{
"model": "llama3",
"prompt": "Hello",
"stream": false
}'
2. Chat API
curl http://localhost:11434/api/chat -d '{
"model": "llama3",
"messages": [
{"role": "system", "content": "You are a helpful assistant"},
{"role": "user", "content": "What is Python?"}
],
"stream": false
}'
3. Embeddings
curl http://localhost:11434/api/embeddings -d '{
"model": "llama3",
"prompt": "Text to embed"
}'
Python Integration
1. Using requests
import requests
import json
def generate(prompt, model="llama3"):
response = requests.post(
"http://localhost:11434/api/generate",
json={
"model": model,
"prompt": prompt,
"stream": False
}
)
return response.json()["response"]
def chat(messages, model="llama3"):
response = requests.post(
"http://localhost:11434/api/chat",
json={
"model": model,
"messages": messages,
"stream": False
}
)
return response.json()["message"]["content"]
Usage
result = generate("Explain quantum computing")
print(result)
messages = [
{"role": "system", "content": "You are a math teacher"},
{"role": "user", "content": "Explain the Pythagorean theorem"}
]
result = chat(messages)
print(result)
2. Using ollama-python Library
pip install ollama
import ollama
Generate
response = ollama.generate(
model='llama3',
prompt='What is a neural network?'
)
print(response['response'])
Chat
response = ollama.chat(
model='llama3',
messages=[
{'role': 'system', 'content': 'You are a coding assistant'},
{'role': 'user', 'content': 'Write a Python function for factorial'}
]
)
print(response['message']['content'])
Streaming
stream = ollama.chat(
model='llama3',
messages=[{'role': 'user', 'content': 'Tell me about AI'}],
stream=True
)
for chunk in stream:
print(chunk['message']['content'], end='', flush=True)
Embeddings
embeddings = ollama.embeddings(
model='llama3',
prompt='Hello world'
)
print(len(embeddings['embedding'])) # Vector dimension
3. Async Support
import asyncio
import ollama
async def asyncchat():
response = await ollama.AsyncClient().chat(
model='llama3',
messages=[{'role': 'user', 'content': 'Hello!'}]
)
return response['message']['content']
Run
result = asyncio.run(asyncchat())
print(result)
OpenAI Compatibility
Ollama can be used as a drop-in replacement for OpenAI API.
1. With OpenAI Python Library
from openai import OpenAI
Point to Ollama server
client = OpenAI(
baseurl="http://localhost:11434/v1",
apikey="ollama" # Required but not used
)
Chat completion
response = client.chat.completions.create(
model="llama3",
messages=[
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": "What is Python?"}
]
)
print(response.choices[0].message.content)
Streaming
stream = client.chat.completions.create(
model="llama3",
messages=[{"role": "user", "content": "Tell me a story"}],
stream=True
)
for chunk in stream:
if chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="")
2. With LangChain
from langchaincommunity.llms import Ollama
from langchain
community.chatmodels import ChatOllama
from langchain.chains import LLMChain
from langchain.prompts import PromptTemplate
LLM
llm = Ollama(model="llama3")
response = llm.invoke("Explain machine learning")
print(response)
Chat Model
chat = ChatOllama(model="llama3", temperature=0.7)
from langchain.schema import HumanMessage, SystemMessage
messages = [
SystemMessage(content="You are a helpful assistant"),
HumanMessage(content="What is deep learning?")
]
response = chat.invoke(messages)
print(response.content)
Chain
template = """Question: {question}
Answer: Let's think step by step."""
prompt = PromptTemplate(template=template, input
variables=["question"])
chain = LLMChain(llm=llm, prompt=prompt)
result = chain.invoke({"question": "What is 25 * 4?"})
print(result["text"])
Custom Models with Modelfile
1. Basic Modelfile
# Modelfile
FROM llama3
Set parameters
PARAMETER temperature 0.7
PARAMETER topp 0.9
PARAMETER topk 40
System prompt
SYSTEM """
You are a helpful AI assistant.
Answer concisely and clearly.
"""
# Create model from Modelfile
ollama create my-assistant -f Modelfile
Run custom model
ollama run my-assistant
2. Advanced Modelfile
# Modelfile for coding assistant
FROM codellama
PARAMETER temperature 0.2
PARAMETER numctx 4096
PARAMETER repeatpenalty 1.1
SYSTEM """
You are an expert programmer. Follow these rules:
Write clean, well-documented code
Follow best practices and design patterns
Explain your code when asked
Use type hints in Python
Write tests when appropriate
"""
Template for formatting
TEMPLATE """{{ if .System }}<|system|>
{{ .System }}<|end|>
{{ end }}{{ if .Prompt }}<|user|>
{{ .Prompt }}<|end|>
{{ end }}<|assistant|>
{{ .Response }}<|end|>
"""
3. Fine-tuned Model Import
# Import GGUF model
FROM ./my-finetuned-model.gguf
PARAMETER temperature 0.8
SYSTEM "Custom fine-tuned model for specific task"
# Create from GGUF file
ollama create my-finetuned -f Modelfile
Multi-Modal Models (Vision)
1. LLaVA (Vision + Language)
# Pull LLaVA model
ollama pull llava
Run with image
ollama run llava "Describe this image: ./photo.jpg"
2. Python with Image
import ollama
import base64
def encodeimage(imagepath):
with open(imagepath, "rb") as f:
return base64.b64encode(f.read()).decode()
Analyze image
response = ollama.chat(
model='llava',
messages=[{
'role': 'user',
'content': 'What do you see in this image?',
'images': [encodeimage('photo.jpg')]
}]
)
print(response['message']['content'])
Performance Optimization
1. GPU Configuration
# Check GPU usage
nvidia-smi
Set specific GPU
CUDAVISIBLEDEVICES=0 ollama serve
Multiple GPUs
CUDAVISIBLEDEVICES=0,1 ollama serve
2. Memory Management
# Set context length (reduce to save memory)
ollama run llama3 --num-ctx 2048
In Modelfile
PARAMETER numctx 2048
PARAMETER numgpu 1 # Layers on GPU
3. Quantization Options
# Pull quantized versions
ollama pull llama3:8b-q40 # 4-bit quantization
ollama pull llama3:8b-q80 # 8-bit quantization
Smaller = faster, larger = better quality
Building Applications
1. Simple Chatbot
import ollama
class Chatbot:
def init(self, model="llama3", systemprompt=None):
self.model = model
self.messages = []
if systemprompt:
self.messages.append({
"role": "system",
"content": systemprompt
})
def chat(self, userinput):
self.messages.append({
"role": "user",
"content": userinput
})
response = ollama.chat(
model=self.model,
messages=self.messages
)
assistantmessage = response["message"]["content"]
self.messages.append({
"role": "assistant",
"content": assistantmessage
})
return assistantmessage
def clear(self):
self.messages = self.messages[:1] if self.messages else []
Usage
bot = Chatbot(systemprompt="You are a friendly assistant")
while True:
userinput = input("You: ")
if userinput.lower() in ['quit', 'exit']:
break
response = bot.chat(userinput)
print(f"Bot: {response}\n")
2. RAG with Ollama
import ollama
from langchaincommunity.embeddings import OllamaEmbeddings
from langchaincommunity.vectorstores import Chroma
from langchain.textsplitter import RecursiveCharacterTextSplitter
Setup embeddings
embeddings = OllamaEmbeddings(model="llama3")
Sample documents
documents = [
"Python is a popular programming language.",
"Machine learning is a branch of AI.",
"Deep learning uses neural networks."
]
Create vector store
textsplitter = RecursiveCharacterTextSplitter(chunksize=500)
texts = textsplitter.createdocuments(documents)
vectorstore = Chroma.fromdocuments(texts, embeddings)
RAG function
def ragquery(question):
# Retrieve relevant docs
docs = vectorstore.similaritysearch(question, k=2)
context = "\n".join([doc.pagecontent for doc in docs])
# Generate answer
prompt = f"""Based on the following context, answer the question.
Context:
{context}
Question: {question}
Answer:"""
response = ollama.generate(model="llama3", prompt=prompt)
return response["response"]
Usage
answer = ragquery("What is machine learning?")
print(answer)
3. FastAPI Server
from fastapi import FastAPI, HTTPException
from pydantic import BaseModel
import ollama
app = FastAPI()
class ChatRequest(BaseModel):
message: str
model: str = "llama3"
system: str = None
class ChatResponse(BaseModel):
response: str
@app.post("/chat", responsemodel=ChatResponse)
async def chat(request: ChatRequest):
try:
messages = []
if request.system:
messages.append({"role": "system", "content": request.system})
messages.append({"role": "user", "content": request.message})
response = ollama.chat(
model=request.model,
messages=messages
)
return ChatResponse(response=response["message"]["content"])
except Exception as e:
raise HTTPException(statuscode=500, detail=str(e))
@app.get("/models")
async def listmodels():
models = ollama.list()
return {"models": [m["name"] for m in models["models"]]}
Run: uvicorn main:app --reload
Troubleshooting
1. Common Issues
# Model not found
ollama pull llama3 # Re-download
Out of memory
Use smaller or quantized model
ollama run llama3:8b-q40
Server not running
ollama serve # Start server manually
Port conflict
OLLAMAHOST=0.0.0.0:11435 ollama serve
2. Logs and Debug
# View logs
journalctl -u ollama -f # Linux with systemd
Environment variables
OLLAMADEBUG=1 ollama serve # Enable debug mode
OLLAMA_HOST=0.0.0.0 ollama serve # Listen on all interfaces
Conclusion
Ollama simplifies local LLM deployment with:
Key takeaways:
- Use quantized models for limited hardware
- Custom Modelfile for specific behavior
- OpenAI compatibility for easy migration
- Ideal for development and privacy-sensitive apps