Tutorial Lengkap Ollama: Deploy LLM Secara Lokal
Ollama adalah tool open-source yang memudahkan Anda menjalankan Large Language Models (LLM) secara lokal di komputer Anda. Dengan Ollama, Anda dapat menggunakan model seperti Llama 3, Mistral, Gemma, dan banyak lagi tanpa memerlukan koneksi internet atau API berbayar.
Mengapa Ollama?
Keuntungan menggunakan Ollama:- Privacy: Data tidak keluar dari komputer Anda
- No API costs: Gratis setelah download model
- Offline capable: Bekerja tanpa internet
- Easy setup: Satu command untuk menjalankan model
- OpenAI-compatible API: Drop-in replacement untuk OpenAI
- Development dan testing aplikasi AI
- Private/sensitive data processing
- Offline AI applications
- Learning dan eksperimen dengan LLM
- Cost-effective inference
Instalasi
1. Install di Linux
# Install dengan script
curl -fsSL https://ollama.com/install.sh | sh
Verify installation
ollama --version
2. Install di macOS
# Download dari website atau gunakan Homebrew
brew install ollama
Atau download .dmg dari https://ollama.com/download
3. Install di Windows
Download installer dari ollama.com/download dan jalankan.
4. Install via Docker
# Pull image
docker pull ollama/ollama
Run container
docker run -d -v ollama:/root/.ollama -p 11434:11434 --name ollama ollama/ollama
Dengan GPU (NVIDIA)
docker run -d --gpus all -v ollama:/root/.ollama -p 11434:11434 --name ollama ollama/ollama
Quick Start
1. Pull dan Run Model
# Download dan run Llama 3
ollama run llama3
Atau model lain
ollama run mistral
ollama run gemma:7b
ollama run phi3
ollama run codellama
Chat langsung di terminal
>> Halo, siapa kamu?
Saya adalah AI assistant...
>> /bye # Keluar dari chat
2. Model yang Tersedia
| Model | Size | Use Case |
|-------|------|----------|
| llama3:8b | 4.7GB | General purpose, balanced |
| llama3:70b | 40GB | High quality responses |
| mistral | 4.1GB | Fast, efficient |
| gemma:7b | 5GB | Google's open model |
| phi3 | 2.2GB | Small, efficient |
| codellama | 3.8GB | Code generation |
| llava | 4.5GB | Vision + Language |
| mixtral | 26GB | Mixture of experts |
# List available models
ollama list
Pull specific version
ollama pull llama3:8b
ollama pull llama3:70b
Remove model
ollama rm llama3:8b
3. Model Commands
# Show model info
ollama show llama3
Copy model (untuk custom)
ollama cp llama3 my-llama3
Push to registry (jika punya akun)
ollama push username/my-model
REST API
Ollama menyediakan REST API yang OpenAI-compatible.
1. Generate Completion
# Simple generation
curl http://localhost:11434/api/generate -d '{
"model": "llama3",
"prompt": "Jelaskan apa itu machine learning dalam 2 kalimat"
}'
Dengan streaming disabled
curl http://localhost:11434/api/generate -d '{
"model": "llama3",
"prompt": "Hello",
"stream": false
}'
2. Chat API
curl http://localhost:11434/api/chat -d '{
"model": "llama3",
"messages": [
{"role": "system", "content": "Kamu adalah asisten yang helpful"},
{"role": "user", "content": "Apa itu Python?"}
],
"stream": false
}'
3. Embeddings
curl http://localhost:11434/api/embeddings -d '{
"model": "llama3",
"prompt": "Teks untuk di-embed"
}'
Python Integration
1. Menggunakan requests
import requests
import json
def generate(prompt, model="llama3"):
response = requests.post(
"http://localhost:11434/api/generate",
json={
"model": model,
"prompt": prompt,
"stream": False
}
)
return response.json()["response"]
def chat(messages, model="llama3"):
response = requests.post(
"http://localhost:11434/api/chat",
json={
"model": model,
"messages": messages,
"stream": False
}
)
return response.json()["message"]["content"]
Usage
result = generate("Jelaskan quantum computing")
print(result)
messages = [
{"role": "system", "content": "Kamu adalah guru matematika"},
{"role": "user", "content": "Jelaskan teorema Pythagoras"}
]
result = chat(messages)
print(result)
2. Menggunakan ollama-python Library
pip install ollama
import ollama
Generate
response = ollama.generate(
model='llama3',
prompt='Apa itu neural network?'
)
print(response['response'])
Chat
response = ollama.chat(
model='llama3',
messages=[
{'role': 'system', 'content': 'Kamu adalah asisten coding'},
{'role': 'user', 'content': 'Tulis fungsi Python untuk factorial'}
]
)
print(response['message']['content'])
Streaming
stream = ollama.chat(
model='llama3',
messages=[{'role': 'user', 'content': 'Ceritakan tentang AI'}],
stream=True
)
for chunk in stream:
print(chunk['message']['content'], end='', flush=True)
Embeddings
embeddings = ollama.embeddings(
model='llama3',
prompt='Hello world'
)
print(len(embeddings['embedding'])) # Vector dimension
3. Async Support
import asyncio
import ollama
async def asyncchat():
response = await ollama.AsyncClient().chat(
model='llama3',
messages=[{'role': 'user', 'content': 'Hello!'}]
)
return response['message']['content']
Run
result = asyncio.run(asyncchat())
print(result)
OpenAI Compatibility
Ollama dapat digunakan sebagai drop-in replacement untuk OpenAI API.
1. Dengan OpenAI Python Library
from openai import OpenAI
Point ke Ollama server
client = OpenAI(
baseurl="http://localhost:11434/v1",
apikey="ollama" # Required tapi tidak digunakan
)
Chat completion
response = client.chat.completions.create(
model="llama3",
messages=[
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": "What is Python?"}
]
)
print(response.choices[0].message.content)
Streaming
stream = client.chat.completions.create(
model="llama3",
messages=[{"role": "user", "content": "Tell me a story"}],
stream=True
)
for chunk in stream:
if chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="")
2. Dengan LangChain
from langchaincommunity.llms import Ollama
from langchain
community.chatmodels import ChatOllama
from langchain.chains import LLMChain
from langchain.prompts import PromptTemplate
LLM
llm = Ollama(model="llama3")
response = llm.invoke("Explain machine learning")
print(response)
Chat Model
chat = ChatOllama(model="llama3", temperature=0.7)
from langchain.schema import HumanMessage, SystemMessage
messages = [
SystemMessage(content="You are a helpful assistant"),
HumanMessage(content="What is deep learning?")
]
response = chat.invoke(messages)
print(response.content)
Chain
template = """Question: {question}
Answer: Let's think step by step."""
prompt = PromptTemplate(template=template, input
variables=["question"])
chain = LLMChain(llm=llm, prompt=prompt)
result = chain.invoke({"question": "What is 25 * 4?"})
print(result["text"])
Custom Models dengan Modelfile
1. Basic Modelfile
# Modelfile
FROM llama3
Set parameters
PARAMETER temperature 0.7
PARAMETER topp 0.9
PARAMETER topk 40
System prompt
SYSTEM """
Kamu adalah asisten AI yang helpful dan berbicara dalam Bahasa Indonesia.
Jawab dengan singkat dan jelas.
"""
# Create model dari Modelfile
ollama create my-assistant -f Modelfile
Run custom model
ollama run my-assistant
2. Advanced Modelfile
# Modelfile untuk coding assistant
FROM codellama
PARAMETER temperature 0.2
PARAMETER numctx 4096
PARAMETER repeatpenalty 1.1
SYSTEM """
You are an expert programmer. Follow these rules:
Write clean, well-documented code
Follow best practices and design patterns
Explain your code when asked
Use type hints in Python
Write tests when appropriate
"""
Template untuk formatting
TEMPLATE """{{ if .System }}<|system|>
{{ .System }}<|end|>
{{ end }}{{ if .Prompt }}<|user|>
{{ .Prompt }}<|end|>
{{ end }}<|assistant|>
{{ .Response }}<|end|>
"""
3. Fine-tuned Model Import
# Import GGUF model
FROM ./my-finetuned-model.gguf
PARAMETER temperature 0.8
SYSTEM "Custom fine-tuned model for specific task"
# Create dari GGUF file
ollama create my-finetuned -f Modelfile
Multi-Modal Models (Vision)
1. LLaVA (Vision + Language)
# Pull LLaVA model
ollama pull llava
Run dengan image
ollama run llava "Describe this image: ./photo.jpg"
2. Python dengan Image
import ollama
import base64
def encodeimage(imagepath):
with open(imagepath, "rb") as f:
return base64.b64encode(f.read()).decode()
Analyze image
response = ollama.chat(
model='llava',
messages=[{
'role': 'user',
'content': 'What do you see in this image?',
'images': [encodeimage('photo.jpg')]
}]
)
print(response['message']['content'])
Performance Optimization
1. GPU Configuration
# Check GPU usage
nvidia-smi
Set specific GPU
CUDAVISIBLEDEVICES=0 ollama serve
Multiple GPUs
CUDAVISIBLEDEVICES=0,1 ollama serve
2. Memory Management
# Set context length (reduce untuk hemat memory)
ollama run llama3 --num-ctx 2048
Dalam Modelfile
PARAMETER numctx 2048
PARAMETER numgpu 1 # Layers di GPU
3. Quantization Options
# Pull quantized versions
ollama pull llama3:8b-q40 # 4-bit quantization
ollama pull llama3:8b-q80 # 8-bit quantization
Smaller = faster, larger = better quality
Building Applications
1. Simple Chatbot
import ollama
class Chatbot:
def init(self, model="llama3", systemprompt=None):
self.model = model
self.messages = []
if systemprompt:
self.messages.append({
"role": "system",
"content": systemprompt
})
def chat(self, userinput):
self.messages.append({
"role": "user",
"content": userinput
})
response = ollama.chat(
model=self.model,
messages=self.messages
)
assistantmessage = response["message"]["content"]
self.messages.append({
"role": "assistant",
"content": assistantmessage
})
return assistantmessage
def clear(self):
self.messages = self.messages[:1] if self.messages else []
Usage
bot = Chatbot(systemprompt="Kamu adalah asisten yang ramah")
while True:
userinput = input("You: ")
if userinput.lower() in ['quit', 'exit']:
break
response = bot.chat(userinput)
print(f"Bot: {response}\n")
2. RAG dengan Ollama
import ollama
from langchaincommunity.embeddings import OllamaEmbeddings
from langchaincommunity.vectorstores import Chroma
from langchain.textsplitter import RecursiveCharacterTextSplitter
Setup embeddings
embeddings = OllamaEmbeddings(model="llama3")
Sample documents
documents = [
"Python adalah bahasa pemrograman yang populer.",
"Machine learning adalah cabang dari AI.",
"Deep learning menggunakan neural networks."
]
Create vector store
textsplitter = RecursiveCharacterTextSplitter(chunksize=500)
texts = textsplitter.createdocuments(documents)
vectorstore = Chroma.fromdocuments(texts, embeddings)
RAG function
def ragquery(question):
# Retrieve relevant docs
docs = vectorstore.similaritysearch(question, k=2)
context = "\n".join([doc.pagecontent for doc in docs])
# Generate answer
prompt = f"""Based on the following context, answer the question.
Context:
{context}
Question: {question}
Answer:"""
response = ollama.generate(model="llama3", prompt=prompt)
return response["response"]
Usage
answer = ragquery("Apa itu machine learning?")
print(answer)
3. FastAPI Server
from fastapi import FastAPI, HTTPException
from pydantic import BaseModel
import ollama
app = FastAPI()
class ChatRequest(BaseModel):
message: str
model: str = "llama3"
system: str = None
class ChatResponse(BaseModel):
response: str
@app.post("/chat", responsemodel=ChatResponse)
async def chat(request: ChatRequest):
try:
messages = []
if request.system:
messages.append({"role": "system", "content": request.system})
messages.append({"role": "user", "content": request.message})
response = ollama.chat(
model=request.model,
messages=messages
)
return ChatResponse(response=response["message"]["content"])
except Exception as e:
raise HTTPException(statuscode=500, detail=str(e))
@app.get("/models")
async def listmodels():
models = ollama.list()
return {"models": [m["name"] for m in models["models"]]}
Run: uvicorn main:app --reload
Troubleshooting
1. Common Issues
# Model not found
ollama pull llama3 # Download ulang
Out of memory
Gunakan model lebih kecil atau quantized
ollama run llama3:8b-q40
Server not running
ollama serve # Start server manually
Port conflict
OLLAMAHOST=0.0.0.0:11435 ollama serve
2. Logs dan Debug
# View logs
journalctl -u ollama -f # Linux dengan systemd
Environment variables
OLLAMADEBUG=1 ollama serve # Enable debug mode
OLLAMA_HOST=0.0.0.0 ollama serve # Listen on all interfaces
Kesimpulan
Ollama memudahkan deployment LLM secara lokal dengan:
Key takeaways:
- Gunakan quantized models untuk hardware terbatas
- Custom Modelfile untuk behavior spesifik
- OpenAI compatibility untuk easy migration
- Ideal untuk development dan privacy-sensitive apps