Tutorial Lengkap Ollama: Deploy LLMs Secara Lokal

# Tutorial Lengkap Ollama: Deploy LLM Secara Lokal Ollama adalah tool open-source yang memudahkan Anda menjalankan Large Language Models (LLM) secara lokal di komputer Anda. Dengan Ollama, Anda dapat...

By Ruby Abdullah · · tutorial
OllamaLLMAILocal AIPythonMachine Learning

Tutorial Lengkap Ollama: Deploy LLM Secara Lokal

Ollama adalah tool open-source yang memudahkan Anda menjalankan Large Language Models (LLM) secara lokal di komputer Anda. Dengan Ollama, Anda dapat menggunakan model seperti Llama 3, Mistral, Gemma, dan banyak lagi tanpa memerlukan koneksi internet atau API berbayar.

Mengapa Ollama?

Keuntungan menggunakan Ollama:
  • Privacy: Data tidak keluar dari komputer Anda
  • No API costs: Gratis setelah download model
  • Offline capable: Bekerja tanpa internet
  • Easy setup: Satu command untuk menjalankan model
  • OpenAI-compatible API: Drop-in replacement untuk OpenAI

Use Cases:
  • Development dan testing aplikasi AI
  • Private/sensitive data processing
  • Offline AI applications
  • Learning dan eksperimen dengan LLM
  • Cost-effective inference

Instalasi

1. Install di Linux

# Install dengan script

curl -fsSL https://ollama.com/install.sh | sh

Verify installation

ollama --version

2. Install di macOS

# Download dari website atau gunakan Homebrew

brew install ollama

Atau download .dmg dari https://ollama.com/download

3. Install di Windows

Download installer dari ollama.com/download dan jalankan.

4. Install via Docker

# Pull image

docker pull ollama/ollama

Run container

docker run -d -v ollama:/root/.ollama -p 11434:11434 --name ollama ollama/ollama

Dengan GPU (NVIDIA)

docker run -d --gpus all -v ollama:/root/.ollama -p 11434:11434 --name ollama ollama/ollama

Quick Start

1. Pull dan Run Model

# Download dan run Llama 3

ollama run llama3

Atau model lain

ollama run mistral

ollama run gemma:7b

ollama run phi3

ollama run codellama

Chat langsung di terminal

>> Halo, siapa kamu?

Saya adalah AI assistant...

>> /bye # Keluar dari chat

2. Model yang Tersedia

| Model | Size | Use Case |

|-------|------|----------|

| llama3:8b | 4.7GB | General purpose, balanced |

| llama3:70b | 40GB | High quality responses |

| mistral | 4.1GB | Fast, efficient |

| gemma:7b | 5GB | Google's open model |

| phi3 | 2.2GB | Small, efficient |

| codellama | 3.8GB | Code generation |

| llava | 4.5GB | Vision + Language |

| mixtral | 26GB | Mixture of experts |

# List available models

ollama list

Pull specific version

ollama pull llama3:8b

ollama pull llama3:70b

Remove model

ollama rm llama3:8b

3. Model Commands

# Show model info

ollama show llama3

Copy model (untuk custom)

ollama cp llama3 my-llama3

Push to registry (jika punya akun)

ollama push username/my-model

REST API

Ollama menyediakan REST API yang OpenAI-compatible.

1. Generate Completion

# Simple generation

curl http://localhost:11434/api/generate -d '{

"model": "llama3",

"prompt": "Jelaskan apa itu machine learning dalam 2 kalimat"

}'

Dengan streaming disabled

curl http://localhost:11434/api/generate -d '{

"model": "llama3",

"prompt": "Hello",

"stream": false

}'

2. Chat API

curl http://localhost:11434/api/chat -d '{

"model": "llama3",

"messages": [

{"role": "system", "content": "Kamu adalah asisten yang helpful"},

{"role": "user", "content": "Apa itu Python?"}

],

"stream": false

}'

3. Embeddings

curl http://localhost:11434/api/embeddings -d '{

"model": "llama3",

"prompt": "Teks untuk di-embed"

}'

Python Integration

1. Menggunakan requests

import requests

import json

def generate(prompt, model="llama3"):

response = requests.post(

"http://localhost:11434/api/generate",

json={

"model": model,

"prompt": prompt,

"stream": False

}

)

return response.json()["response"]

def chat(messages, model="llama3"):

response = requests.post(

"http://localhost:11434/api/chat",

json={

"model": model,

"messages": messages,

"stream": False

}

)

return response.json()["message"]["content"]

Usage

result = generate("Jelaskan quantum computing")

print(result)

messages = [

{"role": "system", "content": "Kamu adalah guru matematika"},

{"role": "user", "content": "Jelaskan teorema Pythagoras"}

]

result = chat(messages)

print(result)

2. Menggunakan ollama-python Library

pip install ollama

import ollama

Generate

response = ollama.generate(

model='llama3',

prompt='Apa itu neural network?'

)

print(response['response'])

Chat

response = ollama.chat(

model='llama3',

messages=[

{'role': 'system', 'content': 'Kamu adalah asisten coding'},

{'role': 'user', 'content': 'Tulis fungsi Python untuk factorial'}

]

)

print(response['message']['content'])

Streaming

stream = ollama.chat(

model='llama3',

messages=[{'role': 'user', 'content': 'Ceritakan tentang AI'}],

stream=True

)

for chunk in stream:

print(chunk['message']['content'], end='', flush=True)

Embeddings

embeddings = ollama.embeddings(

model='llama3',

prompt='Hello world'

)

print(len(embeddings['embedding'])) # Vector dimension

3. Async Support

import asyncio

import ollama

async def asyncchat():

response = await ollama.AsyncClient().chat(

model='llama3',

messages=[{'role': 'user', 'content': 'Hello!'}]

)

return response['message']['content']

Run

result = asyncio.run(asyncchat())

print(result)

OpenAI Compatibility

Ollama dapat digunakan sebagai drop-in replacement untuk OpenAI API.

1. Dengan OpenAI Python Library

from openai import OpenAI

Point ke Ollama server

client = OpenAI(

baseurl="http://localhost:11434/v1",

apikey="ollama" # Required tapi tidak digunakan

)

Chat completion

response = client.chat.completions.create(

model="llama3",

messages=[

{"role": "system", "content": "You are a helpful assistant."},

{"role": "user", "content": "What is Python?"}

]

)

print(response.choices[0].message.content)

Streaming

stream = client.chat.completions.create(

model="llama3",

messages=[{"role": "user", "content": "Tell me a story"}],

stream=True

)

for chunk in stream:

if chunk.choices[0].delta.content:

print(chunk.choices[0].delta.content, end="")

2. Dengan LangChain

from langchaincommunity.llms import Ollama

from langchaincommunity.chatmodels import ChatOllama

from langchain.chains import LLMChain

from langchain.prompts import PromptTemplate

LLM

llm = Ollama(model="llama3")

response = llm.invoke("Explain machine learning")

print(response)

Chat Model

chat = ChatOllama(model="llama3", temperature=0.7)

from langchain.schema import HumanMessage, SystemMessage

messages = [

SystemMessage(content="You are a helpful assistant"),

HumanMessage(content="What is deep learning?")

]

response = chat.invoke(messages)

print(response.content)

Chain

template = """Question: {question}

Answer: Let's think step by step."""

prompt = PromptTemplate(template=template, inputvariables=["question"])

chain = LLMChain(llm=llm, prompt=prompt)

result = chain.invoke({"question": "What is 25 * 4?"})

print(result["text"])

Custom Models dengan Modelfile

1. Basic Modelfile

# Modelfile

FROM llama3

Set parameters

PARAMETER temperature 0.7

PARAMETER topp 0.9

PARAMETER topk 40

System prompt

SYSTEM """

Kamu adalah asisten AI yang helpful dan berbicara dalam Bahasa Indonesia.

Jawab dengan singkat dan jelas.

"""

# Create model dari Modelfile

ollama create my-assistant -f Modelfile

Run custom model

ollama run my-assistant

2. Advanced Modelfile

# Modelfile untuk coding assistant

FROM codellama

PARAMETER temperature 0.2

PARAMETER numctx 4096

PARAMETER repeatpenalty 1.1

SYSTEM """

You are an expert programmer. Follow these rules:

  • Write clean, well-documented code
  • Follow best practices and design patterns
  • Explain your code when asked
  • Use type hints in Python
  • Write tests when appropriate
  • """

    Template untuk formatting

    TEMPLATE """{{ if .System }}<|system|>

    {{ .System }}<|end|>

    {{ end }}{{ if .Prompt }}<|user|>

    {{ .Prompt }}<|end|>

    {{ end }}<|assistant|>

    {{ .Response }}<|end|>

    """

    3. Fine-tuned Model Import

    # Import GGUF model
    

    FROM ./my-finetuned-model.gguf

    PARAMETER temperature 0.8

    SYSTEM "Custom fine-tuned model for specific task"

    # Create dari GGUF file
    

    ollama create my-finetuned -f Modelfile

    Multi-Modal Models (Vision)

    1. LLaVA (Vision + Language)

    # Pull LLaVA model
    

    ollama pull llava

    Run dengan image

    ollama run llava "Describe this image: ./photo.jpg"

    2. Python dengan Image

    import ollama
    

    import base64

    def encodeimage(imagepath):

    with open(imagepath, "rb") as f:

    return base64.b64encode(f.read()).decode()

    Analyze image

    response = ollama.chat(

    model='llava',

    messages=[{

    'role': 'user',

    'content': 'What do you see in this image?',

    'images': [encodeimage('photo.jpg')]

    }]

    )

    print(response['message']['content'])

    Performance Optimization

    1. GPU Configuration

    # Check GPU usage
    

    nvidia-smi

    Set specific GPU

    CUDAVISIBLEDEVICES=0 ollama serve

    Multiple GPUs

    CUDAVISIBLEDEVICES=0,1 ollama serve

    2. Memory Management

    # Set context length (reduce untuk hemat memory)
    

    ollama run llama3 --num-ctx 2048

    Dalam Modelfile

    PARAMETER numctx 2048

    PARAMETER numgpu 1 # Layers di GPU

    3. Quantization Options

    # Pull quantized versions
    

    ollama pull llama3:8b-q40 # 4-bit quantization

    ollama pull llama3:8b-q80 # 8-bit quantization

    Smaller = faster, larger = better quality

    Building Applications

    1. Simple Chatbot

    import ollama
    
    

    class Chatbot:

    def init(self, model="llama3", systemprompt=None):

    self.model = model

    self.messages = []

    if systemprompt:

    self.messages.append({

    "role": "system",

    "content": systemprompt

    })

    def chat(self, userinput):

    self.messages.append({

    "role": "user",

    "content": userinput

    })

    response = ollama.chat(

    model=self.model,

    messages=self.messages

    )

    assistantmessage = response["message"]["content"]

    self.messages.append({

    "role": "assistant",

    "content": assistantmessage

    })

    return assistantmessage

    def clear(self):

    self.messages = self.messages[:1] if self.messages else []

    Usage

    bot = Chatbot(systemprompt="Kamu adalah asisten yang ramah")

    while True:

    userinput = input("You: ")

    if userinput.lower() in ['quit', 'exit']:

    break

    response = bot.chat(userinput)

    print(f"Bot: {response}\n")

    2. RAG dengan Ollama

    import ollama
    

    from langchaincommunity.embeddings import OllamaEmbeddings

    from langchaincommunity.vectorstores import Chroma

    from langchain.textsplitter import RecursiveCharacterTextSplitter

    Setup embeddings

    embeddings = OllamaEmbeddings(model="llama3")

    Sample documents

    documents = [

    "Python adalah bahasa pemrograman yang populer.",

    "Machine learning adalah cabang dari AI.",

    "Deep learning menggunakan neural networks."

    ]

    Create vector store

    textsplitter = RecursiveCharacterTextSplitter(chunksize=500)

    texts = textsplitter.createdocuments(documents)

    vectorstore = Chroma.fromdocuments(texts, embeddings)

    RAG function

    def ragquery(question):

    # Retrieve relevant docs

    docs = vectorstore.similaritysearch(question, k=2)

    context = "\n".join([doc.pagecontent for doc in docs])

    # Generate answer

    prompt = f"""Based on the following context, answer the question.

    Context:

    {context}

    Question: {question}

    Answer:"""

    response = ollama.generate(model="llama3", prompt=prompt)

    return response["response"]

    Usage

    answer = ragquery("Apa itu machine learning?")

    print(answer)

    3. FastAPI Server

    from fastapi import FastAPI, HTTPException
    

    from pydantic import BaseModel

    import ollama

    app = FastAPI()

    class ChatRequest(BaseModel):

    message: str

    model: str = "llama3"

    system: str = None

    class ChatResponse(BaseModel):

    response: str

    @app.post("/chat", responsemodel=ChatResponse)

    async def chat(request: ChatRequest):

    try:

    messages = []

    if request.system:

    messages.append({"role": "system", "content": request.system})

    messages.append({"role": "user", "content": request.message})

    response = ollama.chat(

    model=request.model,

    messages=messages

    )

    return ChatResponse(response=response["message"]["content"])

    except Exception as e:

    raise HTTPException(statuscode=500, detail=str(e))

    @app.get("/models")

    async def listmodels():

    models = ollama.list()

    return {"models": [m["name"] for m in models["models"]]}

    Run: uvicorn main:app --reload

    Troubleshooting

    1. Common Issues

    # Model not found
    

    ollama pull llama3 # Download ulang

    Out of memory

    Gunakan model lebih kecil atau quantized

    ollama run llama3:8b-q40

    Server not running

    ollama serve # Start server manually

    Port conflict

    OLLAMAHOST=0.0.0.0:11435 ollama serve

    2. Logs dan Debug

    # View logs
    

    journalctl -u ollama -f # Linux dengan systemd

    Environment variables

    OLLAMADEBUG=1 ollama serve # Enable debug mode

    OLLAMA_HOST=0.0.0.0 ollama serve # Listen on all interfaces

    Kesimpulan

    Ollama memudahkan deployment LLM secara lokal dengan:

  • Simple CLI: Pull dan run model dengan satu command
  • REST API: OpenAI-compatible untuk easy integration
  • Custom Models: Buat model dengan Modelfile
  • Multi-Modal: Support vision models seperti LLaVA
  • Performance: GPU acceleration dan quantization
  • Key takeaways:

    • Gunakan quantized models untuk hardware terbatas
    • Custom Modelfile untuk behavior spesifik
    • OpenAI compatibility untuk easy migration
    • Ideal untuk development dan privacy-sensitive apps

    Artikel Terkait

    Tutorial MLX: Framework Machine Learning Apple untuk Apple Silicon

    Tutorial MLX: Framework Machine Learning Apple untuk Apple Silicon MLX adalah framework machine learning open-source dar...

    DSPy: Framework untuk Optimasi LLM Secara Programatik

    DSPy: Framework untuk Optimasi LLM Secara Programatik Prompt engineering secara manual adalah proses yang melelahkan dan...

    Tutorial Lengkap LlamaIndex: Membangun Aplikasi RAG dengan LLM

    Tutorial Lengkap LlamaIndex: Membangun Aplikasi RAG dengan LLM LlamaIndex adalah framework data yang powerful untuk memb...

    Tutorial Lengkap vLLM: High-Performance LLM Serving

    Tutorial Lengkap vLLM: High-Performance LLM Serving vLLM adalah library Python untuk inference dan serving LLM dengan pe...