Complete Ollama Tutorial: Deploy LLMs Locally

# Tutorial Lengkap Ollama: Deploy LLM Secara Lokal Ollama adalah tool open-source yang memudahkan Anda menjalankan Large Language Models (LLM) secara lokal di komputer Anda. Dengan Ollama, Anda dapat...

By Ruby Abdullah · · tutorial
OllamaLLMAILocal AIPythonMachine Learning

Complete Ollama Tutorial: Deploy LLMs Locally

Ollama is an open-source tool that makes it easy to run Large Language Models (LLMs) locally on your computer. With Ollama, you can use models like Llama 3, Mistral, Gemma, and many more without requiring internet connection or paid APIs.

Why Ollama?

Benefits of using Ollama:
  • Privacy: Data never leaves your computer
  • No API costs: Free after downloading models
  • Offline capable: Works without internet
  • Easy setup: One command to run models
  • OpenAI-compatible API: Drop-in replacement for OpenAI

Use Cases:
  • Development and testing AI applications
  • Private/sensitive data processing
  • Offline AI applications
  • Learning and experimenting with LLMs
  • Cost-effective inference

Installation

1. Install on Linux

# Install with script

curl -fsSL https://ollama.com/install.sh | sh

Verify installation

ollama --version

2. Install on macOS

# Download from website or use Homebrew

brew install ollama

Or download .dmg from https://ollama.com/download

3. Install on Windows

Download installer from ollama.com/download and run it.

4. Install via Docker

# Pull image

docker pull ollama/ollama

Run container

docker run -d -v ollama:/root/.ollama -p 11434:11434 --name ollama ollama/ollama

With GPU (NVIDIA)

docker run -d --gpus all -v ollama:/root/.ollama -p 11434:11434 --name ollama ollama/ollama

Quick Start

1. Pull and Run Model

# Download and run Llama 3

ollama run llama3

Or other models

ollama run mistral

ollama run gemma:7b

ollama run phi3

ollama run codellama

Chat directly in terminal

>> Hello, who are you?

I am an AI assistant...

>> /bye # Exit chat

2. Available Models

| Model | Size | Use Case |

|-------|------|----------|

| llama3:8b | 4.7GB | General purpose, balanced |

| llama3:70b | 40GB | High quality responses |

| mistral | 4.1GB | Fast, efficient |

| gemma:7b | 5GB | Google's open model |

| phi3 | 2.2GB | Small, efficient |

| codellama | 3.8GB | Code generation |

| llava | 4.5GB | Vision + Language |

| mixtral | 26GB | Mixture of experts |

# List available models

ollama list

Pull specific version

ollama pull llama3:8b

ollama pull llama3:70b

Remove model

ollama rm llama3:8b

3. Model Commands

# Show model info

ollama show llama3

Copy model (for custom)

ollama cp llama3 my-llama3

Push to registry (if you have account)

ollama push username/my-model

REST API

Ollama provides an OpenAI-compatible REST API.

1. Generate Completion

# Simple generation

curl http://localhost:11434/api/generate -d '{

"model": "llama3",

"prompt": "Explain machine learning in 2 sentences"

}'

With streaming disabled

curl http://localhost:11434/api/generate -d '{

"model": "llama3",

"prompt": "Hello",

"stream": false

}'

2. Chat API

curl http://localhost:11434/api/chat -d '{

"model": "llama3",

"messages": [

{"role": "system", "content": "You are a helpful assistant"},

{"role": "user", "content": "What is Python?"}

],

"stream": false

}'

3. Embeddings

curl http://localhost:11434/api/embeddings -d '{

"model": "llama3",

"prompt": "Text to embed"

}'

Python Integration

1. Using requests

import requests

import json

def generate(prompt, model="llama3"):

response = requests.post(

"http://localhost:11434/api/generate",

json={

"model": model,

"prompt": prompt,

"stream": False

}

)

return response.json()["response"]

def chat(messages, model="llama3"):

response = requests.post(

"http://localhost:11434/api/chat",

json={

"model": model,

"messages": messages,

"stream": False

}

)

return response.json()["message"]["content"]

Usage

result = generate("Explain quantum computing")

print(result)

messages = [

{"role": "system", "content": "You are a math teacher"},

{"role": "user", "content": "Explain the Pythagorean theorem"}

]

result = chat(messages)

print(result)

2. Using ollama-python Library

pip install ollama

import ollama

Generate

response = ollama.generate(

model='llama3',

prompt='What is a neural network?'

)

print(response['response'])

Chat

response = ollama.chat(

model='llama3',

messages=[

{'role': 'system', 'content': 'You are a coding assistant'},

{'role': 'user', 'content': 'Write a Python function for factorial'}

]

)

print(response['message']['content'])

Streaming

stream = ollama.chat(

model='llama3',

messages=[{'role': 'user', 'content': 'Tell me about AI'}],

stream=True

)

for chunk in stream:

print(chunk['message']['content'], end='', flush=True)

Embeddings

embeddings = ollama.embeddings(

model='llama3',

prompt='Hello world'

)

print(len(embeddings['embedding'])) # Vector dimension

3. Async Support

import asyncio

import ollama

async def asyncchat():

response = await ollama.AsyncClient().chat(

model='llama3',

messages=[{'role': 'user', 'content': 'Hello!'}]

)

return response['message']['content']

Run

result = asyncio.run(asyncchat())

print(result)

OpenAI Compatibility

Ollama can be used as a drop-in replacement for OpenAI API.

1. With OpenAI Python Library

from openai import OpenAI

Point to Ollama server

client = OpenAI(

baseurl="http://localhost:11434/v1",

apikey="ollama" # Required but not used

)

Chat completion

response = client.chat.completions.create(

model="llama3",

messages=[

{"role": "system", "content": "You are a helpful assistant."},

{"role": "user", "content": "What is Python?"}

]

)

print(response.choices[0].message.content)

Streaming

stream = client.chat.completions.create(

model="llama3",

messages=[{"role": "user", "content": "Tell me a story"}],

stream=True

)

for chunk in stream:

if chunk.choices[0].delta.content:

print(chunk.choices[0].delta.content, end="")

2. With LangChain

from langchaincommunity.llms import Ollama

from langchaincommunity.chatmodels import ChatOllama

from langchain.chains import LLMChain

from langchain.prompts import PromptTemplate

LLM

llm = Ollama(model="llama3")

response = llm.invoke("Explain machine learning")

print(response)

Chat Model

chat = ChatOllama(model="llama3", temperature=0.7)

from langchain.schema import HumanMessage, SystemMessage

messages = [

SystemMessage(content="You are a helpful assistant"),

HumanMessage(content="What is deep learning?")

]

response = chat.invoke(messages)

print(response.content)

Chain

template = """Question: {question}

Answer: Let's think step by step."""

prompt = PromptTemplate(template=template, inputvariables=["question"])

chain = LLMChain(llm=llm, prompt=prompt)

result = chain.invoke({"question": "What is 25 * 4?"})

print(result["text"])

Custom Models with Modelfile

1. Basic Modelfile

# Modelfile

FROM llama3

Set parameters

PARAMETER temperature 0.7

PARAMETER topp 0.9

PARAMETER topk 40

System prompt

SYSTEM """

You are a helpful AI assistant.

Answer concisely and clearly.

"""

# Create model from Modelfile

ollama create my-assistant -f Modelfile

Run custom model

ollama run my-assistant

2. Advanced Modelfile

# Modelfile for coding assistant

FROM codellama

PARAMETER temperature 0.2

PARAMETER numctx 4096

PARAMETER repeatpenalty 1.1

SYSTEM """

You are an expert programmer. Follow these rules:

  • Write clean, well-documented code
  • Follow best practices and design patterns
  • Explain your code when asked
  • Use type hints in Python
  • Write tests when appropriate
  • """

    Template for formatting

    TEMPLATE """{{ if .System }}<|system|>

    {{ .System }}<|end|>

    {{ end }}{{ if .Prompt }}<|user|>

    {{ .Prompt }}<|end|>

    {{ end }}<|assistant|>

    {{ .Response }}<|end|>

    """

    3. Fine-tuned Model Import

    # Import GGUF model
    

    FROM ./my-finetuned-model.gguf

    PARAMETER temperature 0.8

    SYSTEM "Custom fine-tuned model for specific task"

    # Create from GGUF file
    

    ollama create my-finetuned -f Modelfile

    Multi-Modal Models (Vision)

    1. LLaVA (Vision + Language)

    # Pull LLaVA model
    

    ollama pull llava

    Run with image

    ollama run llava "Describe this image: ./photo.jpg"

    2. Python with Image

    import ollama
    

    import base64

    def encodeimage(imagepath):

    with open(imagepath, "rb") as f:

    return base64.b64encode(f.read()).decode()

    Analyze image

    response = ollama.chat(

    model='llava',

    messages=[{

    'role': 'user',

    'content': 'What do you see in this image?',

    'images': [encodeimage('photo.jpg')]

    }]

    )

    print(response['message']['content'])

    Performance Optimization

    1. GPU Configuration

    # Check GPU usage
    

    nvidia-smi

    Set specific GPU

    CUDAVISIBLEDEVICES=0 ollama serve

    Multiple GPUs

    CUDAVISIBLEDEVICES=0,1 ollama serve

    2. Memory Management

    # Set context length (reduce to save memory)
    

    ollama run llama3 --num-ctx 2048

    In Modelfile

    PARAMETER numctx 2048

    PARAMETER numgpu 1 # Layers on GPU

    3. Quantization Options

    # Pull quantized versions
    

    ollama pull llama3:8b-q40 # 4-bit quantization

    ollama pull llama3:8b-q80 # 8-bit quantization

    Smaller = faster, larger = better quality

    Building Applications

    1. Simple Chatbot

    import ollama
    
    

    class Chatbot:

    def init(self, model="llama3", systemprompt=None):

    self.model = model

    self.messages = []

    if systemprompt:

    self.messages.append({

    "role": "system",

    "content": systemprompt

    })

    def chat(self, userinput):

    self.messages.append({

    "role": "user",

    "content": userinput

    })

    response = ollama.chat(

    model=self.model,

    messages=self.messages

    )

    assistantmessage = response["message"]["content"]

    self.messages.append({

    "role": "assistant",

    "content": assistantmessage

    })

    return assistantmessage

    def clear(self):

    self.messages = self.messages[:1] if self.messages else []

    Usage

    bot = Chatbot(systemprompt="You are a friendly assistant")

    while True:

    userinput = input("You: ")

    if userinput.lower() in ['quit', 'exit']:

    break

    response = bot.chat(userinput)

    print(f"Bot: {response}\n")

    2. RAG with Ollama

    import ollama
    

    from langchaincommunity.embeddings import OllamaEmbeddings

    from langchaincommunity.vectorstores import Chroma

    from langchain.textsplitter import RecursiveCharacterTextSplitter

    Setup embeddings

    embeddings = OllamaEmbeddings(model="llama3")

    Sample documents

    documents = [

    "Python is a popular programming language.",

    "Machine learning is a branch of AI.",

    "Deep learning uses neural networks."

    ]

    Create vector store

    textsplitter = RecursiveCharacterTextSplitter(chunksize=500)

    texts = textsplitter.createdocuments(documents)

    vectorstore = Chroma.fromdocuments(texts, embeddings)

    RAG function

    def ragquery(question):

    # Retrieve relevant docs

    docs = vectorstore.similaritysearch(question, k=2)

    context = "\n".join([doc.pagecontent for doc in docs])

    # Generate answer

    prompt = f"""Based on the following context, answer the question.

    Context:

    {context}

    Question: {question}

    Answer:"""

    response = ollama.generate(model="llama3", prompt=prompt)

    return response["response"]

    Usage

    answer = ragquery("What is machine learning?")

    print(answer)

    3. FastAPI Server

    from fastapi import FastAPI, HTTPException
    

    from pydantic import BaseModel

    import ollama

    app = FastAPI()

    class ChatRequest(BaseModel):

    message: str

    model: str = "llama3"

    system: str = None

    class ChatResponse(BaseModel):

    response: str

    @app.post("/chat", responsemodel=ChatResponse)

    async def chat(request: ChatRequest):

    try:

    messages = []

    if request.system:

    messages.append({"role": "system", "content": request.system})

    messages.append({"role": "user", "content": request.message})

    response = ollama.chat(

    model=request.model,

    messages=messages

    )

    return ChatResponse(response=response["message"]["content"])

    except Exception as e:

    raise HTTPException(statuscode=500, detail=str(e))

    @app.get("/models")

    async def listmodels():

    models = ollama.list()

    return {"models": [m["name"] for m in models["models"]]}

    Run: uvicorn main:app --reload

    Troubleshooting

    1. Common Issues

    # Model not found
    

    ollama pull llama3 # Re-download

    Out of memory

    Use smaller or quantized model

    ollama run llama3:8b-q40

    Server not running

    ollama serve # Start server manually

    Port conflict

    OLLAMAHOST=0.0.0.0:11435 ollama serve

    2. Logs and Debug

    # View logs
    

    journalctl -u ollama -f # Linux with systemd

    Environment variables

    OLLAMADEBUG=1 ollama serve # Enable debug mode

    OLLAMA_HOST=0.0.0.0 ollama serve # Listen on all interfaces

    Conclusion

    Ollama simplifies local LLM deployment with:

  • Simple CLI: Pull and run models with one command
  • REST API: OpenAI-compatible for easy integration
  • Custom Models: Create models with Modelfile
  • Multi-Modal: Support vision models like LLaVA
  • Performance: GPU acceleration and quantization
  • Key takeaways:

    • Use quantized models for limited hardware
    • Custom Modelfile for specific behavior
    • OpenAI compatibility for easy migration
    • Ideal for development and privacy-sensitive apps

    Related Articles

    MLX Tutorial: Apple's Machine Learning Framework for Apple Silicon

    Tutorial MLX: Framework Machine Learning Apple untuk Apple Silicon MLX adalah framework machine learning open-source dar...

    DSPy: A Framework for Programmatic LLM Optimization

    DSPy: Framework untuk Optimasi LLM Secara Programatik Prompt engineering secara manual adalah proses yang melelahkan dan...

    Complete LlamaIndex Tutorial: Building RAG Applications with LLMs

    Tutorial Lengkap LlamaIndex: Membangun Aplikasi RAG dengan LLM LlamaIndex adalah framework data yang powerful untuk memb...

    Complete vLLM Tutorial: High-Performance LLM Serving

    Tutorial Lengkap vLLM: High-Performance LLM Serving vLLM adalah library Python untuk inference dan serving LLM dengan pe...