Tutorial MLX: Framework Machine Learning Apple untuk Apple Silicon
MLX adalah framework machine learning open-source dari Apple yang dirancang khusus untuk Apple Silicon (M1, M2, M3, M4). Framework ini menawarkan API yang familiar bagi pengguna NumPy dan PyTorch, dengan performa yang dioptimalkan untuk chip Apple. MLX menjadi pilihan utama bagi developer yang ingin menjalankan model ML secara lokal di perangkat Mac tanpa memerlukan GPU NVIDIA.
Dalam tutorial ini, kita akan mempelajari cara instalasi, penggunaan dasar, fine-tuning model LLM, inferensi model, dan best practices untuk memaksimalkan performa MLX di perangkat Apple Silicon.
Mengapa MLX?
Sebelum memulai, mari kita pahami mengapa MLX layak dipelajari:
Instalasi
Prasyarat
- macOS 13.5 atau lebih baru
- Apple Silicon (M1/M2/M3/M4)
- Python 3.9 atau lebih baru
Instalasi MLX Core
pip install mlx
Instalasi MLX-LM (untuk Language Models)
pip install mlx-lm
Instalasi MLX-VLM (untuk Vision-Language Models)
pip install mlx-vlm
Instalasi dari Source (Opsional)
git clone https://github.com/ml-explore/mlx.git
cd mlx
pip install -e .
Verifikasi Instalasi
import mlx.core as mx
import mlx.nn as nn
print(f"MLX version: {mx.version}")
print(f"Default device: {mx.defaultdevice()}")
Test sederhana
a = mx.array([1, 2, 3, 4, 5])
print(f"Array: {a}")
print(f"Sum: {mx.sum(a)}")
Output yang diharapkan:
MLX version: 0.x.x
Default device: Device(gpu, 0)
Array: array([1, 2, 3, 4, 5], dtype=int32)
Sum: array(15, dtype=int32)
Penggunaan Dasar
Operasi Array
MLX array sangat mirip dengan NumPy, tetapi dioptimalkan untuk Apple Silicon:
import mlx.core as mx
Membuat array
a = mx.array([1.0, 2.0, 3.0, 4.0])
b = mx.ones((3, 4))
c = mx.zeros((2, 3))
d = mx.random.normal((5, 5))
print(f"Shape of b: {b.shape}")
print(f"Dtype of a: {a.dtype}")
Operasi matematika
x = mx.array([[1, 2], [3, 4]], dtype=mx.float32)
y = mx.array([[5, 6], [7, 8]], dtype=mx.float32)
Element-wise operations
print(f"Addition: {x + y}")
print(f"Multiplication: {x y}")
Matrix multiplication
print(f"MatMul: {x @ y}")
Reduction operations
print(f"Sum: {mx.sum(x)}")
print(f"Mean: {mx.mean(x)}")
print(f"Max: {mx.max(x)}")
Lazy Evaluation
Salah satu fitur unik MLX adalah lazy evaluation. Komputasi tidak langsung dieksekusi sampai hasilnya dibutuhkan:
import mlx.core as mx
a = mx.ones((1000, 1000))
b = mx.ones((1000, 1000))
Operasi ini belum dieksekusi
c = a + b
d = c 2
Evaluasi terjadi saat kita membutuhkan hasilnya
mx.eval(d)
print(d)
Atau saat kita mengkonversi ke Python/NumPy
result = d.tolist()
Kontrol Device
import mlx.core as mx
Cek device default
print(f"Default device: {mx.defaultdevice()}")
Jalankan di CPU
mx.setdefaultdevice(mx.cpu)
a = mx.ones((100, 100))
print(f"Device: {mx.defaultdevice()}")
Kembali ke GPU
mx.setdefaultdevice(mx.gpu)
b = mx.ones((100, 100))
print(f"Device: {mx.defaultdevice()}")
Membangun Neural Network dengan MLX
Model Sederhana
MLX menyediakan modul mlx.nn yang mirip dengan PyTorch:
import mlx.core as mx
import mlx.nn as nn
import mlx.optimizers as optim
class SimpleNet(nn.Module):
def init(self):
super().init()
self.layers = [
nn.Linear(784, 256),
nn.Linear(256, 128),
nn.Linear(128, 10),
]
def call(self, x):
for i, layer in enumerate(self.layers[:-1]):
x = nn.relu(layer(x))
return self.layers-1
Inisialisasi model
model = SimpleNet()
Cek parameter
params = model.parameters()
print(f"Total parameters: {sum(p.size for p in model.parameters()['layers'])}")
Training Loop
import mlx.core as mx
import mlx.nn as nn
import mlx.optimizers as optim
import numpy as np
class MLP(nn.Module):
def init(self, inputdim, hiddendim, outputdim):
super().init()
self.linear1 = nn.Linear(inputdim, hiddendim)
self.linear2 = nn.Linear(hiddendim, outputdim)
def call(self, x):
x = nn.relu(self.linear1(x))
return self.linear2(x)
Data sintetis
np.random.seed(42)
X = np.random.randn(1000, 10).astype(np.float32)
y = (X[:, 0] + X[:, 1] > 0).astype(np.int32)
Xtrain = mx.array(X[:800])
ytrain = mx.array(y[:800])
Xtest = mx.array(X[800:])
ytest = mx.array(y[800:])
Model dan optimizer
model = MLP(10, 64, 2)
optimizer = optim.Adam(learningrate=1e-3)
Loss function
def lossfn(model, x, y):
logits = model(x)
return mx.mean(nn.losses.crossentropy(logits, y))
Training
lossandgradfn = nn.valueandgrad(model, lossfn)
for epoch in range(100):
loss, grads = lossandgradfn(model, Xtrain, ytrain)
optimizer.update(model, grads)
mx.eval(model.parameters(), optimizer.state)
if (epoch + 1) % 20 == 0:
# Evaluasi
testlogits = model(Xtest)
testpreds = mx.argmax(testlogits, axis=1)
accuracy = mx.mean(testpreds == ytest)
print(f"Epoch {epoch+1}, Loss: {loss.item():.4f}, Test Acc: {accuracy.item():.4f}")
Convolutional Neural Network
import mlx.core as mx
import mlx.nn as nn
class ConvNet(nn.Module):
def init(self, numclasses=10):
super().init()
self.conv1 = nn.Conv2d(1, 32, kernelsize=3, padding=1)
self.conv2 = nn.Conv2d(32, 64, kernelsize=3, padding=1)
self.pool = nn.MaxPool2d(kernelsize=2, stride=2)
self.fc1 = nn.Linear(64 7 7, 128)
self.fc2 = nn.Linear(128, numclasses)
self.dropout = nn.Dropout(p=0.5)
def call(self, x):
x = self.pool(nn.relu(self.conv1(x)))
x = self.pool(nn.relu(self.conv2(x)))
x = x.reshape(x.shape[0], -1)
x = self.dropout(nn.relu(self.fc1(x)))
return self.fc2(x)
model = ConvNet()
print("ConvNet initialized successfully")
Menggunakan MLX-LM untuk Language Models
MLX-LM adalah library yang memudahkan penggunaan Large Language Models di Apple Silicon.
Inferensi dengan Model dari Hugging Face
# Download dan jalankan model
mlxlm.generate --model mlx-community/Llama-3.2-3B-Instruct-4bit \
--prompt "Explain machine learning in simple terms" \
--max-tokens 200
Inferensi via Python
from mlxlm import load, generate
Load model (otomatis download dari Hugging Face)
model, tokenizer = load("mlx-community/Llama-3.2-3B-Instruct-4bit")
Generate text
prompt = "Jelaskan apa itu machine learning dalam bahasa sederhana:"
messages = [{"role": "user", "content": prompt}]
formattedprompt = tokenizer.applychattemplate(
messages, tokenize=False, addgenerationprompt=True
)
response = generate(
model,
tokenizer,
prompt=formattedprompt,
maxtokens=500,
temp=0.7,
)
print(response)
Streaming Response
from mlxlm import load, streamgenerate
model, tokenizer = load("mlx-community/Mistral-7B-Instruct-v0.3-4bit")
prompt = "Write a Python function to calculate fibonacci numbers:"
messages = [{"role": "user", "content": prompt}]
formatted
prompt = tokenizer.applychattemplate(
messages, tokenize=False, addgenerationprompt=True
)
Streaming output
for token in streamgenerate(
model,
tokenizer,
prompt=formattedprompt,
maxtokens=500,
):
print(token, end="", flush=True)
print()
Konversi Model ke Format MLX
Anda bisa mengkonversi model Hugging Face ke format MLX yang dioptimalkan:
# Konversi dengan quantization 4-bit
mlxlm.convert \
--hf-path meta-llama/Llama-3.2-3B-Instruct \
--mlx-path ./mlx-llama-3.2-3b-4bit \
--quantize \
--q-bits 4
Konversi tanpa quantization
mlxlm.convert \
--hf-path microsoft/Phi-3-mini-4k-instruct \
--mlx-path ./mlx-phi-3-mini
Konversi via Python
from mlxlm import convert
convert(
hfpath="meta-llama/Llama-3.2-1B-Instruct",
mlxpath="./mlx-llama-1b",
quantize=True,
qbits=4,
qgroupsize=64,
)
Fine-Tuning dengan MLX
LoRA Fine-Tuning
MLX-LM mendukung fine-tuning menggunakan LoRA (Low-Rank Adaptation):
# Siapkan data training dalam format JSONL
train.jsonl format:
{"text": "instruction: ... response: ..."}
atau format chat:
{"messages": [{"role": "user", "content": "..."}, {"role": "assistant", "content": "..."}]}
Buat file konfigurasi loraconfig.yaml:
model: "mlx-community/Llama-3.2-3B-Instruct-4bit"
train: true
data: "./trainingdata"
seed: 42
loralayers: 16
batchsize: 4
iters: 1000
valbatches: 25
learningrate: 1e-5
loraparameters:
keys: ["selfattn.qproj", "selfattn.vproj"]
rank: 8
alpha: 16
dropout: 0.0
Jalankan fine-tuning:
mlxlm.lora \
--model mlx-community/Llama-3.2-3B-Instruct-4bit \
--train \
--data ./training
data \
--batch-size 4 \
--lora-layers 16 \
--iters 1000 \
--learning-rate 1e-5
Menyiapkan Data Training
import json
Format data untuk fine-tuning
trainingdata = [
{
"messages": [
{"role": "user", "content": "Apa itu Python?"},
{"role": "assistant", "content": "Python adalah bahasa pemrograman tingkat tinggi yang dirancang untuk kemudahan pembacaan kode. Python mendukung berbagai paradigma pemrograman termasuk prosedural, berorientasi objek, dan fungsional."}
]
},
{
"messages": [
{"role": "user", "content": "Bagaimana cara membuat list di Python?"},
{"role": "assistant", "content": "Untuk membuat list di Python, gunakan tanda kurung siku []. Contoh: mylist = [1, 2, 3, 'hello']. List di Python bisa menyimpan berbagai tipe data."}
]
},
]
Simpan sebagai JSONL
with open("trainingdata/train.jsonl", "w") as f:
for item in trainingdata:
f.write(json.dumps(item, ensureascii=False) + "\n")
Data validasi
with open("trainingdata/valid.jsonl", "w") as f:
for item in trainingdata[:10]:
f.write(json.dumps(item, ensureascii=False) + "\n")
Menggunakan Adapter Hasil Fine-Tuning
from mlxlm import load, generate
Load model dengan adapter
model, tokenizer = load(
"mlx-community/Llama-3.2-3B-Instruct-4bit",
adapter
path="./adapters"
)
Generate dengan model yang sudah di-fine-tune
response = generate(
model,
tokenizer,
prompt="Apa itu Python?",
maxtokens=200
)
print(response)
Fuse Adapter ke Model
Setelah fine-tuning, Anda bisa menggabungkan adapter dengan model dasar:
mlxlm.fuse \
--model mlx-community/Llama-3.2-3B-Instruct-4bit \
--adapter-path ./adapters \
--save-path ./fused-model
MLX untuk Computer Vision
Image Classification
import mlx.core as mx
import mlx.nn as nn
class ResidualBlock(nn.Module):
def init(self, channels):
super().init()
self.conv1 = nn.Conv2d(channels, channels, kernelsize=3, padding=1)
self.bn1 = nn.BatchNorm(channels)
self.conv2 = nn.Conv2d(channels, channels, kernelsize=3, padding=1)
self.bn2 = nn.BatchNorm(channels)
def call(self, x):
residual = x
x = nn.relu(self.bn1(self.conv1(x)))
x = self.bn2(self.conv2(x))
return nn.relu(x + residual)
class SimpleResNet(nn.Module):
def init(self, numclasses=10):
super().init()
self.conv1 = nn.Conv2d(3, 64, kernelsize=7, stride=2, padding=3)
self.bn1 = nn.BatchNorm(64)
self.pool = nn.MaxPool2d(kernelsize=3, stride=2, padding=1)
self.block1 = ResidualBlock(64)
self.block2 = ResidualBlock(64)
self.globalpool = nn.AvgPool2d(kernelsize=7)
self.fc = nn.Linear(64, numclasses)
def call(self, x):
x = self.pool(nn.relu(self.bn1(self.conv1(x))))
x = self.block1(x)
x = self.block2(x)
x = self.globalpool(x)
x = x.reshape(x.shape[0], -1)
return self.fc(x)
model = SimpleResNet(numclasses=10)
dummyinput = mx.random.normal((1, 3, 224, 224))
output = model(dummyinput)
print(f"Output shape: {output.shape}")
MLX Server: OpenAI-Compatible API
MLX-LM menyediakan server yang kompatibel dengan OpenAI API:
# Jalankan server
mlxlm.server --model mlx-community/Llama-3.2-3B-Instruct-4bit --port 8080
Gunakan dengan OpenAI SDK:
from openai import OpenAI
client = OpenAI(
baseurl="http://localhost:8080/v1",
apikey="not-needed"
)
response = client.chat.completions.create(
model="mlx-community/Llama-3.2-3B-Instruct-4bit",
messages=[
{"role": "system", "content": "Kamu adalah asisten yang membantu."},
{"role": "user", "content": "Apa manfaat machine learning?"}
],
maxtokens=300,
temperature=0.7
)
print(response.choices[0].message.content)
Atau gunakan curl:
curl http://localhost:8080/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "mlx-community/Llama-3.2-3B-Instruct-4bit",
"messages": [{"role": "user", "content": "Hello!"}],
"maxtokens": 100
}'
Benchmark dan Performa
Mengukur Throughput
import mlx.core as mx
import time
def benchmarkmatmul(size, iterations=100):
a = mx.random.normal((size, size))
b = mx.random.normal((size, size))
mx.eval(a, b)
start = time.time()
for in range(iterations):
c = a @ b
mx.eval(c)
elapsed = time.time() - start
flops = 2 size3 iterations
gflops = flops / elapsed / 1e9
print(f"Size {size}x{size}: {gflops:.1f} GFLOPS ({elapsed/iterations1000:.2f} ms/iter)")
for size in [512, 1024, 2048, 4096]:
benchmarkmatmul(size)
Perbandingan Kecepatan Inferensi LLM
from mlxlm import load, generate
import time
models
totest = [
"mlx-community/Llama-3.2-1B-Instruct-4bit",
"mlx-community/Llama-3.2-3B-Instruct-4bit",
"mlx-community/Mistral-7B-Instruct-v0.3-4bit",
]
prompt = "Explain the concept of neural networks in detail:"
for model
name in modelstotest:
model, tokenizer = load(modelname)
messages = [{"role": "user", "content": prompt}]
formatted = tokenizer.applychattemplate(
messages, tokenize=False, addgenerationprompt=True
)
start = time.time()
response = generate(model, tokenizer, prompt=formatted, maxtokens=200)
elapsed = time.time() - start
tokens = len(tokenizer.encode(response))
print(f"{modelname}: {tokens/elapsed:.1f} tokens/sec")
Advanced: Custom Operations
Function Transformations
MLX mendukung transformasi fungsi seperti gradient computation dan vectorization:
import mlx.core as mx
Automatic differentiation
def f(x):
return mx.sum(x 2)
Gradient
gradf = mx.grad(f)
x = mx.array([1.0, 2.0, 3.0])
print(f"Gradient: {gradf(x)}")
Value and gradient
valandgradf = mx.valueandgrad(f)
value, gradient = valandgradf(x)
print(f"Value: {value}, Gradient: {gradient}")
Vectorized map
def singlefn(x):
return x 2 + 2 x + 1
batch = mx.array([[1.0], [2.0], [3.0], [4.0]])
results = mx.vmap(singlefn)(batch)
print(f"Vectorized results: {results}")
Custom Layer
import mlx.core as mx
import mlx.nn as nn
class MultiHeadSelfAttention(nn.Module):
def init(self, dims, numheads):
super().init()
self.numheads = numheads
self.headdim = dims // numheads
self.query = nn.Linear(dims, dims)
self.key = nn.Linear(dims, dims)
self.value = nn.Linear(dims, dims)
self.out = nn.Linear(dims, dims)
def call(self, x):
B, T, C = x.shape
q = self.query(x).reshape(B, T, self.numheads, self.headdim).transpose(0, 2, 1, 3)
k = self.key(x).reshape(B, T, self.numheads, self.headdim).transpose(0, 2, 1, 3)
v = self.value(x).reshape(B, T, self.numheads, self.headdim).transpose(0, 2, 1, 3)
scale = self.headdim * -0.5
attn = (q @ k.transpose(0, 1, 3, 2)) scale
attn = mx.softmax(attn, axis=-1)
out = (attn @ v).transpose(0, 2, 1, 3).reshape(B, T, C)
return self.out(out)
Test
attention = MultiHeadSelfAttention(256, 8)
x = mx.random.normal((2, 10, 256))
output = attention(x)
print(f"Attention output shape: {output.shape}")
Best Practices
1. Manfaatkan Lazy Evaluation
import mlx.core as mx
Baik: batch evaluations
a = mx.random.normal((1000, 1000))
b = mx.random.normal((1000, 1000))
c = a @ b
d = c + mx.oneslike(c)
e = mx.sum(d)
mx.eval(e) # Evaluasi sekali di akhir
Kurang optimal: evaluasi terlalu sering
a = mx.random.normal((1000, 1000))
mx.eval(a) # Evaluasi tidak perlu
b = mx.random.normal((1000, 1000))
mx.eval(b) # Evaluasi tidak perlu
c = a @ b
mx.eval(c) # Baru perlu evaluasi
2. Gunakan Quantization untuk Model Besar
from mlxlm import load
4-bit quantization menghemat ~75% memori
model, tokenizer = load("mlx-community/Llama-3.2-3B-Instruct-4bit")
8-bit untuk keseimbangan kualitas dan memori
model, tokenizer = load("mlx-community/Llama-3.2-3B-Instruct-8bit")
3. Kelola Memori dengan Baik
import mlx.core as mx
Hapus array yang tidak diperlukan
largearray = mx.random.normal((10000, 10000))
result = mx.sum(largearray)
mx.eval(result)
del largearray
Gunakan metal.clearcache() jika perlu
mx.metal.clearcache()
4. Batch Processing untuk Efisiensi
from mlxlm import load, generate
model, tokenizer = load("mlx-community/Llama-3.2-3B-Instruct-4bit")
prompts = [
"What is Python?",
"Explain Docker briefly.",
"What is Kubernetes?"
]
for prompt in prompts:
messages = [{"role": "user", "content": prompt}]
formatted = tokenizer.apply
chattemplate(
messages, tokenize=False, add
generationprompt=True
)
response = generate(model, tokenizer, prompt=formatted, max
tokens=100)
print(f"Q: {prompt}")
print(f"A: {response}\n")
5. Monitoring Penggunaan GPU
import mlx.core as mx
Cek memory usage
peakmemory = mx.metal.getpeakmemory()
activememory = mx.metal.getactivememory()
cachememory = mx.metal.getcachememory()
print(f"Peak Memory: {peakmemory / 1e9:.2f} GB")
print(f"Active Memory: {activememory / 1e9:.2f} GB")
print(f"Cache Memory: {cachememory / 1e9:.2f} GB")
Ekosistem MLX
MLX memiliki ekosistem yang terus berkembang:
| Library | Fungsi |
|---------|--------|
| mlx | Core framework (array ops, neural network, optimizers) |
| mlx-lm | Language model inference dan fine-tuning |
| mlx-vlm | Vision-language models |
| mlx-audio | Audio processing dan speech models |
| mlx-data | Data loading dan preprocessing |
| mlx-graphs | Graph neural networks |
| mlx-image | Computer vision utilities |
Model yang Tersedia di MLX Community
Komunitas MLX di Hugging Face menyediakan ratusan model yang sudah dikonversi ke format MLX:
- Llama 3.x (1B, 3B, 8B, 70B) dalam berbagai level quantization
- Mistral/Mixtral untuk model MoE yang efisien
- Phi-3/Phi-4 model kecil dari Microsoft
- Qwen 2.5 dari Alibaba
- Gemma 2 dari Google
- DeepSeek models
- Whisper untuk speech-to-text
- Stable Diffusion untuk image generation
Anda bisa mencari model di: https://huggingface.co/mlx-community
Kesimpulan
MLX adalah framework yang sangat powerful untuk menjalankan machine learning workloads di Apple Silicon. Dengan API yang familiar, dukungan lazy evaluation, dan unified memory architecture, MLX memungkinkan developer untuk memanfaatkan penuh kemampuan chip Apple.
Poin-poin penting yang perlu diingat:
pip install mlx mlx-lm untuk memulaiMLX cocok untuk prototyping, development lokal, penelitian, dan bahkan deployment di perangkat Apple. Dengan semakin banyaknya model yang tersedia di format MLX, framework ini menjadi pilihan utama untuk ML on Apple Silicon.