Groq API Tutorial: Super Fast LLM Inference for AI Applications

# Tutorial Groq API: Inferensi LLM Super Cepat untuk Aplikasi AI ## Introduction Groq telah menjadi salah satu platform inferensi AI paling populer berkat kecepatan luar biasa yang ditawarkan oleh c...

By Ruby Abdullah · · tutorial
GroqLLMAPIAIInference

Groq API Tutorial: Super Fast LLM Inference for AI Applications

Introduction

Groq has become one of the most popular AI inference platforms thanks to the remarkable speed offered by their proprietary Language Processing Unit (LPU) chip. If you have ever been frustrated by high latency when calling large language model APIs, Groq provides a solution with inference speeds that far surpass traditional GPUs.

In this tutorial, we will learn how to use the Groq API comprehensively, from initial setup, basic chat completion usage, to advanced features like streaming, tool use (function calling), vision models, and integration with popular frameworks like LangChain and LlamaIndex. All code examples in this tutorial can be run directly in your local environment.

Groq provides access to various popular open-source models including Llama, Mixtral, and Gemma with output speeds reaching hundreds of tokens per second. Interestingly, the Groq API uses a format compatible with the OpenAI API, making migration from OpenAI to Groq straightforward.

Installation and Setup

Getting an API Key

The first step is to register and obtain an API key from Groq:

  • Visit console.groq.com and create an account
  • Navigate to the API Keys section
  • Click "Create API Key" and save the generated key
  • Installing the Library

    Groq provides an official Python SDK that can be installed via pip:

    pip install groq
    

    For more complete projects, install additional dependencies:

    pip install groq python-dotenv httpx Pillow
    

    Environment Configuration

    Create a .env file to store your API key:

    GROQAPIKEY=gskyourapikeyhere
    

    Verify the installation with a simple script:

    import os
    

    from dotenv import loaddotenv

    from groq import Groq

    loaddotenv()

    client = Groq(apikey=os.environ.get("GROQAPIKEY"))

    Test connection

    models = client.models.list()

    for model in models.data:

    print(f"Model: {model.id}")

    If successful, you will see a list of available models on Groq.

    Basic Usage

    Simple Chat Completion

    The most basic usage of the Groq API is chat completion:

    from groq import Groq
    
    

    client = Groq()

    chatcompletion = client.chat.completions.create(

    messages=[

    {

    "role": "system",

    "content": "You are a helpful and friendly AI assistant."

    },

    {

    "role": "user",

    "content": "Explain what machine learning is in 3 sentences."

    }

    ],

    model="llama-3.3-70b-versatile",

    temperature=0.7,

    maxtokens=1024,

    )

    print(chatcompletion.choices[0].message.content)

    Available Models

    Groq provides several popular models. Here are usage recommendations:

    # Model for general tasks and reasoning
    

    MODELGENERAL = "llama-3.3-70b-versatile"

    Fast model for simple tasks

    MODELFAST = "llama-3.1-8b-instant"

    Mixtral model for multilingual tasks

    MODELMULTILINGUAL = "mixtral-8x7b-32768"

    Model with large context window

    MODELLONGCONTEXT = "llama-3.3-70b-versatile" # 128K context

    Vision model for images

    MODELVISION = "llama-3.2-90b-vision-preview"

    Configuring Parameters

    You can control model output with various parameters:

    response = client.chat.completions.create(
    

    messages=[

    {"role": "user", "content": "Write a short poem about coding."}

    ],

    model="llama-3.3-70b-versatile",

    temperature=0.9, # Creativity (0.0 - 2.0)

    maxtokens=512, # Maximum output tokens

    topp=0.9, # Nucleus sampling

    frequencypenalty=0.5, # Word repetition penalty

    presencepenalty=0.3, # Penalty for already mentioned topics

    stop=["---"], # Stop sequence

    )

    Multi-turn Conversation

    For multi-turn conversations, send the entire conversation history:

    conversationhistory = [
    

    {"role": "system", "content": "You are a patient Python tutor."}

    ]

    def chat(usermessage):

    conversationhistory.append({

    "role": "user",

    "content": usermessage

    })

    response = client.chat.completions.create(

    messages=conversationhistory,

    model="llama-3.3-70b-versatile",

    temperature=0.7,

    maxtokens=1024,

    )

    assistantmessage = response.choices[0].message.content

    conversationhistory.append({

    "role": "assistant",

    "content": assistantmessage

    })

    return assistantmessage

    Example conversation

    print(chat("What is list comprehension in Python?"))

    print(chat("Give an example with filtering."))

    print(chat("How does its performance compare to a regular for loop?"))

    Streaming Response

    Basic Streaming

    Streaming is very useful for better UX in chat applications:

    stream = client.chat.completions.create(
    

    messages=[

    {"role": "user", "content": "Explain the Transformer architecture in detail."}

    ],

    model="llama-3.3-70b-versatile",

    temperature=0.7,

    maxtokens=2048,

    stream=True,

    )

    for chunk in stream:

    content = chunk.choices[0].delta.content

    if content:

    print(content, end="", flush=True)

    print()

    Streaming with Metadata

    You can collect metadata during streaming:

    def streamwithstats(messages, model="llama-3.3-70b-versatile"):
    

    stream = client.chat.completions.create(

    messages=messages,

    model=model,

    stream=True,

    maxtokens=2048,

    )

    fullresponse = ""

    prompttokens = 0

    completiontokens = 0

    for chunk in stream:

    if chunk.choices[0].delta.content:

    content = chunk.choices[0].delta.content

    fullresponse += content

    print(content, end="", flush=True)

    if hasattr(chunk, 'xgroq') and chunk.xgroq:

    if hasattr(chunk.xgroq, 'usage') and chunk.xgroq.usage:

    prompttokens = chunk.xgroq.usage.prompttokens

    completiontokens = chunk.xgroq.usage.completiontokens

    print(f"\n\nToken usage - Prompt: {prompttokens}, "

    f"Completion: {completiontokens}")

    return fullresponse

    result = streamwithstats([

    {"role": "user", "content": "What are the advantages of Groq over GPUs for inference?"}

    ])

    Async Support

    Using the Async Client

    For applications requiring high concurrency:

    import asyncio
    

    from groq import AsyncGroq

    async def asyncchat(prompt):

    client = AsyncGroq()

    response = await client.chat.completions.create(

    messages=[{"role": "user", "content": prompt}],

    model="llama-3.1-8b-instant",

    maxtokens=512,

    )

    return response.choices[0].message.content

    async def batchprocess(prompts):

    tasks = [asyncchat(prompt) for prompt in prompts]

    results = await asyncio.gather(*tasks)

    return results

    Process multiple prompts in parallel

    prompts = [

    "What is Docker?",

    "What is Kubernetes?",

    "What is CI/CD?",

    "What is microservices?",

    ]

    results = asyncio.run(batchprocess(prompts))

    for prompt, result in zip(prompts, results):

    print(f"Q: {prompt}")

    print(f"A: {result[:100]}...")

    print()

    Async Streaming

    async def asyncstream(prompt):
    

    client = AsyncGroq()

    stream = await client.chat.completions.create(

    messages=[{"role": "user", "content": prompt}],

    model="llama-3.3-70b-versatile",

    stream=True,

    maxtokens=1024,

    )

    fullresponse = ""

    async for chunk in stream:

    content = chunk.choices[0].delta.content

    if content:

    fullresponse += content

    print(content, end="", flush=True)

    print()

    return fullresponse

    asyncio.run(asyncstream("Explain the concept of RAG in AI."))

    Tool Use (Function Calling)

    Defining Tools

    Groq supports function calling compatible with the OpenAI format:

    import json
    
    

    tools = [

    {

    "type": "function",

    "function": {

    "name": "getweather",

    "description": "Get weather information for a location",

    "parameters": {

    "type": "object",

    "properties": {

    "location": {

    "type": "string",

    "description": "City name, e.g., New York, London"

    },

    "unit": {

    "type": "string",

    "enum": ["celsius", "fahrenheit"],

    "description": "Temperature unit"

    }

    },

    "required": ["location"]

    }

    }

    },

    {

    "type": "function",

    "function": {

    "name": "searchproducts",

    "description": "Search products by query",

    "parameters": {

    "type": "object",

    "properties": {

    "query": {

    "type": "string",

    "description": "Search keyword"

    },

    "maxprice": {

    "type": "number",

    "description": "Maximum price"

    },

    "category": {

    "type": "string",

    "description": "Product category"

    }

    },

    "required": ["query"]

    }

    }

    }

    ]

    Using Tools in Chat

    def getweather(location, unit="celsius"):
    

    # Simulated weather data

    weatherdata = {

    "New York": {"temp": 22, "condition": "Partly cloudy"},

    "London": {"temp": 16, "condition": "Light rain"},

    "Tokyo": {"temp": 28, "condition": "Sunny"},

    }

    data = weatherdata.get(location, {"temp": 20, "condition": "Unknown"})

    return json.dumps({

    "location": location,

    "temperature": data["temp"],

    "unit": unit,

    "condition": data["condition"]

    })

    def searchproducts(query, maxprice=None, category=None):

    return json.dumps({

    "products": [

    {"name": f"{query} Premium", "price": 99.99, "stock": 25},

    {"name": f"{query} Standard", "price": 49.99, "stock": 100},

    ]

    })

    Function mapping

    availablefunctions = {

    "getweather": getweather,

    "searchproducts": searchproducts,

    }

    def chatwithtools(usermessage):

    messages = [{"role": "user", "content": usermessage}]

    response = client.chat.completions.create(

    messages=messages,

    model="llama-3.3-70b-versatile",

    tools=tools,

    toolchoice="auto",

    maxtokens=1024,

    )

    responsemessage = response.choices[0].message

    if responsemessage.toolcalls:

    messages.append(responsemessage)

    for toolcall in responsemessage.toolcalls:

    functionname = toolcall.function.name

    functionargs = json.loads(toolcall.function.arguments)

    functionresponse = availablefunctionsfunctionname

    messages.append({

    "role": "tool",

    "toolcallid": toolcall.id,

    "name": functionname,

    "content": functionresponse,

    })

    secondresponse = client.chat.completions.create(

    messages=messages,

    model="llama-3.3-70b-versatile",

    maxtokens=1024,

    )

    return secondresponse.choices[0].message.content

    return responsemessage.content

    Test

    print(chatwithtools("What's the weather like in New York today?"))

    print(chatwithtools("Search for laptops under $1000"))

    Vision (Multimodal)

    Image Analysis

    Groq supports vision models for image analysis:

    import base64
    

    import httpx

    def encodeimagefromurl(imageurl):

    response = httpx.get(imageurl)

    return base64.b64encode(response.content).decode("utf-8")

    def encodeimagefromfile(filepath):

    with open(filepath, "rb") as f:

    return base64.b64encode(f.read()).decode("utf-8")

    Analyze image from URL

    def analyzeimage(imageurl, question="Describe this image in detail."):

    response = client.chat.completions.create(

    messages=[

    {

    "role": "user",

    "content": [

    {"type": "text", "text": question},

    {

    "type": "imageurl",

    "imageurl": {

    "url": imageurl,

    },

    },

    ],

    }

    ],

    model="llama-3.2-90b-vision-preview",

    maxtokens=1024,

    )

    return response.choices[0].message.content

    Analyze local image

    def analyzelocalimage(filepath, question="What do you see in this image?"):

    base64image = encodeimagefromfile(filepath)

    response = client.chat.completions.create(

    messages=[

    {

    "role": "user",

    "content": [

    {"type": "text", "text": question},

    {

    "type": "imageurl",

    "imageurl": {

    "url": f"data:image/png;base64,{base64image}",

    },

    },

    ],

    }

    ],

    model="llama-3.2-90b-vision-preview",

    maxtokens=1024,

    )

    return response.choices[0].message.content

    JSON Mode

    Structured Output

    Use JSON mode to get structured output:

    import json
    
    

    def extractentities(text):

    response = client.chat.completions.create(

    messages=[

    {

    "role": "system",

    "content": "Extract entities from the text and return in JSON format "

    "with keys: persons (array), locations (array), "

    "organizations (array), dates (array)."

    },

    {"role": "user", "content": text}

    ],

    model="llama-3.3-70b-versatile",

    temperature=0,

    maxtokens=1024,

    responseformat={"type": "jsonobject"},

    )

    return json.loads(response.choices[0].message.content)

    text = """

    On January 15, 2026, Google CEO Sundar Pichai announced the opening

    of a new office in Singapore. The Minister of Communications welcomed

    this investment.

    """

    entities = extractentities(text)

    print(json.dumps(entities, indent=2))

    Structured Data Extraction

    def analyzesentimentbatch(reviews):
    

    response = client.chat.completions.create(

    messages=[

    {

    "role": "system",

    "content": (

    "Analyze the sentiment of each review. "

    "Return JSON with key 'results' containing an array of objects "

    "with fields: reviewindex (int), sentiment (positive/negative/neutral), "

    "confidence (0.0-1.0), keyphrases (array of string)."

    )

    },

    {

    "role": "user",

    "content": json.dumps(reviews)

    }

    ],

    model="llama-3.3-70b-versatile",

    temperature=0,

    responseformat={"type": "jsonobject"},

    maxtokens=2048,

    )

    return json.loads(response.choices[0].message.content)

    reviews = [

    "Great product, fast shipping!",

    "Disappointing quality, doesn't match the description.",

    "Average, you get what you pay for.",

    ]

    results = analyzesentimentbatch(reviews)

    print(json.dumps(results, indent=2))

    Advanced Usage

    Rate Limiting and Retry

    Implementing robust retry logic:

    import time
    

    from groq import Groq, RateLimitError, APIError

    def robustcompletion(messages, model="llama-3.3-70b-versatile",

    maxretries=3, kwargs):

    client = Groq()

    for attempt in range(maxretries):

    try:

    response = client.chat.completions.create(

    messages=messages,

    model=model,

    kwargs

    )

    return response

    except RateLimitError as e:

    if attempt < maxretries - 1:

    waittime = (2 attempt) + 1

    print(f"Rate limited. Waiting {waittime} seconds...")

    time.sleep(waittime)

    else:

    raise

    except APIError as e:

    if attempt < maxretries - 1:

    print(f"API error: {e}. Retry {attempt + 1}/{maxretries}")

    time.sleep(1)

    else:

    raise

    return None

    Text Chunking for Long Documents

    def processlongdocument(document, chunksize=4000, overlap=200):
    

    chunks = []

    start = 0

    while start < len(document):

    end = start + chunksize

    chunk = document[start:end]

    chunks.append(chunk)

    start = end - overlap

    summaries = []

    for i, chunk in enumerate(chunks):

    response = client.chat.completions.create(

    messages=[

    {

    "role": "system",

    "content": "Summarize the following document section concisely."

    },

    {"role": "user", "content": chunk}

    ],

    model="llama-3.1-8b-instant",

    temperature=0.3,

    maxtokens=512,

    )

    summaries.append(response.choices[0].message.content)

    print(f"Chunk {i+1}/{len(chunks)} processed.")

    # Combine all summaries

    combined = "\n\n".join(summaries)

    finalresponse = client.chat.completions.create(

    messages=[

    {

    "role": "system",

    "content": "Combine the following summaries into a single "

    "coherent and comprehensive summary."

    },

    {"role": "user", "content": combined}

    ],

    model="llama-3.3-70b-versatile",

    temperature=0.3,

    maxtokens=1024,

    )

    return finalresponse.choices[0].message.content

    Integration with LangChain

    from langchaingroq import ChatGroq
    

    from langchaincore.messages import HumanMessage, SystemMessage

    from langchaincore.outputparsers import StrOutputParser

    from langchaincore.prompts import ChatPromptTemplate

    Initialize model

    llm = ChatGroq(

    model="llama-3.3-70b-versatile",

    temperature=0.7,

    maxtokens=1024,

    )

    Simple chain

    prompt = ChatPromptTemplate.frommessages([

    ("system", "You are a {topic} expert who explains concepts simply."),

    ("human", "{question}")

    ])

    chain = prompt | llm | StrOutputParser()

    result = chain.invoke({

    "topic": "machine learning",

    "question": "What's the difference between supervised and unsupervised learning?"

    })

    print(result)

    Integration with LlamaIndex

    from llamaindex.llms.groq import Groq as GroqLLM
    

    from llamaindex.core.llms import ChatMessage

    llm = GroqLLM(model="llama-3.3-70b-versatile", apikey=os.environ["GROQAPIKEY"])

    messages = [

    ChatMessage(role="system", content="You are a coding expert assistant."),

    ChatMessage(role="user", content="Write a Python function for binary search."),

    ]

    response = llm.chat(messages)

    print(response.message.content)

    Building an Application: AI Chatbot with FastAPI

    Here is a complete example of building a chatbot API using Groq and FastAPI:

    from fastapi import FastAPI, HTTPException
    

    from fastapi.responses import StreamingResponse

    from pydantic import BaseModel

    from groq import Groq

    import json

    app = FastAPI(title="Groq Chatbot API")

    client = Groq()

    class ChatRequest(BaseModel):

    message: str

    model: str = "llama-3.3-70b-versatile"

    systemprompt: str = "You are a helpful AI assistant."

    stream: bool = False

    history: list[dict] = []

    @app.post("/chat")

    async def chat(request: ChatRequest):

    messages = [{"role": "system", "content": request.systemprompt}]

    messages.extend(request.history)

    messages.append({"role": "user", "content": request.message})

    if request.stream:

    async def generate():

    stream = client.chat.completions.create(

    messages=messages,

    model=request.model,

    stream=True,

    maxtokens=2048,

    )

    for chunk in stream:

    content = chunk.choices[0].delta.content

    if content:

    yield f"data: {json.dumps({'content': content})}\n\n"

    yield "data: [DONE]\n\n"

    return StreamingResponse(generate(), mediatype="text/event-stream")

    response = client.chat.completions.create(

    messages=messages,

    model=request.model,

    maxtokens=2048,

    )

    return {

    "response": response.choices[0].message.content,

    "usage": {

    "prompttokens": response.usage.prompttokens,

    "completiontokens": response.usage.completiontokens,

    "totaltokens": response.usage.totaltokens,

    }

    }

    @app.get("/models")

    async def listmodels():

    models = client.models.list()

    return {"models": [m.id for m in models.data]}

    Run the server:

    pip install fastapi uvicorn
    

    uvicorn main:app --reload --port 8000

    Best Practices

    1. Choose the Right Model

    Use models according to your needs to save costs and increase speed:

    • Simple tasks (classification, extraction, short QA): llama-3.1-8b-instant
    • Complex tasks (reasoning, analysis, writing): llama-3.3-70b-versatile
    • Multilingual tasks: mixtral-8x7b-32768
    • Image analysis: llama-3.2-90b-vision-preview

    2. Optimize Prompts

    # Bad - too long and ambiguous
    

    badprompt = "Please help me analyze this sales data and give useful insights about what can be improved..."

    Good - specific and structured

    goodprompt = """Analyze the following sales data:

    • Identify the top 3 best-performing products
    • Identify the 3 products with the biggest decline
    • Suggestions for improvement for each declining product

    Output format: JSON with keys topproducts, decliningproducts, recommendations"""

    3. Manage Rate Limits

    Groq has rate limits based on tiers. Management strategy:

    import time
    

    from collections import deque

    class RateLimiter:

    def init(self, maxrequestsperminute=30):

    self.maxrpm = maxrequestsperminute

    self.requests = deque()

    def waitifneeded(self):

    now = time.time()

    # Remove requests older than 1 minute

    while self.requests and self.requests[0] < now - 60:

    self.requests.popleft()

    if len(self.requests) >= self.maxrpm:

    sleeptime = 60 - (now - self.requests[0])

    if sleeptime > 0:

    print(f"Approaching rate limit. Waiting {sleeptime:.1f} seconds...")

    time.sleep(sleeptime)

    self.requests.append(time.time())

    limiter = RateLimiter(maxrequestsperminute=30)

    def safecompletion(messages, kwargs):

    limiter.waitifneeded()

    return client.chat.completions.create(messages=messages, kwargs)

    4. Caching for Efficiency

    import hashlib
    

    import json

    class SimpleCache:

    def init(self):

    self.cache = {}

    def makekey(self, messages, model, temperature):

    content = json.dumps({"messages": messages, "model": model,

    "temperature": temperature}, sortkeys=True)

    return hashlib.md5(content.encode()).hexdigest()

    def getorcreate(self, messages, model="llama-3.3-70b-versatile",

    temperature=0, kwargs):

    key = self.makekey(messages, model, temperature)

    if key in self.cache:

    print("Cache hit!")

    return self.cache[key]

    response = client.chat.completions.create(

    messages=messages,

    model=model,

    temperature=temperature,

    kwargs

    )

    result = response.choices[0].message.content

    self.cache[key] = result

    return result

    cache = SimpleCache()

    5. Proper Error Handling

    from groq import (
    

    Groq,

    APIError,

    AuthenticationError,

    RateLimitError,

    BadRequestError,

    )

    def safechat(messages, model="llama-3.3-70b-versatile"):

    try:

    response = client.chat.completions.create(

    messages=messages,

    model=model,

    max_tokens=1024,

    )

    return {

    "success": True,

    "content": response.choices[0].message.content,

    "usage": response.usage,

    }

    except AuthenticationError:

    return {"success": False, "error": "Invalid API key."}

    except RateLimitError:

    return {"success": False, "error": "Rate limit reached. Try again later."}

    except BadRequestError as e:

    return {"success": False, "error": f"Invalid request: {e}"}

    except APIError as e:

    return {"success": False, "error": f"API error: {e}"}

    Comparing Groq with Other Providers

    | Feature | Groq | OpenAI | Anthropic |

    |---------|------|--------|-----------|

    | Inference Speed | Very fast (LPU) | Standard (GPU) | Standard (GPU) |

    | Models | Open-source (Llama, Mixtral) | GPT-4, GPT-4o | Claude |

    | Pricing | Competitive | Premium | Premium |

    | Context Window | Up to 128K | Up to 128K | Up to 200K |

    | Function Calling | Yes | Yes | Yes |

    | Vision | Yes (Llama Vision) | Yes (GPT-4o) | Yes (Claude) |

    | Streaming | Yes | Yes | Yes |

    | JSON Mode | Yes | Yes | Yes |

    Conclusion

    Groq API offers a blazing-fast LLM inference solution with an easy-to-use API that is compatible with the OpenAI format. The main advantages of Groq are:

  • Speed: Inference using LPU is significantly faster than traditional GPUs
  • Compatibility: API format is compatible with OpenAI, making migration easy
  • Open-source Models: Access to popular models like Llama and Mixtral
  • Full Features: Supports streaming, function calling, vision, and JSON mode
  • Competitive Pricing: More affordable costs for the speed offered
  • With this tutorial, you now have a solid foundation for building AI applications using the Groq API. From basic usage to integration with popular frameworks and best practices for production.

    Next steps you can explore:

    • Implementing RAG (Retrieval-Augmented Generation) using Groq
    • Building multi-agent systems with Groq as the backbone
    • Cost optimization with model routing based on task complexity
    • Monitoring and observability for Groq applications in production

    Related Articles

    Fireworks AI: A Super Fast Inference Platform for Open LLMs and Multimodal Models

    Fireworks AI: Platform Inference Super Cepat buat Open LLM dan Model Multimodal Halo temen-temen, di tutorial kali ini a...

    Prompt Engineering Masterclass Tutorial: Modern Prompting Techniques

    Masterclass Prompt Engineering Daftar Isi Pendahuluan Prasyarat Memahami Dasar-Dasar Prompt Engineering [Ze...

    Complete Azure OpenAI Service Tutorial: GPT and LLMs on Azure

    Tutorial Lengkap Azure OpenAI Service: Enterprise AI dengan Model GPT Azure OpenAI Service menyediakan akses REST API ke...

    Complete LlamaIndex Tutorial: Building RAG Applications with LLMs

    Tutorial Lengkap LlamaIndex: Membangun Aplikasi RAG dengan LLM LlamaIndex adalah framework data yang powerful untuk memb...