Groq API Tutorial: Super Fast LLM Inference for AI Applications
Introduction
Groq has become one of the most popular AI inference platforms thanks to the remarkable speed offered by their proprietary Language Processing Unit (LPU) chip. If you have ever been frustrated by high latency when calling large language model APIs, Groq provides a solution with inference speeds that far surpass traditional GPUs.
In this tutorial, we will learn how to use the Groq API comprehensively, from initial setup, basic chat completion usage, to advanced features like streaming, tool use (function calling), vision models, and integration with popular frameworks like LangChain and LlamaIndex. All code examples in this tutorial can be run directly in your local environment.
Groq provides access to various popular open-source models including Llama, Mixtral, and Gemma with output speeds reaching hundreds of tokens per second. Interestingly, the Groq API uses a format compatible with the OpenAI API, making migration from OpenAI to Groq straightforward.
Installation and Setup
Getting an API Key
The first step is to register and obtain an API key from Groq:
Installing the Library
Groq provides an official Python SDK that can be installed via pip:
pip install groq
For more complete projects, install additional dependencies:
pip install groq python-dotenv httpx Pillow
Environment Configuration
Create a .env file to store your API key:
GROQAPIKEY=gskyourapikeyhere
Verify the installation with a simple script:
import os
from dotenv import loaddotenv
from groq import Groq
loaddotenv()
client = Groq(apikey=os.environ.get("GROQAPIKEY"))
Test connection
models = client.models.list()
for model in models.data:
print(f"Model: {model.id}")
If successful, you will see a list of available models on Groq.
Basic Usage
Simple Chat Completion
The most basic usage of the Groq API is chat completion:
from groq import Groq
client = Groq()
chatcompletion = client.chat.completions.create(
messages=[
{
"role": "system",
"content": "You are a helpful and friendly AI assistant."
},
{
"role": "user",
"content": "Explain what machine learning is in 3 sentences."
}
],
model="llama-3.3-70b-versatile",
temperature=0.7,
maxtokens=1024,
)
print(chatcompletion.choices[0].message.content)
Available Models
Groq provides several popular models. Here are usage recommendations:
# Model for general tasks and reasoning
MODELGENERAL = "llama-3.3-70b-versatile"
Fast model for simple tasks
MODELFAST = "llama-3.1-8b-instant"
Mixtral model for multilingual tasks
MODELMULTILINGUAL = "mixtral-8x7b-32768"
Model with large context window
MODELLONGCONTEXT = "llama-3.3-70b-versatile" # 128K context
Vision model for images
MODELVISION = "llama-3.2-90b-vision-preview"
Configuring Parameters
You can control model output with various parameters:
response = client.chat.completions.create(
messages=[
{"role": "user", "content": "Write a short poem about coding."}
],
model="llama-3.3-70b-versatile",
temperature=0.9, # Creativity (0.0 - 2.0)
maxtokens=512, # Maximum output tokens
topp=0.9, # Nucleus sampling
frequencypenalty=0.5, # Word repetition penalty
presencepenalty=0.3, # Penalty for already mentioned topics
stop=["---"], # Stop sequence
)
Multi-turn Conversation
For multi-turn conversations, send the entire conversation history:
conversationhistory = [
{"role": "system", "content": "You are a patient Python tutor."}
]
def chat(user
message):
conversationhistory.append({
"role": "user",
"content": usermessage
})
response = client.chat.completions.create(
messages=conversationhistory,
model="llama-3.3-70b-versatile",
temperature=0.7,
maxtokens=1024,
)
assistantmessage = response.choices[0].message.content
conversationhistory.append({
"role": "assistant",
"content": assistantmessage
})
return assistantmessage
Example conversation
print(chat("What is list comprehension in Python?"))
print(chat("Give an example with filtering."))
print(chat("How does its performance compare to a regular for loop?"))
Streaming Response
Basic Streaming
Streaming is very useful for better UX in chat applications:
stream = client.chat.completions.create(
messages=[
{"role": "user", "content": "Explain the Transformer architecture in detail."}
],
model="llama-3.3-70b-versatile",
temperature=0.7,
maxtokens=2048,
stream=True,
)
for chunk in stream:
content = chunk.choices[0].delta.content
if content:
print(content, end="", flush=True)
print()
Streaming with Metadata
You can collect metadata during streaming:
def streamwithstats(messages, model="llama-3.3-70b-versatile"):
stream = client.chat.completions.create(
messages=messages,
model=model,
stream=True,
max
tokens=2048,
)
fullresponse = ""
prompttokens = 0
completiontokens = 0
for chunk in stream:
if chunk.choices[0].delta.content:
content = chunk.choices[0].delta.content
fullresponse += content
print(content, end="", flush=True)
if hasattr(chunk, 'xgroq') and chunk.xgroq:
if hasattr(chunk.xgroq, 'usage') and chunk.xgroq.usage:
prompttokens = chunk.xgroq.usage.prompttokens
completiontokens = chunk.xgroq.usage.completiontokens
print(f"\n\nToken usage - Prompt: {prompttokens}, "
f"Completion: {completiontokens}")
return fullresponse
result = streamwithstats([
{"role": "user", "content": "What are the advantages of Groq over GPUs for inference?"}
])
Async Support
Using the Async Client
For applications requiring high concurrency:
import asyncio
from groq import AsyncGroq
async def asyncchat(prompt):
client = AsyncGroq()
response = await client.chat.completions.create(
messages=[{"role": "user", "content": prompt}],
model="llama-3.1-8b-instant",
maxtokens=512,
)
return response.choices[0].message.content
async def batchprocess(prompts):
tasks = [asyncchat(prompt) for prompt in prompts]
results = await asyncio.gather(*tasks)
return results
Process multiple prompts in parallel
prompts = [
"What is Docker?",
"What is Kubernetes?",
"What is CI/CD?",
"What is microservices?",
]
results = asyncio.run(batchprocess(prompts))
for prompt, result in zip(prompts, results):
print(f"Q: {prompt}")
print(f"A: {result[:100]}...")
print()
Async Streaming
async def asyncstream(prompt):
client = AsyncGroq()
stream = await client.chat.completions.create(
messages=[{"role": "user", "content": prompt}],
model="llama-3.3-70b-versatile",
stream=True,
max
tokens=1024,
)
fullresponse = ""
async for chunk in stream:
content = chunk.choices[0].delta.content
if content:
fullresponse += content
print(content, end="", flush=True)
print()
return fullresponse
asyncio.run(asyncstream("Explain the concept of RAG in AI."))
Tool Use (Function Calling)
Defining Tools
Groq supports function calling compatible with the OpenAI format:
import json
tools = [
{
"type": "function",
"function": {
"name": "getweather",
"description": "Get weather information for a location",
"parameters": {
"type": "object",
"properties": {
"location": {
"type": "string",
"description": "City name, e.g., New York, London"
},
"unit": {
"type": "string",
"enum": ["celsius", "fahrenheit"],
"description": "Temperature unit"
}
},
"required": ["location"]
}
}
},
{
"type": "function",
"function": {
"name": "searchproducts",
"description": "Search products by query",
"parameters": {
"type": "object",
"properties": {
"query": {
"type": "string",
"description": "Search keyword"
},
"maxprice": {
"type": "number",
"description": "Maximum price"
},
"category": {
"type": "string",
"description": "Product category"
}
},
"required": ["query"]
}
}
}
]
Using Tools in Chat
def getweather(location, unit="celsius"):
# Simulated weather data
weatherdata = {
"New York": {"temp": 22, "condition": "Partly cloudy"},
"London": {"temp": 16, "condition": "Light rain"},
"Tokyo": {"temp": 28, "condition": "Sunny"},
}
data = weatherdata.get(location, {"temp": 20, "condition": "Unknown"})
return json.dumps({
"location": location,
"temperature": data["temp"],
"unit": unit,
"condition": data["condition"]
})
def searchproducts(query, maxprice=None, category=None):
return json.dumps({
"products": [
{"name": f"{query} Premium", "price": 99.99, "stock": 25},
{"name": f"{query} Standard", "price": 49.99, "stock": 100},
]
})
Function mapping
availablefunctions = {
"getweather": getweather,
"searchproducts": searchproducts,
}
def chatwithtools(usermessage):
messages = [{"role": "user", "content": usermessage}]
response = client.chat.completions.create(
messages=messages,
model="llama-3.3-70b-versatile",
tools=tools,
toolchoice="auto",
maxtokens=1024,
)
responsemessage = response.choices[0].message
if responsemessage.toolcalls:
messages.append(responsemessage)
for toolcall in responsemessage.toolcalls:
functionname = toolcall.function.name
functionargs = json.loads(toolcall.function.arguments)
functionresponse = availablefunctionsfunctionname
messages.append({
"role": "tool",
"toolcallid": toolcall.id,
"name": functionname,
"content": functionresponse,
})
secondresponse = client.chat.completions.create(
messages=messages,
model="llama-3.3-70b-versatile",
maxtokens=1024,
)
return secondresponse.choices[0].message.content
return responsemessage.content
Test
print(chatwithtools("What's the weather like in New York today?"))
print(chatwithtools("Search for laptops under $1000"))
Vision (Multimodal)
Image Analysis
Groq supports vision models for image analysis:
import base64
import httpx
def encodeimagefromurl(imageurl):
response = httpx.get(imageurl)
return base64.b64encode(response.content).decode("utf-8")
def encodeimagefromfile(filepath):
with open(filepath, "rb") as f:
return base64.b64encode(f.read()).decode("utf-8")
Analyze image from URL
def analyzeimage(imageurl, question="Describe this image in detail."):
response = client.chat.completions.create(
messages=[
{
"role": "user",
"content": [
{"type": "text", "text": question},
{
"type": "imageurl",
"imageurl": {
"url": imageurl,
},
},
],
}
],
model="llama-3.2-90b-vision-preview",
maxtokens=1024,
)
return response.choices[0].message.content
Analyze local image
def analyzelocalimage(filepath, question="What do you see in this image?"):
base64image = encodeimagefromfile(filepath)
response = client.chat.completions.create(
messages=[
{
"role": "user",
"content": [
{"type": "text", "text": question},
{
"type": "imageurl",
"imageurl": {
"url": f"data:image/png;base64,{base64image}",
},
},
],
}
],
model="llama-3.2-90b-vision-preview",
maxtokens=1024,
)
return response.choices[0].message.content
JSON Mode
Structured Output
Use JSON mode to get structured output:
import json
def extractentities(text):
response = client.chat.completions.create(
messages=[
{
"role": "system",
"content": "Extract entities from the text and return in JSON format "
"with keys: persons (array), locations (array), "
"organizations (array), dates (array)."
},
{"role": "user", "content": text}
],
model="llama-3.3-70b-versatile",
temperature=0,
maxtokens=1024,
responseformat={"type": "jsonobject"},
)
return json.loads(response.choices[0].message.content)
text = """
On January 15, 2026, Google CEO Sundar Pichai announced the opening
of a new office in Singapore. The Minister of Communications welcomed
this investment.
"""
entities = extractentities(text)
print(json.dumps(entities, indent=2))
Structured Data Extraction
def analyzesentimentbatch(reviews):
response = client.chat.completions.create(
messages=[
{
"role": "system",
"content": (
"Analyze the sentiment of each review. "
"Return JSON with key 'results' containing an array of objects "
"with fields: reviewindex (int), sentiment (positive/negative/neutral), "
"confidence (0.0-1.0), keyphrases (array of string)."
)
},
{
"role": "user",
"content": json.dumps(reviews)
}
],
model="llama-3.3-70b-versatile",
temperature=0,
responseformat={"type": "jsonobject"},
maxtokens=2048,
)
return json.loads(response.choices[0].message.content)
reviews = [
"Great product, fast shipping!",
"Disappointing quality, doesn't match the description.",
"Average, you get what you pay for.",
]
results = analyzesentimentbatch(reviews)
print(json.dumps(results, indent=2))
Advanced Usage
Rate Limiting and Retry
Implementing robust retry logic:
import time
from groq import Groq, RateLimitError, APIError
def robustcompletion(messages, model="llama-3.3-70b-versatile",
maxretries=3, kwargs):
client = Groq()
for attempt in range(maxretries):
try:
response = client.chat.completions.create(
messages=messages,
model=model,
kwargs
)
return response
except RateLimitError as e:
if attempt < maxretries - 1:
waittime = (2 attempt) + 1
print(f"Rate limited. Waiting {waittime} seconds...")
time.sleep(waittime)
else:
raise
except APIError as e:
if attempt < maxretries - 1:
print(f"API error: {e}. Retry {attempt + 1}/{maxretries}")
time.sleep(1)
else:
raise
return None
Text Chunking for Long Documents
def processlongdocument(document, chunksize=4000, overlap=200):
chunks = []
start = 0
while start < len(document):
end = start + chunk
size
chunk = document[start:end]
chunks.append(chunk)
start = end - overlap
summaries = []
for i, chunk in enumerate(chunks):
response = client.chat.completions.create(
messages=[
{
"role": "system",
"content": "Summarize the following document section concisely."
},
{"role": "user", "content": chunk}
],
model="llama-3.1-8b-instant",
temperature=0.3,
maxtokens=512,
)
summaries.append(response.choices[0].message.content)
print(f"Chunk {i+1}/{len(chunks)} processed.")
# Combine all summaries
combined = "\n\n".join(summaries)
finalresponse = client.chat.completions.create(
messages=[
{
"role": "system",
"content": "Combine the following summaries into a single "
"coherent and comprehensive summary."
},
{"role": "user", "content": combined}
],
model="llama-3.3-70b-versatile",
temperature=0.3,
maxtokens=1024,
)
return finalresponse.choices[0].message.content
Integration with LangChain
from langchaingroq import ChatGroq
from langchain
core.messages import HumanMessage, SystemMessage
from langchaincore.outputparsers import StrOutputParser
from langchaincore.prompts import ChatPromptTemplate
Initialize model
llm = ChatGroq(
model="llama-3.3-70b-versatile",
temperature=0.7,
maxtokens=1024,
)
Simple chain
prompt = ChatPromptTemplate.frommessages([
("system", "You are a {topic} expert who explains concepts simply."),
("human", "{question}")
])
chain = prompt | llm | StrOutputParser()
result = chain.invoke({
"topic": "machine learning",
"question": "What's the difference between supervised and unsupervised learning?"
})
print(result)
Integration with LlamaIndex
from llamaindex.llms.groq import Groq as GroqLLM
from llamaindex.core.llms import ChatMessage
llm = GroqLLM(model="llama-3.3-70b-versatile", apikey=os.environ["GROQAPIKEY"])
messages = [
ChatMessage(role="system", content="You are a coding expert assistant."),
ChatMessage(role="user", content="Write a Python function for binary search."),
]
response = llm.chat(messages)
print(response.message.content)
Building an Application: AI Chatbot with FastAPI
Here is a complete example of building a chatbot API using Groq and FastAPI:
from fastapi import FastAPI, HTTPException
from fastapi.responses import StreamingResponse
from pydantic import BaseModel
from groq import Groq
import json
app = FastAPI(title="Groq Chatbot API")
client = Groq()
class ChatRequest(BaseModel):
message: str
model: str = "llama-3.3-70b-versatile"
systemprompt: str = "You are a helpful AI assistant."
stream: bool = False
history: list[dict] = []
@app.post("/chat")
async def chat(request: ChatRequest):
messages = [{"role": "system", "content": request.systemprompt}]
messages.extend(request.history)
messages.append({"role": "user", "content": request.message})
if request.stream:
async def generate():
stream = client.chat.completions.create(
messages=messages,
model=request.model,
stream=True,
maxtokens=2048,
)
for chunk in stream:
content = chunk.choices[0].delta.content
if content:
yield f"data: {json.dumps({'content': content})}\n\n"
yield "data: [DONE]\n\n"
return StreamingResponse(generate(), mediatype="text/event-stream")
response = client.chat.completions.create(
messages=messages,
model=request.model,
maxtokens=2048,
)
return {
"response": response.choices[0].message.content,
"usage": {
"prompttokens": response.usage.prompttokens,
"completiontokens": response.usage.completiontokens,
"totaltokens": response.usage.totaltokens,
}
}
@app.get("/models")
async def listmodels():
models = client.models.list()
return {"models": [m.id for m in models.data]}
Run the server:
pip install fastapi uvicorn
uvicorn main:app --reload --port 8000
Best Practices
1. Choose the Right Model
Use models according to your needs to save costs and increase speed:
- Simple tasks (classification, extraction, short QA):
llama-3.1-8b-instant - Complex tasks (reasoning, analysis, writing):
llama-3.3-70b-versatile - Multilingual tasks:
mixtral-8x7b-32768 - Image analysis:
llama-3.2-90b-vision-preview
2. Optimize Prompts
# Bad - too long and ambiguous
badprompt = "Please help me analyze this sales data and give useful insights about what can be improved..."
Good - specific and structured
goodprompt = """Analyze the following sales data:
- Identify the top 3 best-performing products
- Identify the 3 products with the biggest decline
- Suggestions for improvement for each declining product
Output format: JSON with keys topproducts, decliningproducts, recommendations"""
3. Manage Rate Limits
Groq has rate limits based on tiers. Management strategy:
import time
from collections import deque
class RateLimiter:
def init(self, maxrequestsperminute=30):
self.maxrpm = maxrequestsperminute
self.requests = deque()
def waitifneeded(self):
now = time.time()
# Remove requests older than 1 minute
while self.requests and self.requests[0] < now - 60:
self.requests.popleft()
if len(self.requests) >= self.maxrpm:
sleeptime = 60 - (now - self.requests[0])
if sleeptime > 0:
print(f"Approaching rate limit. Waiting {sleeptime:.1f} seconds...")
time.sleep(sleeptime)
self.requests.append(time.time())
limiter = RateLimiter(maxrequestsperminute=30)
def safecompletion(messages, kwargs):
limiter.waitifneeded()
return client.chat.completions.create(messages=messages, kwargs)
4. Caching for Efficiency
import hashlib
import json
class SimpleCache:
def init(self):
self.cache = {}
def makekey(self, messages, model, temperature):
content = json.dumps({"messages": messages, "model": model,
"temperature": temperature}, sortkeys=True)
return hashlib.md5(content.encode()).hexdigest()
def getorcreate(self, messages, model="llama-3.3-70b-versatile",
temperature=0, kwargs):
key = self.makekey(messages, model, temperature)
if key in self.cache:
print("Cache hit!")
return self.cache[key]
response = client.chat.completions.create(
messages=messages,
model=model,
temperature=temperature,
kwargs
)
result = response.choices[0].message.content
self.cache[key] = result
return result
cache = SimpleCache()
5. Proper Error Handling
from groq import (
Groq,
APIError,
AuthenticationError,
RateLimitError,
BadRequestError,
)
def safechat(messages, model="llama-3.3-70b-versatile"):
try:
response = client.chat.completions.create(
messages=messages,
model=model,
max_tokens=1024,
)
return {
"success": True,
"content": response.choices[0].message.content,
"usage": response.usage,
}
except AuthenticationError:
return {"success": False, "error": "Invalid API key."}
except RateLimitError:
return {"success": False, "error": "Rate limit reached. Try again later."}
except BadRequestError as e:
return {"success": False, "error": f"Invalid request: {e}"}
except APIError as e:
return {"success": False, "error": f"API error: {e}"}
Comparing Groq with Other Providers
| Feature | Groq | OpenAI | Anthropic |
|---------|------|--------|-----------|
| Inference Speed | Very fast (LPU) | Standard (GPU) | Standard (GPU) |
| Models | Open-source (Llama, Mixtral) | GPT-4, GPT-4o | Claude |
| Pricing | Competitive | Premium | Premium |
| Context Window | Up to 128K | Up to 128K | Up to 200K |
| Function Calling | Yes | Yes | Yes |
| Vision | Yes (Llama Vision) | Yes (GPT-4o) | Yes (Claude) |
| Streaming | Yes | Yes | Yes |
| JSON Mode | Yes | Yes | Yes |
Conclusion
Groq API offers a blazing-fast LLM inference solution with an easy-to-use API that is compatible with the OpenAI format. The main advantages of Groq are:
With this tutorial, you now have a solid foundation for building AI applications using the Groq API. From basic usage to integration with popular frameworks and best practices for production.
Next steps you can explore:
- Implementing RAG (Retrieval-Augmented Generation) using Groq
- Building multi-agent systems with Groq as the backbone
- Cost optimization with model routing based on task complexity
- Monitoring and observability for Groq applications in production