Marker Tutorial: Converting PDFs and Documents to Markdown for RAG Pipelines
In the world of AI and data engineering, extracting text from PDF documents, EPUBs, and images into structured formats is a fundamental requirement. Marker is an open-source library by VikParuchuri that converts documents into high-quality Markdown with remarkable accuracy, including recognition of tables, mathematical formulas, and complex document layouts.
This tutorial covers how to use Marker comprehensively, from installation to integrating it into production RAG (Retrieval-Augmented Generation) pipelines.
Why Marker?
Before Marker, developers typically relied on PyPDF2, pdfminer, or Tesseract OCR for text extraction from PDFs. However, these tools often produce messy output, lose table formatting, and fail to recognize multi-column layouts.
Marker solves these problems with a different approach. The library uses a combination of deep learning models to:
- Detect page layout (headers, footers, columns, tables)
- Recognize text with Surya-based OCR
- Convert tables into clean Markdown format
- Process mathematical formulas into LaTeX
- Handle multilingual documents
Marker supports PDF, EPUB, MOBI, and various image formats as input. The output is clean Markdown ready for indexing, RAG pipelines, or further analysis.
Installation
System Prerequisites
Marker requires Python 3.9 or later. For optimal performance, a GPU with CUDA is highly recommended, although CPU is also supported.
python --version
Make sure your Python version is at least 3.9.
Installation via pip
The easiest way to install Marker is using pip:
pip install marker-pdf
This will install Marker along with all required dependencies, including Surya models for OCR and layout detection.
Installation from Source
If you want the latest version or wish to contribute to development:
git clone https://github.com/VikParuchuri/marker.git
cd marker
pip install -e .
Installation with GPU Support
To enable GPU acceleration with CUDA:
pip install marker-pdf[gpu]
Make sure the CUDA toolkit is installed on your system. Verify with:
python -c "import torch; print(torch.cuda.isavailable())"
Model Download
When run for the first time, Marker will automatically download the required models from Hugging Face. These models include:
- Layout detection model
- OCR model (Surya)
- Table recognition model
- Mathematical formula conversion model
Total model size is approximately 2-3 GB. Ensure a stable internet connection during the first run.
Basic Usage
Single File Conversion via CLI
Marker provides an easy-to-use command-line interface:
markersingle /path/to/document.pdf /path/to/output/folder
This command generates a Markdown file in the output folder along with images extracted from the document.
Batch Conversion
To convert multiple documents at once:
marker /path/to/input/folder /path/to/output/folder
Marker will process all PDF, EPUB, and MOBI files in the input folder in parallel.
Using the Python API
For integration into Python applications, use the API directly:
from marker.converters.pdf import PdfConverter
from marker.models import createmodeldict
from marker.output import textfromrendered
Initialize converter with models
converter = PdfConverter(
artifactdict=createmodeldict(),
)
Convert document
rendered = converter("document.pdf")
Get Markdown text and metadata
text, metadata, images = textfromrendered(rendered)
print(text) # Markdown output
print(metadata) # Document metadata
Example Output
For instance, if you have a financial report PDF with tables, Marker will produce output like:
# Financial Report Q1 2026
Revenue Summary
| Category | Q1 2025 | Q1 2026 | Growth |
|----------|---------|---------|--------|
| Product A | $500K | $750K | 50% |
| Product B | $300K | $420K | 40% |
| Total | $800K | $1.17M | 46% |
Company revenue increased by 46% year-over-year...
Notice how the table is neatly converted to Markdown format with proper alignment.
Advanced Configuration
Setting Conversion Parameters
Marker provides various parameters to control the conversion process:
from marker.converters.pdf import PdfConverter
from marker.models import createmodeldict
converter = PdfConverter(
artifactdict=createmodeldict(),
)
Configuration through settings
import marker.settings as settings
Enable debug mode for detailed processing info
settings.DEBUG = True
Set number of workers for batch processing
settings.WORKERNUM = 4
Set device (cuda or cpu)
settings.TORCHDEVICE = "cuda"
Processing Specific Pages
If you only need certain pages from a document:
from marker.converters.pdf import PdfConverter
from marker.models import createmodeldict
from marker.output import textfromrendered
converter = PdfConverter(
artifactdict=createmodeldict(),
)
Convert only pages 1-5
rendered = converter("document.pdf", pagerange=[0, 5])
text, metadata, images = textfromrendered(rendered)
Handling OCR for Scanned Documents
For scanned documents without a text layer:
from marker.converters.pdf import PdfConverter
from marker.models import createmodeldict
from marker.output import textfromrendered
import marker.settings as settings
Force OCR on all pages
settings.OCRALLPAGES = True
converter = PdfConverter(
artifactdict=createmodeldict(),
)
rendered = converter("scanneddocument.pdf")
text, metadata, images = textfromrendered(rendered)
OCR Language Configuration
Marker supports multilingual OCR through Surya:
import marker.settings as settings
Set supported OCR languages
settings.OCRLANGUAGES = ["en", "id"] # English and Indonesian
Advanced Usage
Integration with RAG Pipelines
One of Marker's primary use cases is preparing documents for RAG pipelines. Here's an example integration with LangChain:
import os
from marker.converters.pdf import PdfConverter
from marker.models import createmodeldict
from marker.output import textfromrendered
from langchain.textsplitter import RecursiveCharacterTextSplitter
from langchaincommunity.vectorstores import Chroma
from langchaincommunity.embeddings import HuggingFaceEmbeddings
def processdocuments(pdffolder: str) -> list:
"""Convert all PDFs to Markdown using Marker."""
converter = PdfConverter(
artifactdict=createmodeldict(),
)
documents = []
for filename in os.listdir(pdffolder):
if filename.endswith(".pdf"):
filepath = os.path.join(pdffolder, filename)
rendered = converter(filepath)
text, metadata, images = textfromrendered(rendered)
documents.append({
"content": text,
"metadata": {
"source": filename,
*metadata
}
})
return documents
def buildragindex(documents: list):
"""Build RAG index from converted documents."""
textsplitter = RecursiveCharacterTextSplitter(
chunksize=1000,
chunkoverlap=200,
separators=["\n## ", "\n### ", "\n\n", "\n", " "]
)
chunks = []
for doc in documents:
splits = textsplitter.splittext(doc["content"])
for split in splits:
chunks.append({
"text": split,
"metadata": doc["metadata"]
})
embeddings = HuggingFaceEmbeddings(
modelname="sentence-transformers/all-MiniLM-L6-v2"
)
texts = [c["text"] for c in chunks]
metadatas = [c["metadata"] for c in chunks]
vectorstore = Chroma.fromtexts(
texts=texts,
metadatas=metadatas,
embedding=embeddings,
persistdirectory="./chromadb"
)
return vectorstore
Usage
pdffolder = "./documents"
docs = processdocuments(pdffolder)
vectorstore = buildragindex(docs)
Query
results = vectorstore.similaritysearch("Q1 2026 revenue", k=3)
for r in results:
print(r.pagecontent)
print(r.metadata)
print("---")
Batch Processing with Monitoring
For processing large numbers of documents with progress monitoring:
import os
import json
import time
from marker.converters.pdf import PdfConverter
from marker.models import createmodeldict
from marker.output import textfromrendered
def batchconvertwithmonitoring(
inputfolder: str,
outputfolder: str,
logfile: str = "conversionlog.json"
):
"""Batch convert with logging and error handling."""
os.makedirs(outputfolder, existok=True)
converter = PdfConverter(
artifactdict=createmodeldict(),
)
pdffiles = [f for f in os.listdir(inputfolder) if f.endswith(".pdf")]
total = len(pdffiles)
results = []
print(f"Processing {total} documents...")
for i, filename in enumerate(pdffiles, 1):
filepath = os.path.join(inputfolder, filename)
starttime = time.time()
try:
rendered = converter(filepath)
text, metadata, images = textfromrendered(rendered)
# Save Markdown output
outputname = filename.replace(".pdf", ".md")
outputpath = os.path.join(outputfolder, outputname)
with open(outputpath, "w", encoding="utf-8") as f:
f.write(text)
# Save images if any
if images:
imgfolder = os.path.join(
outputfolder,
filename.replace(".pdf", "images")
)
os.makedirs(imgfolder, existok=True)
for imgname, imgdata in images.items():
imgpath = os.path.join(imgfolder, imgname)
with open(imgpath, "wb") as f:
f.write(imgdata)
elapsed = time.time() - starttime
result = {
"file": filename,
"status": "success",
"output": outputpath,
"pages": metadata.get("pages", 0),
"timeseconds": round(elapsed, 2),
"imagesextracted": len(images) if images else 0
}
print(f"[{i}/{total}] OK: {filename} ({elapsed:.1f}s)")
except Exception as e:
elapsed = time.time() - starttime
result = {
"file": filename,
"status": "error",
"error": str(e),
"timeseconds": round(elapsed, 2)
}
print(f"[{i}/{total}] ERROR: {filename} - {e}")
results.append(result)
# Save log
with open(logfile, "w") as f:
json.dump(results, f, indent=2)
success = sum(1 for r in results if r["status"] == "success")
print(f"\nComplete: {success}/{total} successfully converted")
print(f"Log saved to: {logfile}")
return results
Run
results = batchconvertwithmonitoring(
inputfolder="./pdfdocuments",
outputfolder="./markdownoutput"
)
Building a REST API with FastAPI
To serve Marker as a microservice:
import os
import tempfile
from fastapi import FastAPI, UploadFile, File, HTTPException
from fastapi.responses import JSONResponse
from marker.converters.pdf import PdfConverter
from marker.models import createmodeldict
from marker.output import textfromrendered
app = FastAPI(title="Marker PDF Converter API")
converter = None
@app.onevent("startup")
async def loadmodels():
global converter
converter = PdfConverter(
artifactdict=createmodeldict(),
)
@app.post("/convert")
async def convertpdf(file: UploadFile = File(...)):
"""Convert PDF to Markdown."""
if not file.filename.endswith((".pdf", ".epub", ".mobi")):
raise HTTPException(
statuscode=400,
detail="Unsupported file format. Use PDF, EPUB, or MOBI."
)
with tempfile.NamedTemporaryFile(
suffix=os.path.splitext(file.filename)[1],
delete=False
) as tmp:
content = await file.read()
tmp.write(content)
tmppath = tmp.name
try:
rendered = converter(tmppath)
text, metadata, images = textfromrendered(rendered)
return JSONResponse({
"filename": file.filename,
"markdown": text,
"metadata": metadata,
"imagescount": len(images) if images else 0
})
except Exception as e:
raise HTTPException(statuscode=500, detail=str(e))
finally:
os.unlink(tmppath)
@app.get("/health")
async def health():
return {"status": "ok", "modelloaded": converter is not None}
Run the server:
pip install fastapi uvicorn python-multipart
uvicorn markerapi:app --host 0.0.0.0 --port 8000
Test with curl:
curl -X POST http://localhost:8000/convert \
-F "file=@document.pdf" \
| python -m json.tool
Quality Comparison with Other Tools
Here's a comparison of Marker with other text extraction tools on documents with complex tables:
import time
def compareextraction(pdfpath: str):
"""Compare extraction results across different tools."""
# Marker
from marker.converters.pdf import PdfConverter
from marker.models import createmodeldict
from marker.output import textfromrendered
start = time.time()
converter = PdfConverter(artifactdict=createmodeldict())
rendered = converter(pdfpath)
markertext, , = textfromrendered(rendered)
markertime = time.time() - start
# PyPDF2
from PyPDF2 import PdfReader
start = time.time()
reader = PdfReader(pdfpath)
pypdftext = ""
for page in reader.pages:
pypdftext += page.extracttext() or ""
pypdftime = time.time() - start
# pdfminer
from pdfminer.highlevel import extracttext
start = time.time()
pdfminertext = extracttext(pdfpath)
pdfminertime = time.time() - start
print("=== Comparison Results ===")
print(f"Marker : {len(markertext)} chars, {markertime:.2f}s")
print(f"PyPDF2 : {len(pypdftext)} chars, {pypdftime:.2f}s")
print(f"pdfminer : {len(pdfminertext)} chars, {pdfminertime:.2f}s")
print()
print("--- Marker Output (first 500 chars) ---")
print(markertext[:500])
print()
print("--- PyPDF2 Output (first 500 chars) ---")
print(pypdftext[:500])
compareextraction("complexreport.pdf")
Best Practices
1. Memory Optimization for Large Documents
When processing large documents (hundreds of pages), pay attention to memory usage:
import gc
import torch
def processlargedocument(pdfpath: str, batchsize: int = 20):
"""Process large documents in page batches."""
from marker.converters.pdf import PdfConverter
from marker.models import createmodeldict
from marker.output import textfromrendered
import fitz # PyMuPDF
doc = fitz.open(pdfpath)
totalpages = len(doc)
doc.close()
alltext = []
converter = PdfConverter(artifactdict=createmodeldict())
for start in range(0, totalpages, batchsize):
end = min(start + batchsize, totalpages)
rendered = converter(pdfpath, pagerange=[start, end])
text, , = textfromrendered(rendered)
alltext.append(text)
gc.collect()
if torch.cuda.isavailable():
torch.cuda.emptycache()
return "\n\n".join(alltext)
2. Caching Conversion Results
Avoid re-converting the same documents:
import hashlib
import json
import os
CACHEDIR = "./markercache"
def getfilehash(filepath: str) -> str:
"""Calculate SHA256 hash of a file."""
sha256 = hashlib.sha256()
with open(filepath, "rb") as f:
for chunk in iter(lambda: f.read(8192), b""):
sha256.update(chunk)
return sha256.hexdigest()
def convertwithcache(pdfpath: str) -> str:
"""Convert with file-hash-based caching."""
os.makedirs(CACHEDIR, existok=True)
filehash = getfilehash(pdfpath)
cachepath = os.path.join(CACHEDIR, f"{filehash}.json")
if os.path.exists(cachepath):
with open(cachepath, "r") as f:
cached = json.load(f)
print(f"Cache hit: {pdfpath}")
return cached["text"]
from marker.converters.pdf import PdfConverter
from marker.models import createmodeldict
from marker.output import textfromrendered
converter = PdfConverter(artifactdict=createmodeldict())
rendered = converter(pdfpath)
text, metadata, = textfromrendered(rendered)
with open(cachepath, "w") as f:
json.dump({"text": text, "metadata": metadata}, f)
print(f"Converted and cached: {pdfpath}")
return text
3. Output Validation
Always validate output quality before using it:
def validateconversion(text: str, minlength: int = 100) -> dict:
"""Validate conversion output quality."""
issues = []
if len(text) < minlength:
issues.append(f"Text too short ({len(text)} chars)")
lines = text.split("\n")
emptyratio = sum(1 for l in lines if not l.strip()) / max(len(lines), 1)
if emptyratio > 0.7:
issues.append(f"Too many empty lines ({emptyratio:.0%})")
garbled = sum(1 for c in text if ord(c) > 65535) / max(len(text), 1)
if garbled > 0.05:
issues.append(f"Possible garbled characters ({garbled:.1%})")
tablecount = text.count("|---|")
headingcount = text.count("\n#")
return {
"valid": len(issues) == 0,
"issues": issues,
"stats": {
"length": len(text),
"lines": len(lines),
"tables": tablecount,
"headings": headingcount,
"emptylineratio": round(emptyratio, 2)
}
}
4. Proper Error Handling
from pathlib import Path
def safeconvert(pdfpath: str, fallbackocr: bool = True) -> dict:
"""Convert with comprehensive error handling."""
path = Path(pdfpath)
if not path.exists():
return {"status": "error", "error": f"File not found: {pdfpath}"}
if not path.suffix.lower() in [".pdf", ".epub", ".mobi"]:
return {"status": "error", "error": f"Unsupported format: {path.suffix}"}
filesizemb = path.stat().stsize / (1024 1024)
if filesizemb > 500:
return {"status": "error", "error": f"File too large: {filesizemb:.0f}MB (max 500MB)"}
try:
from marker.converters.pdf import PdfConverter
from marker.models import createmodeldict
from marker.output import textfromrendered
import marker.settings as settings
if fallbackocr:
settings.OCRALLPAGES = True
converter = PdfConverter(artifactdict=createmodeldict())
rendered = converter(pdfpath)
text, metadata, images = textfromrendered(rendered)
validation = validateconversion(text)
return {
"status": "success",
"text": text,
"metadata": metadata,
"imagescount": len(images) if images else 0,
"validation": validation
}
except Exception as e:
return {"status": "error", "error": str(e)}
5. Performance Tips
Here are some tips to optimize Marker's performance:
- Use GPU: Conversion is 3-5x faster with a CUDA GPU
- Batch processing: Use the
markerCLI for multiple files instead of loopingmarkersingle - Page range: If you only need specific pages, use the
pagerangeparameter - SSD storage: Marker models are quite large; SSD speeds up model loading
- Caching: Implement caching for frequently accessed documents
Conclusion
Marker is an extremely powerful solution for converting PDF, EPUB, and MOBI documents into high-quality Markdown. Its key advantages are:
- High accuracy in recognizing layouts, tables, and mathematical formulas
- Multi-format support for PDF, EPUB, MOBI, and images
- Multilingual with Surya OCR support
- Open-source and actively maintained
- Easy to integrate into RAG pipelines and Python applications
With Marker, you can build reliable document processing pipelines for various needs, from corporate knowledge bases and document search systems to RAG-based chatbots that use internal documents as knowledge sources.
For more information, visit the official Marker repository on GitHub and the Surya OCR documentation for additional language configurations.