Marker Tutorial: Converting PDFs and Documents to Markdown for RAG Pipelines

# Tutorial Marker: Mengubah PDF dan Dokumen Menjadi Markdown untuk Pipeline RAG Dalam dunia AI dan data engineering, kemampuan mengekstrak teks dari dokumen PDF, EPUB, dan gambar menjadi format terst...

By Ruby Abdullah · · tutorial
MarkerPDFOCRRAGDocument Processing

Marker Tutorial: Converting PDFs and Documents to Markdown for RAG Pipelines

In the world of AI and data engineering, extracting text from PDF documents, EPUBs, and images into structured formats is a fundamental requirement. Marker is an open-source library by VikParuchuri that converts documents into high-quality Markdown with remarkable accuracy, including recognition of tables, mathematical formulas, and complex document layouts.

This tutorial covers how to use Marker comprehensively, from installation to integrating it into production RAG (Retrieval-Augmented Generation) pipelines.

Why Marker?

Before Marker, developers typically relied on PyPDF2, pdfminer, or Tesseract OCR for text extraction from PDFs. However, these tools often produce messy output, lose table formatting, and fail to recognize multi-column layouts.

Marker solves these problems with a different approach. The library uses a combination of deep learning models to:

  • Detect page layout (headers, footers, columns, tables)
  • Recognize text with Surya-based OCR
  • Convert tables into clean Markdown format
  • Process mathematical formulas into LaTeX
  • Handle multilingual documents

Marker supports PDF, EPUB, MOBI, and various image formats as input. The output is clean Markdown ready for indexing, RAG pipelines, or further analysis.

Installation

System Prerequisites

Marker requires Python 3.9 or later. For optimal performance, a GPU with CUDA is highly recommended, although CPU is also supported.

python --version

Make sure your Python version is at least 3.9.

Installation via pip

The easiest way to install Marker is using pip:

pip install marker-pdf

This will install Marker along with all required dependencies, including Surya models for OCR and layout detection.

Installation from Source

If you want the latest version or wish to contribute to development:

git clone https://github.com/VikParuchuri/marker.git

cd marker

pip install -e .

Installation with GPU Support

To enable GPU acceleration with CUDA:

pip install marker-pdf[gpu]

Make sure the CUDA toolkit is installed on your system. Verify with:

python -c "import torch; print(torch.cuda.isavailable())"

Model Download

When run for the first time, Marker will automatically download the required models from Hugging Face. These models include:

  • Layout detection model
  • OCR model (Surya)
  • Table recognition model
  • Mathematical formula conversion model

Total model size is approximately 2-3 GB. Ensure a stable internet connection during the first run.

Basic Usage

Single File Conversion via CLI

Marker provides an easy-to-use command-line interface:

markersingle /path/to/document.pdf /path/to/output/folder

This command generates a Markdown file in the output folder along with images extracted from the document.

Batch Conversion

To convert multiple documents at once:

marker /path/to/input/folder /path/to/output/folder

Marker will process all PDF, EPUB, and MOBI files in the input folder in parallel.

Using the Python API

For integration into Python applications, use the API directly:

from marker.converters.pdf import PdfConverter

from marker.models import createmodeldict

from marker.output import textfromrendered

Initialize converter with models

converter = PdfConverter(

artifactdict=createmodeldict(),

)

Convert document

rendered = converter("document.pdf")

Get Markdown text and metadata

text, metadata, images = textfromrendered(rendered)

print(text) # Markdown output

print(metadata) # Document metadata

Example Output

For instance, if you have a financial report PDF with tables, Marker will produce output like:

# Financial Report Q1 2026

Revenue Summary

| Category | Q1 2025 | Q1 2026 | Growth |

|----------|---------|---------|--------|

| Product A | $500K | $750K | 50% |

| Product B | $300K | $420K | 40% |

| Total | $800K | $1.17M | 46% |

Company revenue increased by 46% year-over-year...

Notice how the table is neatly converted to Markdown format with proper alignment.

Advanced Configuration

Setting Conversion Parameters

Marker provides various parameters to control the conversion process:

from marker.converters.pdf import PdfConverter

from marker.models import createmodeldict

converter = PdfConverter(

artifactdict=createmodeldict(),

)

Configuration through settings

import marker.settings as settings

Enable debug mode for detailed processing info

settings.DEBUG = True

Set number of workers for batch processing

settings.WORKERNUM = 4

Set device (cuda or cpu)

settings.TORCHDEVICE = "cuda"

Processing Specific Pages

If you only need certain pages from a document:

from marker.converters.pdf import PdfConverter

from marker.models import createmodeldict

from marker.output import textfromrendered

converter = PdfConverter(

artifactdict=createmodeldict(),

)

Convert only pages 1-5

rendered = converter("document.pdf", pagerange=[0, 5])

text, metadata, images = textfromrendered(rendered)

Handling OCR for Scanned Documents

For scanned documents without a text layer:

from marker.converters.pdf import PdfConverter

from marker.models import createmodeldict

from marker.output import textfromrendered

import marker.settings as settings

Force OCR on all pages

settings.OCRALLPAGES = True

converter = PdfConverter(

artifactdict=createmodeldict(),

)

rendered = converter("scanneddocument.pdf")

text, metadata, images = textfromrendered(rendered)

OCR Language Configuration

Marker supports multilingual OCR through Surya:

import marker.settings as settings

Set supported OCR languages

settings.OCRLANGUAGES = ["en", "id"] # English and Indonesian

Advanced Usage

Integration with RAG Pipelines

One of Marker's primary use cases is preparing documents for RAG pipelines. Here's an example integration with LangChain:

import os

from marker.converters.pdf import PdfConverter

from marker.models import createmodeldict

from marker.output import textfromrendered

from langchain.textsplitter import RecursiveCharacterTextSplitter

from langchaincommunity.vectorstores import Chroma

from langchaincommunity.embeddings import HuggingFaceEmbeddings

def processdocuments(pdffolder: str) -> list:

"""Convert all PDFs to Markdown using Marker."""

converter = PdfConverter(

artifactdict=createmodeldict(),

)

documents = []

for filename in os.listdir(pdffolder):

if filename.endswith(".pdf"):

filepath = os.path.join(pdffolder, filename)

rendered = converter(filepath)

text, metadata, images = textfromrendered(rendered)

documents.append({

"content": text,

"metadata": {

"source": filename,

*metadata

}

})

return documents

def buildragindex(documents: list):

"""Build RAG index from converted documents."""

textsplitter = RecursiveCharacterTextSplitter(

chunksize=1000,

chunkoverlap=200,

separators=["\n## ", "\n### ", "\n\n", "\n", " "]

)

chunks = []

for doc in documents:

splits = textsplitter.splittext(doc["content"])

for split in splits:

chunks.append({

"text": split,

"metadata": doc["metadata"]

})

embeddings = HuggingFaceEmbeddings(

modelname="sentence-transformers/all-MiniLM-L6-v2"

)

texts = [c["text"] for c in chunks]

metadatas = [c["metadata"] for c in chunks]

vectorstore = Chroma.fromtexts(

texts=texts,

metadatas=metadatas,

embedding=embeddings,

persistdirectory="./chromadb"

)

return vectorstore

Usage

pdffolder = "./documents"

docs = processdocuments(pdffolder)

vectorstore = buildragindex(docs)

Query

results = vectorstore.similaritysearch("Q1 2026 revenue", k=3)

for r in results:

print(r.pagecontent)

print(r.metadata)

print("---")

Batch Processing with Monitoring

For processing large numbers of documents with progress monitoring:

import os

import json

import time

from marker.converters.pdf import PdfConverter

from marker.models import createmodeldict

from marker.output import textfromrendered

def batchconvertwithmonitoring(

inputfolder: str,

outputfolder: str,

logfile: str = "conversionlog.json"

):

"""Batch convert with logging and error handling."""

os.makedirs(outputfolder, existok=True)

converter = PdfConverter(

artifactdict=createmodeldict(),

)

pdffiles = [f for f in os.listdir(inputfolder) if f.endswith(".pdf")]

total = len(pdffiles)

results = []

print(f"Processing {total} documents...")

for i, filename in enumerate(pdffiles, 1):

filepath = os.path.join(inputfolder, filename)

starttime = time.time()

try:

rendered = converter(filepath)

text, metadata, images = textfromrendered(rendered)

# Save Markdown output

outputname = filename.replace(".pdf", ".md")

outputpath = os.path.join(outputfolder, outputname)

with open(outputpath, "w", encoding="utf-8") as f:

f.write(text)

# Save images if any

if images:

imgfolder = os.path.join(

outputfolder,

filename.replace(".pdf", "images")

)

os.makedirs(imgfolder, existok=True)

for imgname, imgdata in images.items():

imgpath = os.path.join(imgfolder, imgname)

with open(imgpath, "wb") as f:

f.write(imgdata)

elapsed = time.time() - starttime

result = {

"file": filename,

"status": "success",

"output": outputpath,

"pages": metadata.get("pages", 0),

"timeseconds": round(elapsed, 2),

"imagesextracted": len(images) if images else 0

}

print(f"[{i}/{total}] OK: {filename} ({elapsed:.1f}s)")

except Exception as e:

elapsed = time.time() - starttime

result = {

"file": filename,

"status": "error",

"error": str(e),

"timeseconds": round(elapsed, 2)

}

print(f"[{i}/{total}] ERROR: {filename} - {e}")

results.append(result)

# Save log

with open(logfile, "w") as f:

json.dump(results, f, indent=2)

success = sum(1 for r in results if r["status"] == "success")

print(f"\nComplete: {success}/{total} successfully converted")

print(f"Log saved to: {logfile}")

return results

Run

results = batchconvertwithmonitoring(

inputfolder="./pdfdocuments",

outputfolder="./markdownoutput"

)

Building a REST API with FastAPI

To serve Marker as a microservice:

import os

import tempfile

from fastapi import FastAPI, UploadFile, File, HTTPException

from fastapi.responses import JSONResponse

from marker.converters.pdf import PdfConverter

from marker.models import createmodeldict

from marker.output import textfromrendered

app = FastAPI(title="Marker PDF Converter API")

converter = None

@app.onevent("startup")

async def loadmodels():

global converter

converter = PdfConverter(

artifactdict=createmodeldict(),

)

@app.post("/convert")

async def convertpdf(file: UploadFile = File(...)):

"""Convert PDF to Markdown."""

if not file.filename.endswith((".pdf", ".epub", ".mobi")):

raise HTTPException(

statuscode=400,

detail="Unsupported file format. Use PDF, EPUB, or MOBI."

)

with tempfile.NamedTemporaryFile(

suffix=os.path.splitext(file.filename)[1],

delete=False

) as tmp:

content = await file.read()

tmp.write(content)

tmppath = tmp.name

try:

rendered = converter(tmppath)

text, metadata, images = textfromrendered(rendered)

return JSONResponse({

"filename": file.filename,

"markdown": text,

"metadata": metadata,

"imagescount": len(images) if images else 0

})

except Exception as e:

raise HTTPException(statuscode=500, detail=str(e))

finally:

os.unlink(tmppath)

@app.get("/health")

async def health():

return {"status": "ok", "modelloaded": converter is not None}

Run the server:

pip install fastapi uvicorn python-multipart

uvicorn markerapi:app --host 0.0.0.0 --port 8000

Test with curl:

curl -X POST http://localhost:8000/convert \

-F "file=@document.pdf" \

| python -m json.tool

Quality Comparison with Other Tools

Here's a comparison of Marker with other text extraction tools on documents with complex tables:

import time

def compareextraction(pdfpath: str):

"""Compare extraction results across different tools."""

# Marker

from marker.converters.pdf import PdfConverter

from marker.models import createmodeldict

from marker.output import textfromrendered

start = time.time()

converter = PdfConverter(artifactdict=createmodeldict())

rendered = converter(pdfpath)

markertext, , = textfromrendered(rendered)

markertime = time.time() - start

# PyPDF2

from PyPDF2 import PdfReader

start = time.time()

reader = PdfReader(pdfpath)

pypdftext = ""

for page in reader.pages:

pypdftext += page.extracttext() or ""

pypdftime = time.time() - start

# pdfminer

from pdfminer.highlevel import extracttext

start = time.time()

pdfminertext = extracttext(pdfpath)

pdfminertime = time.time() - start

print("=== Comparison Results ===")

print(f"Marker : {len(markertext)} chars, {markertime:.2f}s")

print(f"PyPDF2 : {len(pypdftext)} chars, {pypdftime:.2f}s")

print(f"pdfminer : {len(pdfminertext)} chars, {pdfminertime:.2f}s")

print()

print("--- Marker Output (first 500 chars) ---")

print(markertext[:500])

print()

print("--- PyPDF2 Output (first 500 chars) ---")

print(pypdftext[:500])

compareextraction("complexreport.pdf")

Best Practices

1. Memory Optimization for Large Documents

When processing large documents (hundreds of pages), pay attention to memory usage:

import gc

import torch

def processlargedocument(pdfpath: str, batchsize: int = 20):

"""Process large documents in page batches."""

from marker.converters.pdf import PdfConverter

from marker.models import createmodeldict

from marker.output import textfromrendered

import fitz # PyMuPDF

doc = fitz.open(pdfpath)

totalpages = len(doc)

doc.close()

alltext = []

converter = PdfConverter(artifactdict=createmodeldict())

for start in range(0, totalpages, batchsize):

end = min(start + batchsize, totalpages)

rendered = converter(pdfpath, pagerange=[start, end])

text, , = textfromrendered(rendered)

alltext.append(text)

gc.collect()

if torch.cuda.isavailable():

torch.cuda.emptycache()

return "\n\n".join(alltext)

2. Caching Conversion Results

Avoid re-converting the same documents:

import hashlib

import json

import os

CACHEDIR = "./markercache"

def getfilehash(filepath: str) -> str:

"""Calculate SHA256 hash of a file."""

sha256 = hashlib.sha256()

with open(filepath, "rb") as f:

for chunk in iter(lambda: f.read(8192), b""):

sha256.update(chunk)

return sha256.hexdigest()

def convertwithcache(pdfpath: str) -> str:

"""Convert with file-hash-based caching."""

os.makedirs(CACHEDIR, existok=True)

filehash = getfilehash(pdfpath)

cachepath = os.path.join(CACHEDIR, f"{filehash}.json")

if os.path.exists(cachepath):

with open(cachepath, "r") as f:

cached = json.load(f)

print(f"Cache hit: {pdfpath}")

return cached["text"]

from marker.converters.pdf import PdfConverter

from marker.models import createmodeldict

from marker.output import textfromrendered

converter = PdfConverter(artifactdict=createmodeldict())

rendered = converter(pdfpath)

text, metadata, = textfromrendered(rendered)

with open(cachepath, "w") as f:

json.dump({"text": text, "metadata": metadata}, f)

print(f"Converted and cached: {pdfpath}")

return text

3. Output Validation

Always validate output quality before using it:

def validateconversion(text: str, minlength: int = 100) -> dict:

"""Validate conversion output quality."""

issues = []

if len(text) < minlength:

issues.append(f"Text too short ({len(text)} chars)")

lines = text.split("\n")

emptyratio = sum(1 for l in lines if not l.strip()) / max(len(lines), 1)

if emptyratio > 0.7:

issues.append(f"Too many empty lines ({emptyratio:.0%})")

garbled = sum(1 for c in text if ord(c) > 65535) / max(len(text), 1)

if garbled > 0.05:

issues.append(f"Possible garbled characters ({garbled:.1%})")

tablecount = text.count("|---|")

headingcount = text.count("\n#")

return {

"valid": len(issues) == 0,

"issues": issues,

"stats": {

"length": len(text),

"lines": len(lines),

"tables": tablecount,

"headings": headingcount,

"emptylineratio": round(emptyratio, 2)

}

}

4. Proper Error Handling

from pathlib import Path

def safeconvert(pdfpath: str, fallbackocr: bool = True) -> dict:

"""Convert with comprehensive error handling."""

path = Path(pdfpath)

if not path.exists():

return {"status": "error", "error": f"File not found: {pdfpath}"}

if not path.suffix.lower() in [".pdf", ".epub", ".mobi"]:

return {"status": "error", "error": f"Unsupported format: {path.suffix}"}

filesizemb = path.stat().stsize / (1024 1024)

if filesizemb > 500:

return {"status": "error", "error": f"File too large: {filesizemb:.0f}MB (max 500MB)"}

try:

from marker.converters.pdf import PdfConverter

from marker.models import createmodeldict

from marker.output import textfromrendered

import marker.settings as settings

if fallbackocr:

settings.OCRALLPAGES = True

converter = PdfConverter(artifactdict=createmodeldict())

rendered = converter(pdfpath)

text, metadata, images = textfromrendered(rendered)

validation = validateconversion(text)

return {

"status": "success",

"text": text,

"metadata": metadata,

"imagescount": len(images) if images else 0,

"validation": validation

}

except Exception as e:

return {"status": "error", "error": str(e)}

5. Performance Tips

Here are some tips to optimize Marker's performance:

  • Use GPU: Conversion is 3-5x faster with a CUDA GPU
  • Batch processing: Use the marker CLI for multiple files instead of looping markersingle
  • Page range: If you only need specific pages, use the pagerange parameter
  • SSD storage: Marker models are quite large; SSD speeds up model loading
  • Caching: Implement caching for frequently accessed documents

Conclusion

Marker is an extremely powerful solution for converting PDF, EPUB, and MOBI documents into high-quality Markdown. Its key advantages are:

  • High accuracy in recognizing layouts, tables, and mathematical formulas
  • Multi-format support for PDF, EPUB, MOBI, and images
  • Multilingual with Surya OCR support
  • Open-source and actively maintained
  • Easy to integrate into RAG pipelines and Python applications

With Marker, you can build reliable document processing pipelines for various needs, from corporate knowledge bases and document search systems to RAG-based chatbots that use internal documents as knowledge sources.

For more information, visit the official Marker repository on GitHub and the Surya OCR documentation for additional language configurations.

Related Articles

Docling: Smart Document Parsing for AI and RAG Pipelines

Docling: Document Parsing Cerdas untuk Pipeline AI dan RAG Dalam era AI generatif, kemampuan untuk mengekstrak informasi...

PaddleOCR: High Accuracy Text Extraction from Images and Documents

PaddleOCR: Ekstraksi Teks dari Gambar dan Dokumen dengan Akurasi Tinggi Halo temen-temen, kali ini kita bahas salah satu...

DSPy: Stop Hand-Tuning Prompts, Let the Compiler Optimize Them

DSPy: Berhenti Ngoprek Prompt Manual, Biarkan Compiler yang Optimasi Halo temen-temen, kali ini aku mau ngenalin satu li...

FastEmbed: Fast and Lightweight Embeddings Without Torch, by Qdrant

FastEmbed: Bikin Embedding Cepat dan Ringan Tanpa Torch dari Qdrant Halo temen-temen, balik lagi sama aku Ruby Abdullah....