Pandera: Statistical Data Validation for pandas and Polars DataFrames
Data pipelines fail quietly. A column that should never be negative slips through, a category typo breaks a downstream join, and you only notice when a dashboard looks wrong three days later. Pandera lets you declare what your DataFrames should look like directly in Python and enforce those rules where the data actually flows. This tutorial walks through Pandera from first install to production pipeline integration, using a sales transactions dataset as the running example.
What Pandera Is
Pandera is a data validation library for tabular data. You define a schema, a description of the columns, their types, and the constraints they must satisfy, and then you validate a DataFrame against it. If the data conforms, you get the DataFrame back unchanged. If it does not, Pandera raises an error that tells you exactly which rows and columns failed and why.
The core idea is that schemas are ordinary Python objects. They live in your codebase next to the functions that produce and consume the data. This is sometimes called "schema as code": there is no separate configuration file, no external validation service, and no YAML to keep in sync. You import a schema the same way you import any other module.
Pandera started as a pandas-focused tool and now supports several backends, including Polars and PySpark, through a shared schema model. The validation logic you write is largely the same regardless of the DataFrame engine underneath.
When to Use Pandera versus Great Expectations
Both libraries validate tabular data, but they target different workflows.
Great Expectations is a heavier framework. It maintains a data context, generates HTML data documentation, stores validation results, and is designed around the idea of a centralized data quality platform that analysts and engineers share. It is a good fit when you need auditable validation reports, a catalog of expectations, and tooling that non-developers interact with.
Pandera is lightweight and code-first. It adds validation inline in the same Python process that transforms the data. There is no context to configure and no artifact store. You reach for Pandera when you want assertions that live inside your ETL functions, type-checked function boundaries, and validation that runs as a normal part of your pipeline without extra infrastructure.
A practical rule: if validation is a developer concern embedded in code, choose Pandera. If validation is an organizational concern with shared documentation and reporting, Great Expectations earns its weight. The two are not mutually exclusive; some teams use Pandera for fast inline checks and Great Expectations for the documented contract.
Installation
Install the core package with pip.
pip install pandera
Backend support and optional features are distributed as extras. Install only what you need.
# Polars support
pip install 'pandera[polars]'
PySpark support
pip install 'pandera[pyspark]'
Hypothesis-based statistical checks
pip install 'pandera[hypotheses]'
Everything
pip install 'pandera[all]'
Verify the install and check the version.
import pandera as pa
print(pa.version)
A First Schema with DataFrameSchema
The object-based API centers on DataFrameSchema, which holds a mapping of column names to Column definitions. Here is a schema for a sales transactions table.
import pandas as pd
import pandera as pa
from pandera import Column, Check, DataFrameSchema
schema = DataFrameSchema(
{
"transactionid": Column(int, unique=True),
"product": Column(str),
"category": Column(str),
"quantity": Column(int, Check.greaterthan(0)),
"unitprice": Column(float, Check.greaterthanorequalto(0)),
"region": Column(str),
"customeremail": Column(str, nullable=True),
},
strict=True,
coerce=True,
)
Each Column takes a dtype and optional constraints. A few parameters do most of the work:
nullableallows missing values in the column. It defaults toFalse, so by default a column may not contain nulls.uniquerequires every value in the column to be distinct.coercecasts the column to the declared dtype before validating, instead of failing when the type does not match exactly. It can be set per column or for the whole schema.
The strict=True argument on the schema rejects any DataFrame that contains columns not declared in the schema. Without it, extra columns are ignored.
Validate a DataFrame by calling schema.validate or simply calling the schema.
df = pd.DataFrame(
{
"transactionid": [1, 2, 3],
"product": ["Keyboard", "Mouse", "Monitor"],
"category": ["Peripherals", "Peripherals", "Displays"],
"quantity": [2, 1, 1],
"unitprice": [29.99, 14.50, 199.00],
"region": ["West", "East", "West"],
"customeremail": ["a@example.com", None, "c@example.com"],
}
)
validated = schema.validate(df)
When validation succeeds, validated is the same data, possibly with coerced dtypes. When it fails, Pandera raises a SchemaError describing the first failing check.
Built-in Checks
The Check namespace provides a library of common constraints. They cover the validations you write most often.
schema = DataFrameSchema(
{
"quantity": Column(int, Check.greaterthan(0)),
"discount": Column(float, Check.inrange(0.0, 1.0)),
"category": Column(
str, Check.isin(["Peripherals", "Displays", "Cables", "Audio"])
),
"sku": Column(str, Check.strmatches(r"^SKU-\d{6}$")),
"rating": Column(int, Check.lessthanorequalto(5)),
}
)
Frequently used built-ins include greaterthan, greaterthanorequalto, lessthan, lessthanorequalto, inrange, isin, notin, strmatches, strcontains, strlength, and uniquevalueseq. You can pass a list of checks to a single column when more than one constraint applies.
Column(
float,
checks=[
Check.greaterthan(0),
Check.lessthan(100000),
],
)
Custom Checks with Lambdas
When a built-in does not fit, write a custom Check. The simplest form takes a function that receives the column as a pandas Series and returns a boolean Series.
schema = DataFrameSchema(
{
"unitprice": Column(
float,
Check(lambda s: s.round(2).eq(s).all(), error="prices must have at most 2 decimals"),
),
"quantity": Column(
int,
Check(lambda s: s % 1 == 0, error="quantity must be whole"),
),
}
)
A check can also express a relationship across the whole column, such as requiring the mean to fall in a sensible range.
Check(lambda s: s.mean() < 1000, error="average unit price too high")
Element-wise versus Vectorized Checks
By default a Check is vectorized: the function receives the entire column as a Series and should return a boolean Series (one result per row) or a single boolean. Vectorized checks are fast because they use pandas operations.
Sometimes it is clearer to reason about one value at a time. Set elementwise=True and the function receives individual scalar values instead.
# Vectorized: operates on the whole Series at once (preferred for speed)
Check(lambda s: s.str.startswith("SKU-"))
Element-wise: called once per value, receives a single scalar
Check(lambda v: v.startswith("SKU-"), elementwise=True)
Prefer vectorized checks for performance. Reach for elementwise only when the per-row logic is awkward to vectorize, such as parsing a complex string with a try/except.
The Class-Based API with DataFrameModel
For larger projects the class-based API is easier to read and integrates with type hints. You subclass DataFrameModel and declare columns as typed attributes using Series and Field.
import pandera as pa
from pandera.typing import Series
class TransactionSchema(pa.DataFrameModel):
transactionid: Series[int] = pa.Field(unique=True)
product: Series[str]
category: Series[str] = pa.Field(isin=["Peripherals", "Displays", "Cables", "Audio"])
quantity: Series[int] = pa.Field(gt=0)
unitprice: Series[float] = pa.Field(ge=0)
region: Series[str] = pa.Field(isin=["West", "East", "North", "South"])
customeremail: Series[str] = pa.Field(nullable=True)
class Config:
strict = True
coerce = True
Field accepts the same constraints as built-in checks, expressed as keyword arguments: gt, ge, lt, le, isin, strmatches, inrange, unique, nullable, and so on. The nested Config class holds schema-level options such as strict and coerce.
Validate by calling .validate on the model class.
validated = TransactionSchema.validate(df)
You can also attach custom checks as methods using the @pa.check and @pa.dataframecheck decorators.
class TransactionSchema(pa.DataFrameModel):
quantity: Series[int] = pa.Field(gt=0)
unitprice: Series[float] = pa.Field(ge=0)
@pa.check("unitprice")
def priceprecision(cls, s: Series[float]) -> Series[bool]:
return s.round(2).eq(s)
@pa.dataframecheck
def totalispositive(cls, df: pd.DataFrame) -> Series[bool]:
return (df["quantity"] df["unitprice"]) > 0
The @pa.check decorator validates a single column, while @pa.dataframecheck validates a relationship across multiple columns.
Validating Function Boundaries with @pa.checktypes
One of the most useful features of the class-based API is validating the inputs and outputs of a function automatically. Annotate parameters and the return value with DataFrame[Schema] and decorate the function with @pa.checktypes.
from pandera.typing import DataFrame
class RawTransactions(pa.DataFrameModel):
transactionid: Series[int]
quantity: Series[int]
unitprice: Series[float]
class Config:
coerce = True
class EnrichedTransactions(RawTransactions):
total: Series[float] = pa.Field(ge=0)
@pa.checktypes
def addtotal(df: DataFrame[RawTransactions]) -> DataFrame[EnrichedTransactions]:
return df.assign(total=df["quantity"] df["unitprice"])
When addtotal is called, Pandera validates the argument against RawTransactions before the body runs and validates the return value against EnrichedTransactions afterward. This turns schema conformance into a checked contract at every function boundary, which is especially valuable in long transformation chains.
Reusable and Registered Checks
When the same custom logic appears in several schemas, register it once and reference it by name. The @pa.extensions.registercheckmethod decorator adds a check to the Check namespace.
import pandera.extensions as extensions
@extensions.registercheckmethod(statistics=["maxdecimals"])
def hasmaxdecimals(pandasobj, , maxdecimals):
factor = 10 maxdecimals
return (pandasobj factor).round() == (pandasobj factor)
schema = DataFrameSchema(
{
"unitprice": Column(float, Check.hasmaxdecimals(maxdecimals=2)),
}
)
The statistics list names the parameters the check accepts, which lets Pandera display them in error messages and serialize the schema correctly.
Lazy Validation and Inspecting Failures
By default Pandera stops at the first failing check and raises a SchemaError. During data exploration or batch validation you usually want the full picture. Pass lazy=True to collect every failure and raise a single SchemaErrors exception.
import pandera as pa
bad = pd.DataFrame(
{
"transactionid": [1, 1, 3], # duplicate id
"product": ["Keyboard", "Mouse", "Monitor"],
"category": ["Peripherals", "Unknown", "Displays"], # invalid category
"quantity": [2, -1, 1], # negative quantity
"unitprice": [29.99, 14.50, 199.00],
"region": ["West", "East", "West"],
"customeremail": ["a@example.com", None, "c@example.com"],
}
)
try:
schema.validate(bad, lazy=True)
except pa.errors.SchemaErrors as exc:
print(exc.failurecases) # a DataFrame of every failing case
print(exc.data) # the original data that was validated
The failurecases attribute is itself a DataFrame with columns for the schema element, the check that failed, the failing value, and the row index. This is far more useful than a single error when you are cleaning a large dataset, because you see every problem in one pass.
Statistical Checks with Hypotheses
Beyond row-level constraints, Pandera can run statistical hypothesis tests through the Hypothesis class. These verify properties of the distribution rather than individual values. For example, you can assert that the mean unit price of one region is significantly higher than another using a two-sample t-test.
from pandera import Hypothesis
schema = DataFrameSchema(
{
"unitprice": Column(float),
"region": Column(str),
},
checks=Hypothesis.twosamplettest(
sample1="West",
sample2="East",
groupby="region",
relationship="greaterthan",
alpha=0.05,
),
)
Hypothesis checks require the hypotheses extra and the underlying SciPy stack. Use them sparingly; they are best for monitoring data drift or validating assumptions about distributions rather than as routine field checks.
Validating the Index and MultiIndex
Schemas can constrain the index as well as the columns. Use the Index and MultiIndex classes.
from pandera import Index, MultiIndex
Single index: a unique, sorted integer index
schema = DataFrameSchema(
columns={"unitprice": Column(float)},
index=Index(int, Check.greaterthanorequalto(0), unique=True),
)
MultiIndex: validate each level
multi = DataFrameSchema(
columns={"unitprice": Column(float)},
index=MultiIndex(
[
Index(str, name="region"),
Index(pd.Timestamp, name="date"),
]
),
)
In the class-based API, declare index fields with Index from pandera.typing.
from pandera.typing import Index, Series
class IndexedTransactions(pa.DataFrameModel):
idx: Index[int] = pa.Field(ge=0, unique=True)
unitprice: Series[float]
Integrating Validation into a Pipeline
The strongest use of Pandera is validating data between ETL stages, so that a defect is caught at the boundary where it appears rather than far downstream. Decorating each stage with @pa.checktypes makes the schema the contract for that stage.
from pandera.typing import DataFrame
class RawSales(pa.DataFrameModel):
transactionid: Series[int] = pa.Field(unique=True)
product: Series[str]
quantity: Series[int] = pa.Field(gt=0)
unitprice: Series[float] = pa.Field(ge=0)
class Config:
coerce = True
class CleanSales(RawSales):
revenue: Series[float] = pa.Field(ge=0)
class RegionSummary(pa.DataFrameModel):
region: Series[str]
totalrevenue: Series[float] = pa.Field(ge=0)
@pa.checktypes
def extract(path: str) -> DataFrame[RawSales]:
return pd.readcsv(path)
@pa.checktypes
def transform(df: DataFrame[RawSales]) -> DataFrame[CleanSales]:
return df.assign(revenue=df["quantity"] df["unitprice"])
@pa.checktypes
def summarize(df: DataFrame[CleanSales]) -> DataFrame[RegionSummary]:
return (
df.groupby("region", asindex=False)["revenue"]
.sum()
.rename(columns={"revenue": "totalrevenue"})
)
If extract reads a CSV with a duplicate transactionid, the failure surfaces immediately at extraction, not after the data has been aggregated and the original rows are gone. Each schema also documents what the stage expects, so the pipeline is self-describing.
Validating Polars and PySpark DataFrames
Pandera shares one schema model across backends. For Polars, import the schema classes from pandera.polars and annotate with the Polars typing module.
import polars as pl
import pandera.polars as pa
from pandera.typing.polars import Series
class PolarsTransactions(pa.DataFrameModel):
transactionid: Series[int] = pa.Field(unique=True)
quantity: Series[int] = pa.Field(gt=0)
unitprice: Series[float] = pa.Field(ge=0)
df = pl.DataFrame(
{"transactionid": [1, 2], "quantity": [1, 2], "unitprice": [9.9, 19.9]}
)
validated = PolarsTransactions.validate(df)
PySpark support follows the same shape through pandera.pyspark. The check vocabulary is consistent across engines, so a schema you understand for pandas reads the same for Polars. The main differences are the import paths and that some pandas-specific checks may not have an equivalent on every backend.
Inferring and Exporting Schemas
When you face an unfamiliar dataset, let Pandera draft a schema from a sample with pa.inferschema. The result is a starting point you refine by hand; treat the inferred constraints as suggestions, not final rules.
import pandera as pa
inferred = pa.inferschema(df)
print(inferred)
You can serialize a schema to a Python script or to YAML for review and version control.
# Write a runnable Python script that reconstructs the schema
inferred.toscript("transactionschema.py")
Or serialize to YAML
inferred.toyaml("transactionschema.yaml")
Load a schema back from YAML
schema = pa.DataFrameSchema.fromyaml("transactionschema.yaml")
Inference is most useful as a bootstrapping step. Generated schemas tend to be too permissive in some places and too strict in others, so always review them before relying on them.
Best Practices
Keep schemas close to the code that produces the data. A schema defined next to its transformation function is easier to keep correct than one in a distant configuration file.
Use the class-based DataFrameModel for anything beyond a quick script. It reads better, supports inheritance for related schemas, and works with @pa.checktypes to validate function boundaries.
Validate lazily when cleaning data and eagerly in production. During exploration, lazy=True shows every problem at once. In a running pipeline, failing fast at the first error is usually what you want.
Set coerce deliberately. Coercion is convenient but can hide upstream type problems. Turn it on where you genuinely expect to normalize types and leave it off where a wrong dtype signals a real defect.
Enable strict=True to catch unexpected columns. Silent extra columns often indicate a schema drift or a join gone wrong.
Prefer vectorized checks over elementwise for performance, and reserve hypothesis checks for distribution-level monitoring rather than per-record validation.
Version your schemas alongside your code. Because they are plain Python, they belong in the same repository and review process as the pipeline they protect.
Conclusion and Key Takeaways
Pandera brings data validation into the same place your data is transformed: ordinary Python code. By declaring DataFrameSchema or DataFrameModel definitions, attaching built-in and custom checks, and validating at function boundaries with @pa.checktypes, you turn implicit assumptions about your data into explicit, enforced contracts.
The key points to remember:
- Pandera is lightweight and code-first; choose it for inline pipeline validation, and consider Great Expectations when you need a shared, documented data quality platform.
DataFrameSchemaand the class-basedDataFrameModelexpress the same constraints; the class API scales better and integrates with type hints.- Built-in
Checks cover common rules, custom checks handle the rest, andelementwisetrades speed for per-value clarity. lazy=Truecollects every failure into aSchemaErrorsobject whosefailurecasesDataFrame pinpoints each problem.- The same schema model validates pandas, Polars, and PySpark, and
pa.inferschemaplus YAML or script export help you bootstrap and version schemas.
Start small: add one schema to the most fragile boundary in your pipeline, validate it lazily to see what reality looks like, then tighten the rules until the schema is an honest description of your data.