Kedro: Reproducible, Maintainable Data Science Pipelines
Most data science projects start in a single notebook and slowly turn into a tangle of cells, hardcoded file paths, and code that only runs on one laptop. Kedro is an open-source Python framework that applies software-engineering discipline to data science work, giving your project a standard structure, a declarative way to manage data, and a clean separation between code, configuration, and data. This tutorial walks through the concepts and shows a complete, runnable example.
What Kedro Is and the Problems It Solves
Kedro is not a scheduler and not a notebook tool. It is a project framework: an opinionated way to organise a data science codebase so that it is reproducible, testable, and ready to hand off to other engineers. It was created at QuantumBlack (part of McKinsey) and is now maintained under the LF AI & Data Foundation.
Anyone who has maintained data science code recognises the recurring problems:
- Notebook chaos. Logic lives in cells that must be run in the right order. Hidden state makes results impossible to reproduce.
- Hardcoded paths.
pd.readcsv("/Users/ruby/Downloads/datav3final.csv")works on exactly one machine and breaks the moment a colleague clones the repository. - Tangled concerns. Connection strings, model hyperparameters, and business logic are interleaved in the same file, so changing an environment means editing source code.
- No clear lineage. It is unclear which function produces which artifact and what depends on what.
Kedro addresses these through a few core principles:
- Separation of code, configuration, and data. Code lives in
src/, configuration inconf/, data indata/. Each can change independently. - A declarative Data Catalog. Datasets are named entities described in YAML, never raw paths buried in code.
- Modularity. Work is broken into small pure functions (nodes) composed into pipelines, which makes testing and reuse straightforward.
- Reproducibility. Anyone can clone the repository, install dependencies, and run
kedro runto get the same result.
Importantly, Kedro provides structure, not scheduling. When you need cron-like orchestration, retries, or a production scheduler, you deploy a Kedro pipeline onto Airflow, Dagster, Argo, Databricks, or similar. Kedro gives shape to the project; those tools run it on a schedule.
Installation and Creating a Project
Kedro requires Python 3.9 or newer. Always install into a virtual environment.
python -m venv .venv
source .venv/bin/activate # on Windows: .venv\Scripts\activate
pip install kedro
kedro info
Create a new project. Kedro uses starters (templates); the example below uses the standard tooling prompts.
kedro new --name price-prediction
You will be asked which tooling to include (linting, testing, logging, documentation, data structure). For this tutorial we assume a project named price-prediction for a house price regression model.
The Standard Directory Structure
price-prediction/
├── conf/
│ ├── base/
│ │ ├── catalog.yml # dataset definitions
│ │ ├── parameters.yml # pipeline parameters
│ │ └── logging.yml
│ └── local/
│ └── credentials.yml # secrets, never committed
├── data/
│ ├── 01raw/ # immutable source data
│ ├── 02intermediate/ # typed, cleaned data
│ ├── 03primary/
│ ├── 04feature/
│ ├── 05modelinput/
│ ├── 06models/
│ ├── 07modeloutput/
│ └── 08reporting/
├── notebooks/
├── src/
│ └── priceprediction/
│ ├── pipelines/
│ ├── pipelineregistry.py
│ └── settings.py
├── pyproject.toml
└── requirements.txt
The layered data/ folders express a data-engineering convention: raw data is immutable, and each subsequent layer is derived from the previous one. You are free to use these layers as labels; Kedro does not force a fixed meaning, but the convention makes lineage obvious at a glance.
The split between conf/base and conf/local is central. base holds configuration shared by everyone and is committed to version control. local holds machine-specific overrides and secrets and is git-ignored. This is how Kedro keeps credentials out of code.
The Data Catalog
The Data Catalog is the heart of Kedro. Instead of opening files in code, you declare each dataset once in conf/base/catalog.yml and refer to it by name everywhere else. Kedro handles loading and saving with the right connector.
# conf/base/catalog.yml
housesraw:
type: pandas.CSVDataset
filepath: data/01raw/houses.csv
housescleaned:
type: pandas.ParquetDataset
filepath: data/02intermediate/housescleaned.parquet
modelinputtable:
type: pandas.ParquetDataset
filepath: data/05modelinput/modelinput.parquet
pricemodel:
type: pickle.PickleDataset
filepath: data/06models/pricemodel.pkl
versioned: true
A few things are worth highlighting:
- No hardcoded paths in code. A node receives a loaded DataFrame and returns one; it never knows where the data lives. To move from CSV to Parquet, or from local disk to S3, you edit YAML, not Python.
- Many dataset types. Kedro ships connectors for CSV, Parquet, Excel, JSON, pickle, SQL tables and queries, Spark, image data, and more, through the
kedro-datasetspackage. - Versioning. Setting
versioned: truemakes Kedro write each run's output into a timestamped subfolder, so model artifacts are never silently overwritten. - MemoryDataset. Any output that is not declared in the catalog is automatically held in memory as a
MemoryDatasetand passed to downstream nodes within the same run. This is convenient for transient intermediate results you do not need to persist.
A SQL example and a cloud example:
# Reading from a database
customertable:
type: pandas.SQLTableDataset
credentials: dbcredentials
tablename: customers
loadargs:
schema: public
Writing to cloud storage
predictions:
type: pandas.CSVDataset
filepath: s3://my-bucket/predictions/output.csv
credentials: awscreds
Credentials referenced by name (dbcredentials, awscreds) are resolved from conf/local/credentials.yml, keeping secrets out of the committed catalog.
Nodes
A node is a thin wrapper around a pure Python function. The function knows nothing about Kedro; the node maps the function's inputs and outputs to catalog dataset names.
# src/priceprediction/pipelines/dataprocessing/nodes.py
import pandas as pd
def clean
houses(housesraw: pd.DataFrame) -> pd.DataFrame:
df = houses
raw.copy()
df = df.dropna(subset=["price", "areasqm"])
df["pricepersqm"] = df["price"] / df["areasqm"]
df["hasgarage"] = df["garage"].fillna(0).astype(int)
return df
def buildmodelinput(housescleaned: pd.DataFrame) -> pd.DataFrame:
features = ["areasqm", "bedrooms", "bathrooms", "hasgarage", "pricepersqm"]
return housescleaned[features + ["price"]]
Keeping the functions pure (no I/O, no global state) means they are trivial to unit test in isolation, with no Kedro machinery involved.
Pipelines
A pipeline connects nodes into a directed acyclic graph. You do not specify the order explicitly; Kedro infers it from the inputs and outputs each node declares. If node B consumes what node A produces, Kedro runs A before B.
# src/priceprediction/pipelines/dataprocessing/pipeline.py
from kedro.pipeline import Pipeline, node, pipeline
from .nodes import clean
houses, buildmodelinput
def createpipeline(kwargs) -> Pipeline:
return pipeline(
[
node(
func=cleanhouses,
inputs="housesraw",
outputs="housescleaned",
name="cleanhousesnode",
),
node(
func=buildmodelinput,
inputs="housescleaned",
outputs="modelinputtable",
name="buildmodelinputnode",
),
]
)
The strings housesraw, housescleaned, and modelinputtable are the same names defined in the catalog. That is the wiring: catalog names connect nodes to data and nodes to each other.
A Data Science Pipeline
Now a second pipeline for training and evaluation. Note how parameters flow in via the params: prefix.
# src/priceprediction/pipelines/datascience/nodes.py
import logging
import pandas as pd
from sklearn.ensemble import RandomForestRegressor
from sklearn.modelselection import traintestsplit
from sklearn.metrics import meanabsoluteerror, r2score
logger = logging.getLogger(name)
def splitdata(modelinputtable: pd.DataFrame, parameters: dict):
X = modelinputtable.drop(columns=["price"])
y = modelinputtable["price"]
return traintestsplit(
X, y,
testsize=parameters["testsize"],
randomstate=parameters["randomstate"],
)
def trainmodel(Xtrain, ytrain, parameters: dict) -> RandomForestRegressor:
model = RandomForestRegressor(
nestimators=parameters["nestimators"],
maxdepth=parameters["maxdepth"],
randomstate=parameters["randomstate"],
)
model.fit(Xtrain, ytrain)
return model
def evaluatemodel(model: RandomForestRegressor, Xtest, ytest) -> dict:
preds = model.predict(Xtest)
mae = meanabsoluteerror(ytest, preds)
r2 = r2score(ytest, preds)
logger.info("Model MAE: %.2f, R2: %.3f", mae, r2)
return {"mae": mae, "r2": r2}
# src/priceprediction/pipelines/datascience/pipeline.py
from kedro.pipeline import Pipeline, node, pipeline
from .nodes import split
data, trainmodel, evaluatemodel
def createpipeline(kwargs) -> Pipeline:
return pipeline(
[
node(
func=splitdata,
inputs=["modelinputtable", "params:modeloptions"],
outputs=["Xtrain", "Xtest", "ytrain", "ytest"],
name="splitdatanode",
),
node(
func=trainmodel,
inputs=["Xtrain", "ytrain", "params:modeloptions"],
outputs="pricemodel",
name="trainmodelnode",
),
node(
func=evaluatemodel,
inputs=["pricemodel", "Xtest", "ytest"],
outputs="modelmetrics",
name="evaluatemodelnode",
),
]
)
Xtrain, Xtest, ytrain, and ytest are not in the catalog, so Kedro treats them as in-memory datasets passed between nodes. pricemodel is in the catalog, so it gets persisted (and versioned).
Registering Pipelines and the Default Pipeline
Kedro discovers pipelines through pipelineregistry.py. You assemble the full project pipeline here, including a default that runs when no specific pipeline is named.
# src/priceprediction/pipelineregistry.py
from kedro.pipeline import Pipeline
from price
prediction.pipelines import dataprocessing as dp
from price
prediction.pipelines import datascience as ds
def register
pipelines() -> dict[str, Pipeline]:
dataprocessing = dp.createpipeline()
datascience = ds.createpipeline()
return {
"dp": dataprocessing,
"ds": datascience,
"default": dataprocessing + datascience,
}
Adding two pipelines with + concatenates them into one graph. Because modelinputtable produced by the first pipeline is consumed by the second, Kedro automatically orders the whole thing correctly end to end.
Running Pipelines
The basic command runs the default pipeline:
kedro run
Kedro offers fine-grained control over what runs:
# Run only the data science pipeline
kedro run --pipeline ds
Run specific nodes by name
kedro run --nodes "trainmodelnode,evaluatemodelnode"
Run from one node onward, or up to a node
kedro run --from-nodes splitdatanode
kedro run --to-nodes buildmodelinputnode
Run everything tagged for a particular concern
kedro run --tags training
These options make iteration cheap: when you change only the evaluation logic, you can re-run from that point instead of recomputing the entire graph. Tags are attached at node definition time and group nodes across pipelines.
Parameters and Configuration
Hyperparameters and tunable settings live in conf/base/parameters.yml, never hardcoded in nodes.
# conf/base/parameters.yml
modeloptions:
testsize: 0.2
randomstate: 42
nestimators: 200
maxdepth: 12
The params:modeloptions reference in a node's inputs injects this dictionary into the function. To reference a single value you can use dotted access, for example params:modeloptions.testsize.
Configuration Environments
Kedro merges configuration from conf/base (shared) and conf/local (your machine), with local overriding base. You can create additional environments and select them at runtime:
kedro run --env staging
This loads conf/staging on top of conf/base. A typical pattern: base points at sample data, while local or staging points at a real warehouse table, with no code change.
Credentials and OmegaConf Templating
Secrets go in conf/local/credentials.yml, which is git-ignored:
# conf/local/credentials.yml
dbcredentials:
con: postgresql://user:password@host:5432/prod
awscreds:
clientkwargs:
awsaccesskeyid: AKIA...
awssecretaccesskey: ...
Kedro uses OmegaConf as its default config loader, which supports variable interpolation and reusable templates. You can define an anchor once and reuse it across the catalog:
# conf/base/catalog.yml
parquet: &parquet
type: pandas.ParquetDataset
save
args:
compression: snappy
housescleaned:
<<: parquet
filepath: data/02intermediate/housescleaned.parquet
modelinputtable:
<<: parquet
filepath: data/05modelinput/modelinput.parquet
You can also interpolate environment variables and parameters into the catalog, which keeps configuration concise and consistent.
Visualizing the Pipeline with Kedro-Viz
Kedro-Viz renders your DAG as an interactive diagram, showing nodes, datasets, and how data flows between them.
pip install kedro-viz
kedro viz run
This opens a browser at http://127.0.0.1:4141 with the full graph. It is useful for onboarding new team members, spotting unintended dependencies, and explaining the workflow to non-engineers. Kedro-Viz can also display experiment-tracking metrics over time when you log them through the catalog.
Working Interactively: Catalog, IPython, and Jupyter
Kedro is designed to coexist with notebooks rather than ban them. Several commands bridge the two worlds:
# List and inspect declared datasets
kedro catalog list
Launch an IPython session with catalog, context, and pipelines preloaded
kedro ipython
Launch Jupyter with the Kedro session available
kedro jupyter notebook
Inside an IPython or Jupyter session you get a ready-made catalog object, so exploratory work uses the same datasets as the pipeline:
houses = catalog.load("housesraw")
houses.describe()
Save an experiment result back through the catalog
catalog.save("houses
cleaned", cleaneddf)
The recommended workflow is to prototype in a notebook, then move stable logic into nodes so it becomes part of the reproducible pipeline.
Extending Kedro: Hooks and Plugins
Hooks let you run custom code at defined points in the lifecycle: before or after a node runs, before or after a pipeline runs, after the catalog is created, and so on. They are the supported mechanism for cross-cutting concerns such as custom logging, data validation, or sending notifications.# src/priceprediction/hooks.py
import logging
from kedro.framework.hooks import hookimpl
logger = logging.getLogger(name)
class ModelTrackingHooks:
@hookimpl
def afternoderun(self, node, outputs):
if node.name == "evaluatemodelnode":
metrics = outputs.get("modelmetrics", {})
logger.info("Logging metrics to tracking system: %s", metrics)
Register hooks in settings.py:
# src/priceprediction/settings.py
from priceprediction.hooks import ModelTrackingHooks
HOOKS = (ModelTrackingHooks(),)
Plugins extend the kedro command itself or add integrations. Common ones include kedro-viz, kedro-datasets, the Airflow and Argo deployment plugins, and MLflow integrations such as kedro-mlflow.
Deployment: Structure Here, Scheduling There
Kedro packages your project as a standard Python distribution:
kedro package
This produces a wheel in dist/ containing your pipelines, plus the configuration needed to run them. From there, Kedro pipelines run on whatever orchestrator you already use. The framing is consistent: Kedro defines what runs and how it is wired; the orchestrator handles when it runs, retries, and scaling.
- Airflow. The
kedro-airflowplugin converts a Kedro pipeline into an Airflow DAG, mapping each Kedro node to an Airflow task.
pip install kedro-airflow
kedro airflow create
- Dagster. Community and plugin integrations expose Kedro nodes as Dagster ops/assets, so the Kedro graph runs inside Dagster's scheduling and observability layer.
- Argo Workflows / Kubeflow. Templates translate the pipeline into Argo workflow steps for Kubernetes-native execution.
- Databricks. Kedro runs on Databricks either as a packaged job or through the Databricks workflow tooling, which suits Spark-based catalogs.
In every case you write the pipeline once in Kedro and choose a deployment target without rewriting business logic.
Experiment Tracking
For tracking runs, metrics, and parameters, Kedro integrates with MLflow through kedro-mlflow, which logs parameters, metrics, and model artifacts automatically as the pipeline runs. Kedro also has lightweight built-in experiment tracking that surfaces metrics in Kedro-Viz when you declare metric datasets in the catalog. Either approach keeps experiment records tied to the exact pipeline version that produced them.
Best Practices
- Keep node functions pure. No file I/O, no global state, no hidden side effects. The catalog handles persistence; nodes handle logic.
- Name everything. Give every node a
nameso it is addressable from the CLI and readable in Kedro-Viz. - One concern per pipeline. Separate data processing from data science, and keep pipelines small and composable.
- Never commit secrets. Credentials belong in
conf/local, which stays out of version control. - Use the data layers. Following the
01rawthrough08_reportingconvention makes lineage self-documenting. - Version your models. Set
versioned: trueon model and output datasets so runs are auditable. - Test nodes directly. Because functions are pure, unit tests need no Kedro context.
- Promote notebook code into nodes. Explore in notebooks, but move anything that matters into the pipeline.
Conclusion and Key Takeaways
Kedro brings the structure of software engineering to data science without forcing you to abandon the exploratory style that makes the field productive. By separating code, configuration, and data, and by composing pure functions into declarative pipelines, it turns one-off notebooks into projects that any teammate can clone, run, and trust.
Key takeaways:
- Kedro is a project framework, not a scheduler. It defines structure; tools like Airflow, Dagster, and Argo provide scheduling.
- The Data Catalog removes hardcoded paths and centralises every dataset definition in YAML.
- Nodes are pure functions; pipelines compose them into a graph whose order Kedro infers automatically.
- Configuration, parameters, and credentials are separated by environment, never hardcoded.
- Kedro-Viz visualises the DAG, and hooks/plugins extend behaviour and integrate with MLflow and deployment targets.
kedro packageproduces a portable artifact you can run on the orchestrator of your choice.
Start small: model a single pipeline with two or three nodes, declare its datasets in the catalog, and run it with kedro run. The discipline pays off as soon as a second person needs to touch the project.