Kedro Tutorial: Reproducible and Maintainable Data Science Pipelines

# Kedro: Pipeline Data Science yang Reproducible dan Mudah Dirawat Sebagian besar proyek data science dimulai dari satu notebook dan perlahan berubah menjadi kekusutan sel, path file yang ditulis lan...

By Ruby Abdullah · · tutorial
KedroData SciencePipelineMLOpsReproducibilityPython

Kedro: Reproducible, Maintainable Data Science Pipelines

Most data science projects start in a single notebook and slowly turn into a tangle of cells, hardcoded file paths, and code that only runs on one laptop. Kedro is an open-source Python framework that applies software-engineering discipline to data science work, giving your project a standard structure, a declarative way to manage data, and a clean separation between code, configuration, and data. This tutorial walks through the concepts and shows a complete, runnable example.

What Kedro Is and the Problems It Solves

Kedro is not a scheduler and not a notebook tool. It is a project framework: an opinionated way to organise a data science codebase so that it is reproducible, testable, and ready to hand off to other engineers. It was created at QuantumBlack (part of McKinsey) and is now maintained under the LF AI & Data Foundation.

Anyone who has maintained data science code recognises the recurring problems:

  • Notebook chaos. Logic lives in cells that must be run in the right order. Hidden state makes results impossible to reproduce.
  • Hardcoded paths. pd.readcsv("/Users/ruby/Downloads/datav3final.csv") works on exactly one machine and breaks the moment a colleague clones the repository.
  • Tangled concerns. Connection strings, model hyperparameters, and business logic are interleaved in the same file, so changing an environment means editing source code.
  • No clear lineage. It is unclear which function produces which artifact and what depends on what.

Kedro addresses these through a few core principles:

  • Separation of code, configuration, and data. Code lives in src/, configuration in conf/, data in data/. Each can change independently.
  • A declarative Data Catalog. Datasets are named entities described in YAML, never raw paths buried in code.
  • Modularity. Work is broken into small pure functions (nodes) composed into pipelines, which makes testing and reuse straightforward.
  • Reproducibility. Anyone can clone the repository, install dependencies, and run kedro run to get the same result.

Importantly, Kedro provides structure, not scheduling. When you need cron-like orchestration, retries, or a production scheduler, you deploy a Kedro pipeline onto Airflow, Dagster, Argo, Databricks, or similar. Kedro gives shape to the project; those tools run it on a schedule.

Installation and Creating a Project

Kedro requires Python 3.9 or newer. Always install into a virtual environment.

python -m venv .venv

source .venv/bin/activate # on Windows: .venv\Scripts\activate

pip install kedro

kedro info

Create a new project. Kedro uses starters (templates); the example below uses the standard tooling prompts.

kedro new --name price-prediction

You will be asked which tooling to include (linting, testing, logging, documentation, data structure). For this tutorial we assume a project named price-prediction for a house price regression model.

The Standard Directory Structure

price-prediction/

├── conf/

│ ├── base/

│ │ ├── catalog.yml # dataset definitions

│ │ ├── parameters.yml # pipeline parameters

│ │ └── logging.yml

│ └── local/

│ └── credentials.yml # secrets, never committed

├── data/

│ ├── 01raw/ # immutable source data

│ ├── 02intermediate/ # typed, cleaned data

│ ├── 03primary/

│ ├── 04feature/

│ ├── 05modelinput/

│ ├── 06models/

│ ├── 07modeloutput/

│ └── 08reporting/

├── notebooks/

├── src/

│ └── priceprediction/

│ ├── pipelines/

│ ├── pipelineregistry.py

│ └── settings.py

├── pyproject.toml

└── requirements.txt

The layered data/ folders express a data-engineering convention: raw data is immutable, and each subsequent layer is derived from the previous one. You are free to use these layers as labels; Kedro does not force a fixed meaning, but the convention makes lineage obvious at a glance.

The split between conf/base and conf/local is central. base holds configuration shared by everyone and is committed to version control. local holds machine-specific overrides and secrets and is git-ignored. This is how Kedro keeps credentials out of code.

The Data Catalog

The Data Catalog is the heart of Kedro. Instead of opening files in code, you declare each dataset once in conf/base/catalog.yml and refer to it by name everywhere else. Kedro handles loading and saving with the right connector.

# conf/base/catalog.yml

housesraw:

type: pandas.CSVDataset

filepath: data/01raw/houses.csv

housescleaned:

type: pandas.ParquetDataset

filepath: data/02intermediate/housescleaned.parquet

modelinputtable:

type: pandas.ParquetDataset

filepath: data/05modelinput/modelinput.parquet

pricemodel:

type: pickle.PickleDataset

filepath: data/06models/pricemodel.pkl

versioned: true

A few things are worth highlighting:

  • No hardcoded paths in code. A node receives a loaded DataFrame and returns one; it never knows where the data lives. To move from CSV to Parquet, or from local disk to S3, you edit YAML, not Python.
  • Many dataset types. Kedro ships connectors for CSV, Parquet, Excel, JSON, pickle, SQL tables and queries, Spark, image data, and more, through the kedro-datasets package.
  • Versioning. Setting versioned: true makes Kedro write each run's output into a timestamped subfolder, so model artifacts are never silently overwritten.
  • MemoryDataset. Any output that is not declared in the catalog is automatically held in memory as a MemoryDataset and passed to downstream nodes within the same run. This is convenient for transient intermediate results you do not need to persist.

A SQL example and a cloud example:

# Reading from a database

customertable:

type: pandas.SQLTableDataset

credentials: dbcredentials

tablename: customers

loadargs:

schema: public

Writing to cloud storage

predictions:

type: pandas.CSVDataset

filepath: s3://my-bucket/predictions/output.csv

credentials: awscreds

Credentials referenced by name (dbcredentials, awscreds) are resolved from conf/local/credentials.yml, keeping secrets out of the committed catalog.

Nodes

A node is a thin wrapper around a pure Python function. The function knows nothing about Kedro; the node maps the function's inputs and outputs to catalog dataset names.

# src/priceprediction/pipelines/dataprocessing/nodes.py

import pandas as pd

def cleanhouses(housesraw: pd.DataFrame) -> pd.DataFrame:

df = housesraw.copy()

df = df.dropna(subset=["price", "areasqm"])

df["pricepersqm"] = df["price"] / df["areasqm"]

df["hasgarage"] = df["garage"].fillna(0).astype(int)

return df

def buildmodelinput(housescleaned: pd.DataFrame) -> pd.DataFrame:

features = ["areasqm", "bedrooms", "bathrooms", "hasgarage", "pricepersqm"]

return housescleaned[features + ["price"]]

Keeping the functions pure (no I/O, no global state) means they are trivial to unit test in isolation, with no Kedro machinery involved.

Pipelines

A pipeline connects nodes into a directed acyclic graph. You do not specify the order explicitly; Kedro infers it from the inputs and outputs each node declares. If node B consumes what node A produces, Kedro runs A before B.

# src/priceprediction/pipelines/dataprocessing/pipeline.py

from kedro.pipeline import Pipeline, node, pipeline

from .nodes import cleanhouses, buildmodelinput

def createpipeline(kwargs) -> Pipeline:

return pipeline(

[

node(

func=cleanhouses,

inputs="housesraw",

outputs="housescleaned",

name="cleanhousesnode",

),

node(

func=buildmodelinput,

inputs="housescleaned",

outputs="modelinputtable",

name="buildmodelinputnode",

),

]

)

The strings housesraw, housescleaned, and modelinputtable are the same names defined in the catalog. That is the wiring: catalog names connect nodes to data and nodes to each other.

A Data Science Pipeline

Now a second pipeline for training and evaluation. Note how parameters flow in via the params: prefix.

# src/priceprediction/pipelines/datascience/nodes.py

import logging

import pandas as pd

from sklearn.ensemble import RandomForestRegressor

from sklearn.modelselection import traintestsplit

from sklearn.metrics import meanabsoluteerror, r2score

logger = logging.getLogger(name)

def splitdata(modelinputtable: pd.DataFrame, parameters: dict):

X = modelinputtable.drop(columns=["price"])

y = modelinputtable["price"]

return traintestsplit(

X, y,

testsize=parameters["testsize"],

randomstate=parameters["randomstate"],

)

def trainmodel(Xtrain, ytrain, parameters: dict) -> RandomForestRegressor:

model = RandomForestRegressor(

nestimators=parameters["nestimators"],

maxdepth=parameters["maxdepth"],

randomstate=parameters["randomstate"],

)

model.fit(Xtrain, ytrain)

return model

def evaluatemodel(model: RandomForestRegressor, Xtest, ytest) -> dict:

preds = model.predict(Xtest)

mae = meanabsoluteerror(ytest, preds)

r2 = r2score(ytest, preds)

logger.info("Model MAE: %.2f, R2: %.3f", mae, r2)

return {"mae": mae, "r2": r2}

# src/priceprediction/pipelines/datascience/pipeline.py

from kedro.pipeline import Pipeline, node, pipeline

from .nodes import splitdata, trainmodel, evaluatemodel

def createpipeline(kwargs) -> Pipeline:

return pipeline(

[

node(

func=splitdata,

inputs=["modelinputtable", "params:modeloptions"],

outputs=["Xtrain", "Xtest", "ytrain", "ytest"],

name="splitdatanode",

),

node(

func=trainmodel,

inputs=["Xtrain", "ytrain", "params:modeloptions"],

outputs="pricemodel",

name="trainmodelnode",

),

node(

func=evaluatemodel,

inputs=["pricemodel", "Xtest", "ytest"],

outputs="modelmetrics",

name="evaluatemodelnode",

),

]

)

Xtrain, Xtest, ytrain, and ytest are not in the catalog, so Kedro treats them as in-memory datasets passed between nodes. pricemodel is in the catalog, so it gets persisted (and versioned).

Registering Pipelines and the Default Pipeline

Kedro discovers pipelines through pipelineregistry.py. You assemble the full project pipeline here, including a default that runs when no specific pipeline is named.

# src/priceprediction/pipelineregistry.py

from kedro.pipeline import Pipeline

from priceprediction.pipelines import dataprocessing as dp

from priceprediction.pipelines import datascience as ds

def registerpipelines() -> dict[str, Pipeline]:

dataprocessing = dp.createpipeline()

datascience = ds.createpipeline()

return {

"dp": dataprocessing,

"ds": datascience,

"default": dataprocessing + datascience,

}

Adding two pipelines with + concatenates them into one graph. Because modelinputtable produced by the first pipeline is consumed by the second, Kedro automatically orders the whole thing correctly end to end.

Running Pipelines

The basic command runs the default pipeline:

kedro run

Kedro offers fine-grained control over what runs:

# Run only the data science pipeline

kedro run --pipeline ds

Run specific nodes by name

kedro run --nodes "trainmodelnode,evaluatemodelnode"

Run from one node onward, or up to a node

kedro run --from-nodes splitdatanode

kedro run --to-nodes buildmodelinputnode

Run everything tagged for a particular concern

kedro run --tags training

These options make iteration cheap: when you change only the evaluation logic, you can re-run from that point instead of recomputing the entire graph. Tags are attached at node definition time and group nodes across pipelines.

Parameters and Configuration

Hyperparameters and tunable settings live in conf/base/parameters.yml, never hardcoded in nodes.

# conf/base/parameters.yml

modeloptions:

testsize: 0.2

randomstate: 42

nestimators: 200

maxdepth: 12

The params:modeloptions reference in a node's inputs injects this dictionary into the function. To reference a single value you can use dotted access, for example params:modeloptions.testsize.

Configuration Environments

Kedro merges configuration from conf/base (shared) and conf/local (your machine), with local overriding base. You can create additional environments and select them at runtime:

kedro run --env staging

This loads conf/staging on top of conf/base. A typical pattern: base points at sample data, while local or staging points at a real warehouse table, with no code change.

Credentials and OmegaConf Templating

Secrets go in conf/local/credentials.yml, which is git-ignored:

# conf/local/credentials.yml

dbcredentials:

con: postgresql://user:password@host:5432/prod

awscreds:

clientkwargs:

awsaccesskeyid: AKIA...

awssecretaccesskey: ...

Kedro uses OmegaConf as its default config loader, which supports variable interpolation and reusable templates. You can define an anchor once and reuse it across the catalog:

# conf/base/catalog.yml
parquet: &parquet

type: pandas.ParquetDataset

saveargs:

compression: snappy

housescleaned:

<<: parquet

filepath: data/02intermediate/housescleaned.parquet

modelinputtable:

<<: parquet

filepath: data/05modelinput/modelinput.parquet

You can also interpolate environment variables and parameters into the catalog, which keeps configuration concise and consistent.

Visualizing the Pipeline with Kedro-Viz

Kedro-Viz renders your DAG as an interactive diagram, showing nodes, datasets, and how data flows between them.

pip install kedro-viz

kedro viz run

This opens a browser at http://127.0.0.1:4141 with the full graph. It is useful for onboarding new team members, spotting unintended dependencies, and explaining the workflow to non-engineers. Kedro-Viz can also display experiment-tracking metrics over time when you log them through the catalog.

Working Interactively: Catalog, IPython, and Jupyter

Kedro is designed to coexist with notebooks rather than ban them. Several commands bridge the two worlds:

# List and inspect declared datasets

kedro catalog list

Launch an IPython session with catalog, context, and pipelines preloaded

kedro ipython

Launch Jupyter with the Kedro session available

kedro jupyter notebook

Inside an IPython or Jupyter session you get a ready-made catalog object, so exploratory work uses the same datasets as the pipeline:

houses = catalog.load("housesraw")

houses.describe()

Save an experiment result back through the catalog

catalog.save("housescleaned", cleaneddf)

The recommended workflow is to prototype in a notebook, then move stable logic into nodes so it becomes part of the reproducible pipeline.

Extending Kedro: Hooks and Plugins

Hooks let you run custom code at defined points in the lifecycle: before or after a node runs, before or after a pipeline runs, after the catalog is created, and so on. They are the supported mechanism for cross-cutting concerns such as custom logging, data validation, or sending notifications.
# src/priceprediction/hooks.py

import logging

from kedro.framework.hooks import hookimpl

logger = logging.getLogger(name)

class ModelTrackingHooks:

@hookimpl

def afternoderun(self, node, outputs):

if node.name == "evaluatemodelnode":

metrics = outputs.get("modelmetrics", {})

logger.info("Logging metrics to tracking system: %s", metrics)

Register hooks in settings.py:

# src/priceprediction/settings.py

from priceprediction.hooks import ModelTrackingHooks

HOOKS = (ModelTrackingHooks(),)

Plugins extend the kedro command itself or add integrations. Common ones include kedro-viz, kedro-datasets, the Airflow and Argo deployment plugins, and MLflow integrations such as kedro-mlflow.

Deployment: Structure Here, Scheduling There

Kedro packages your project as a standard Python distribution:

kedro package

This produces a wheel in dist/ containing your pipelines, plus the configuration needed to run them. From there, Kedro pipelines run on whatever orchestrator you already use. The framing is consistent: Kedro defines what runs and how it is wired; the orchestrator handles when it runs, retries, and scaling.

  • Airflow. The kedro-airflow plugin converts a Kedro pipeline into an Airflow DAG, mapping each Kedro node to an Airflow task.

  pip install kedro-airflow

kedro airflow create

  • Dagster. Community and plugin integrations expose Kedro nodes as Dagster ops/assets, so the Kedro graph runs inside Dagster's scheduling and observability layer.
  • Argo Workflows / Kubeflow. Templates translate the pipeline into Argo workflow steps for Kubernetes-native execution.
  • Databricks. Kedro runs on Databricks either as a packaged job or through the Databricks workflow tooling, which suits Spark-based catalogs.

In every case you write the pipeline once in Kedro and choose a deployment target without rewriting business logic.

Experiment Tracking

For tracking runs, metrics, and parameters, Kedro integrates with MLflow through kedro-mlflow, which logs parameters, metrics, and model artifacts automatically as the pipeline runs. Kedro also has lightweight built-in experiment tracking that surfaces metrics in Kedro-Viz when you declare metric datasets in the catalog. Either approach keeps experiment records tied to the exact pipeline version that produced them.

Best Practices

  • Keep node functions pure. No file I/O, no global state, no hidden side effects. The catalog handles persistence; nodes handle logic.
  • Name everything. Give every node a name so it is addressable from the CLI and readable in Kedro-Viz.
  • One concern per pipeline. Separate data processing from data science, and keep pipelines small and composable.
  • Never commit secrets. Credentials belong in conf/local, which stays out of version control.
  • Use the data layers. Following the 01raw through 08_reporting convention makes lineage self-documenting.
  • Version your models. Set versioned: true on model and output datasets so runs are auditable.
  • Test nodes directly. Because functions are pure, unit tests need no Kedro context.
  • Promote notebook code into nodes. Explore in notebooks, but move anything that matters into the pipeline.

Conclusion and Key Takeaways

Kedro brings the structure of software engineering to data science without forcing you to abandon the exploratory style that makes the field productive. By separating code, configuration, and data, and by composing pure functions into declarative pipelines, it turns one-off notebooks into projects that any teammate can clone, run, and trust.

Key takeaways:

  • Kedro is a project framework, not a scheduler. It defines structure; tools like Airflow, Dagster, and Argo provide scheduling.
  • The Data Catalog removes hardcoded paths and centralises every dataset definition in YAML.
  • Nodes are pure functions; pipelines compose them into a graph whose order Kedro infers automatically.
  • Configuration, parameters, and credentials are separated by environment, never hardcoded.
  • Kedro-Viz visualises the DAG, and hooks/plugins extend behaviour and integrate with MLflow and deployment targets.
  • kedro package produces a portable artifact you can run on the orchestrator of your choice.

Start small: model a single pipeline with two or three nodes, declare its datasets in the catalog, and run it with kedro run. The discipline pays off as soon as a second person needs to touch the project.

Related Articles

ClearML Tutorial: Open-Source MLOps Platform for Experiment Tracking and Pipeline Automation

Tutorial ClearML: Platform MLOps Open-Source untuk Experiment Tracking dan Pipeline Automation ClearML adalah platform M...

Metaflow Tutorial: Netflix's MLOps Framework for Data Science

Tutorial Metaflow: Framework MLOps dari Netflix untuk Data Science Metaflow adalah framework open-source yang dikembangk...

Marimo Tutorial: Reactive and Reproducible Python Notebooks

Marimo: Notebook Python yang Reaktif dan Reproducible Marimo adalah notebook Python yang menyimpan isinya sebagai berkas...

ZenML: Modular and Cloud-Agnostic MLOps Pipeline Framework

ZenML: Framework Pipeline MLOps yang Modular dan Cloud-Agnostic Pendahuluan Membangun model machine learning yang akurat...