dlt Tutorial: Python-First Data Ingestion Pipelines
# Membangun Pipeline EL Berbasis Python dengan dlt (data load tool)
Sebagian besar tim data menghabiskan waktu yang tidak sedikit pada bagian pipeline yang kurang menarik: menarik data dari sebuah AP...
Building Python-First EL Pipelines with dlt (data load tool)
Most data teams spend a surprising amount of time on the unglamorous part of the pipeline: pulling data out of an API and landing it in a warehouse with a sane schema. dlt (short for data load tool, from dltHub) is a lightweight Python library built specifically for that job. This tutorial walks through how it works and how it fits alongside the dbt and Dagster tooling you may already be using.
A quick note on naming, because the abbreviation is overloaded: the dlt discussed here is the open-source Python ingestion library from dltHub. It is not Delta Lake (the storage format) and not PyTorch's DLT. When we say dlt, we mean pip install dlt.
Where dlt Fits in the Modern Data Stack
The modern analytics stack is commonly described as EL + T: Extract and Load raw data first, then Transform it in the warehouse.
Extract + Load is dlt's territory. It reads from a source (REST API, database, file, generator) and writes the data into a destination warehouse, handling schema inference, normalization, and incremental state for you.
Transform is where dbt lives. Once raw tables land in the warehouse, dbt models reshape them into clean, tested, documented marts.
Orchestration is where Dagster (or Airflow, cron, GitHub Actions) lives. It decides when the dlt extract-load runs and when the dbt transform runs, and wires up the dependency between them.
So dlt does not replace dbt or Dagster. It fills the gap they deliberately leave open: getting raw data into the warehouse reliably, in plain Python, without standing up a heavy ingestion platform. Unlike connector platforms such as Fivetran or Airbyte — great when a connector exists, awkward when it does not — dlt lets you write a normal Python function that yields records, decorate it, and get a robust, schema-aware, incremental pipeline. If you can call an API in Python, you can load it with dlt.
Installation and Project Setup
dlt runs anywhere Python 3.8+ runs. Install the core package plus the extras for your destination.
# Core library
pip install dlt
With a specific destination's dependencies
pip install "dlt[duckdb]" # local analytics, great for development
pip install "dlt[bigquery]" # Google BigQuery
pip install "dlt[postgres]" # PostgreSQL
pip install "dlt[snowflake]" # Snowflake
It is good practice to work inside a virtual environment.
The .dlt/secrets.toml file is added to .gitignore automatically. Never commit it.
Core Concepts: Resource, Source, Pipeline
dlt has three building blocks worth internalizing before writing code.
Resource (@dlt.resource): a function that yields data — typically rows or pages of records. Each resource maps to one table in the destination.
Source (@dlt.source): a function that groups several related resources together (for example, all endpoints of one API). It returns the resources it manages.
Pipeline (dlt.pipeline(...)): the object that connects a source to a destination, runs the load, and tracks state (schema, incremental cursors, run history).
A First Pipeline: A Paginated REST API into DuckDB
Let's load issues from a paginated REST API into DuckDB. We'll use GitHub's REST API as the running example because its pagination is representative of most APIs. Note how the resource, source-style grouping, and pipeline come together: dlt creates the DuckDB file, infers the schema from the data, and writes the rows with no DDL or manual table creation.
yield page # yield the whole page; dlt flattens it
params["page"] += 1
pipeline = dlt.pipeline(
pipelinename="github",
destination="duckdb",
datasetname="githubdata",
)
loadinfo = pipeline.run(githubissues())
print(loadinfo)
A few things worth noting:
dlt.sources.helpers.requests is a thin wrapper around requests that adds sensible retries and timeouts. You can use plain requests too.
We yield an entire page (a list). dlt automatically iterates the list, so each element becomes a row.
writedisposition="replace" means each run rebuilds the table from scratch. We'll improve this with incremental loading later.
Inspecting the result
After a run, inspect what happened through the pipeline object (pipeline.lasttrace for run metrics, pipeline.lasttrace.lastnormalizeinfo for per-step detail) and query the loaded data via SQL.
print(pipeline.lasttrace)
with pipeline.sql
client() as client:
with client.executequery("SELECT count() AS n FROM issues") as cursor:
print(cursor.fetchall())
For DuckDB you can also open the generated .duckdb file directly in any DuckDB client and run SQL such as SELECT id, title, state FROM githubdata.issues ORDER BY createdat DESC LIMIT 10;.
Automatic Schema Inference, Normalization, and Evolution
This is where dlt earns its keep.
Inference
dlt reads your records, infers column names and types, and creates the destination table for you. Python dict keys become columns; types are inferred from the values (with sensible coercion rules for dates, decimals, and so on).
Normalization of nested JSON
API responses are rarely flat. A GitHub issue contains a nested user object and a labels array. dlt normalizes these automatically:
Nested objects are flattened into prefixed columns, e.g. user.login becomes userlogin.
Nested lists are unnested into child tables linked back to the parent by a generated foreign key.
So an issues payload with a labels array produces both an issues table and an issueslabels child table, linked by a generated key.
SELECT l.name, i.title
FROM issueslabels l
JOIN issues i ON l.dltparentid = i.dltid;
dlt adds bookkeeping columns prefixed with dlt (dltid, dltloadid, dltparentid) to make these relationships and load batches traceable.
Evolution
When the source adds a new field, dlt detects it on the next run and adds the column to the destination table automatically — without breaking the load. Removed fields simply produce NULLs for new rows. This schema evolution is what makes dlt resilient to upstream API changes that would otherwise break a hand-rolled loader.
You can inspect the inferred schema at any time.
print(pipeline.defaultschema.toprettyyaml())
Write Dispositions: replace, append, merge
The writedisposition controls how new data interacts with what is already in the table.
# replace: drop and reload the table every run (full refresh)
@dlt.resource(writedisposition="replace")
def dimcountry(): ...
append: add new rows, never touch existing ones (good for immutable events/logs)
@dlt.resource(writedisposition="append")
def pageviews(): ...
merge: upsert — update existing rows by key, insert new ones
replace is simplest and correct for small dimension tables you can afford to rebuild.
append is for facts that never change, like event logs. Beware of duplicates on re-runs.
merge is the upsert. You declare a primarykey (or a mergekey); dlt deletes matching rows and inserts the incoming batch, so re-running is idempotent.
Use mergekey instead of primarykey when the natural key spans multiple columns or when you want to control the delete predicate independently of the row identity.
Incremental Loading with dlt.sources.incremental
Full reloads do not scale. Incremental loading fetches only records that changed since the last run, using a cursor field (a monotonically increasing column such as updatedat or id).
"since": updatedat.lastvalue, # ask the API for only new data
"sort": "updated",
"direction": "asc",
}
while True:
response = requests.get(url, params=params)
response.raiseforstatus()
page = response.json()
if not page:
break
yield page
params["page"] += 1
How this behaves:
On the first run, updatedat.lastvalue equals initialvalue. dlt loads everything from that point.
dlt stores the maximum updatedat it saw in the pipeline state.
On the next run, lastvalue is that stored maximum, so you fetch only newer records.
Combined with writedisposition="merge" and primarykey="id", records that were updated upstream are upserted rather than duplicated.
dlt also deduplicates rows at the boundary value (records with exactly lastvalue) so you do not double-load the edge. State is persisted in the destination, so it survives across separate process runs — no external state store required.
Config-Driven Extraction with the REST API Source
Writing pagination loops by hand gets repetitive. dlt ships a declarative restapisource helper that describes an API as configuration. It handles pagination, auth, and incremental wiring for you.
The declarative form is the recommended starting point for typical REST APIs; drop down to a hand-written @dlt.resource only when an API does something unusual.
Secrets and Configuration
dlt keeps configuration and secrets out of your code. It resolves values from .dlt/secrets.toml, .dlt/config.toml, and environment variables, in that order of specificity.
Reference a secret in code with dlt.secrets["..."] or, more idiomatically, by giving your function a typed argument defaulting to dlt.secrets.value — dlt injects the resolved value automatically at call time.
By default dlt evolves the schema freely, which is convenient in development but risky in production. The schemacontract setting lets you constrain that behavior per table, column, and data type.
@dlt.resource(
schemacontract={
"tables": "evolve", # allow new tables
"columns": "freeze", # reject new columns -> raises on unexpected fields
"datatype": "freeze" # reject type changes
}
)
def orders():
...
The available modes are evolve (allow, the default), freeze (raise an error), discardrow (drop the offending row), and discardvalue (drop the offending value but keep the row). A common production pattern is evolve columns but freeze data types, so additive changes pass while type drift is caught early.
Transformers and Parallelism
A @dlt.transformer consumes a resource's output and produces a derived resource — useful for fan-out, such as fetching details for each ID returned by a list endpoint.
For throughput, mark a resource parallelized=True and tune worker counts per stage (extract, normalize, load) in .dlt/config.toml.
@dlt.resource(parallelized=True)
def heavyendpoint():
...
Start with the defaults and increase workers only when a stage is measurably the bottleneck.
Running and Deploying
A dlt pipeline is just a Python script, so deployment is whatever runs Python on a schedule.
Cron — the simplest option: 0 .venv/bin/python githubpipeline.py.
GitHub Actions — dlt deploy githubpipeline.py github-action --schedule "0 *" scaffolds a workflow and lists the repository secrets to set (it never embeds credentials).
Airflow / Dagster — call the pipeline from inside a task or asset. The common Dagster pattern: an asset runs the dlt extract-load, a downstream dbt asset transforms it, and Dagster manages the dependency and scheduling.
This is exactly the seam where dlt hands off to dbt: dlt lands raw, dbt builds the marts on top.
Best Practices
Develop against DuckDB, deploy to your warehouse. DuckDB needs no infrastructure and runs locally, so iterate fast there, then change only the destination argument for production.
Prefer merge + incremental for anything large. Full replace is fine for small dimensions but wasteful and slow for big, growing tables.
Pick a real cursor field. Use a reliably increasing column (updatedat, an auto-increment id). Avoid fields the source can backfill out of order.
Keep secrets in secrets.toml or env vars. Never hardcode tokens; ensure .dlt/secrets.toml stays gitignored.
Freeze schema in production where it matters. Use schemacontract to catch unexpected upstream changes instead of silently absorbing them.
Inspect lasttrace in CI. Logging the trace and row counts makes failures and silent zero-row loads easy to spot.
Let dlt own EL, dbt own T. Resist transforming inside the resource. Land raw, transform in dbt — it keeps lineage clean and reprocessing cheap.
Conclusion and Key Takeaways
dlt gives data teams a Python-native way to handle the Extract-and-Load half of the modern stack without adopting a heavyweight connector platform. Its automatic schema inference, JSON normalization into child tables, and schema evolution remove most of the brittle plumbing that hand-written ingestion scripts accumulate.
Key takeaways:
dlt (dltHub's data load tool) is a Python library for EL; it complements dbt (T) and Dagster (orchestration), and is unrelated to Delta Lake or PyTorch.
The three core abstractions are @dlt.resource, @dlt.source, and dlt.pipeline with a destination.
Write dispositions (replace, append, merge) plus dlt.sources.incremental give you idempotent, incremental loads with a primarykey.
The declarative restapisource handles pagination, auth, and incremental wiring with configuration instead of loops.
Secrets live in .dlt/secrets.toml or environment variables; schemacontract enforces discipline in production.
Because a pipeline is plain Python, you deploy it with cron, GitHub Actions, Airflow, or Dagster — handing raw data off to dbt for transformation.
Start with pip install "dlt[duckdb]" and a paginated API you already know; you will have a working, incremental, schema-aware pipeline in well under an hour.