Skip to content
Getting Digital

Data Engineering

Also: data pipelines, data infrastructure

Data engineering is the discipline of moving records from the systems that produce them into a store where they can be queried, and of keeping that supply correct while the sources underneath keep changing.

Assessment. Almost no team should be writing its own connectors. Adopt managed ingestion, put the saved effort into the transformation layer where a company's numbers are decided, and introduce an orchestrator only when something breaks that a schedule and a retry cannot fix.

Moving data is not the job. Moving data is a solved problem with vendors attached to it. The job is keeping a figure correct while everything beneath it changes without notice: a checkout is redeployed, a second currency appears, a vendor retires an endpoint, a well-meaning colleague corrects a typo in a category label that forty reports group by. None of those events are visible from the warehouse and none of them raise an error, which is the defining property of failure in this field. Bad data does not crash. It renders. A dashboard fed by a pipeline that stopped running last Tuesday looks exactly like a dashboard fed by one that is working, and it will go on looking that way until somebody who knows the business notices that a number is implausible. So the deliverable is not a pipeline, it is a warranty, and everything the role is known for follows from that. Loads are made idempotent because a retry must not double the rows. Incremental logic needs a deliberate answer for deletions, because most incremental strategies retain records the source has thrown away. Tests assert that a key is unique and a reference resolves, on every run rather than on the day someone complains. Alerting watches freshness, which catches the most common outage of all: nothing happening. The stack itself has converged further than job adverts suggest. A managed connector lands raw tables in a warehouse such as BigQuery, Snowflake or Redshift, dbt turns those into modelled tables the rest of the company is allowed to read, a scheduler triggers the sequence and retries it, and Python fills the gaps where no tool fits. Most of the transformation is written in SQL, not in Python, which surprises people arriving from a course.

Which stage breaks, and whose problem it is

StageHow it failsWho owns the fix
Ingestion from source systemsA schema changes upstream, an API starts throttling, or an incremental sync keeps rows the source has deleted.Data engineer, and realistically only if somebody warned them the source was moving.
Landing in the warehouseA retried run duplicates yesterday's rows; timestamps arrive in whatever zone the source felt like using.Data engineer.
Transformation into modelled tablesTwo departments count revenue differently and both are technically correct.An analytics engineer writes it, a business owner decides it. Leaving the definition unowned is the most expensive failure on this list.
Orchestration and schedulingThe run never happened and nothing announced it.Data engineer or platform team.
Serving to reporting tools and modelsStale figures presented as current, because the report has no way of showing its own age.Shared: the engineer publishes the freshness, the analyst has to put it on the screen.
Testing and monitoringNobody finds out for a week, then everybody finds out at once.Nominally shared, in practice nobody, which is exactly why it should be assigned by name.

In practice

A team syncs its CRM into BigQuery with a managed connector, models it in dbt and reports from Power BI. The connector runs incrementally: each pass asks the source for records whose updated timestamp is later than the previous pass. Sales then deletes a duplicate account. The deletion leaves no row, therefore no updated timestamp, therefore nothing for the next sync to notice, and the dead account sits in the warehouse indefinitely, inflating every count built on that table. The available fixes are all deliberate choices: ask the source to mark records as deleted instead of removing them, reconcile periodically against a full extract, or read the database's change log rather than polling a column. Picking one of those is data engineering. Discovering the problem six months later, from a figure that never once looked wrong, is what the discipline exists to prevent.

Often confused with

Data Analysis
An engineer is accountable for whether the table is correct and current. An analyst is accountable for what follows from it. The skills overlap heavily; the questions asked at a post-mortem do not.
Data Wrangling
Same operations, different lifespan. Wrangling is repair done once for one question; engineering is the automated version that has to survive the source changing shape.
MLOps
The same instinct applied to models rather than tables: versioning, scheduled runs, monitoring and a rollback path, except that the artefact being kept healthy is a trained model.

Key takeaways

  • →Bad data does not raise an error, it renders normally, so detection has to be designed in rather than waited for.
  • →The tooling has converged on managed ingestion, a cloud warehouse and a SQL transformation layer; the difficulty has moved to the definitions.
  • →Assign freshness alerts and data tests to a named person, or they belong to nobody and fire into an empty room.

Related concepts

  • RAG sits on top of well-built data infrastructure.

  • Data engineering is the industrial-scale form of wrangling.

  • RelatedMLOps

    Every production model sits downstream of data pipelines; the two disciplines share most of their plumbing.

  • IncludesETL and ELT

    The central pattern of the discipline.

Courses that teach this

Where this concept sits in the field

Certifications that test this

Vendor exams whose syllabus covers this concept: facts, cost and a preparation path on each page.

More courses from these categories

Courses from the categories where this concept is taught. Details, price and the provider link are on each course page.

FAQ

Do I need to learn Airflow?
Not first, and possibly not at all. A managed connector plus dbt plus whatever scheduling your warehouse or your dbt platform already offers covers a great many companies completely. Reach for a real orchestrator when you have genuine dependencies between jobs that a fixed timetable cannot express, or when a failure halfway through needs to resume rather than restart. Adopting one before that buys you a second system to keep alive.
Warehouse or lakehouse? Does the choice matter?
For most companies, far less than the debate suggests. The big cloud warehouses and the table formats underneath lakehouses have converged on similar capabilities, and the sane default is whichever one your existing cloud contract already covers, because the integration work is what costs you. The choice that does matter is how disciplined your modelling layer is, and no vendor sells that.

Sources

The primary text this definition rests on. Read it before relying on this one.

Last reviewed 26 September 2026 · Getting Digital