Moving data is not the job. Moving data is a solved problem with vendors attached to it. The job is keeping a figure correct while everything beneath it changes without notice: a checkout is redeployed, a second currency appears, a vendor retires an endpoint, a well-meaning colleague corrects a typo in a category label that forty reports group by. None of those events are visible from the warehouse and none of them raise an error, which is the defining property of failure in this field. Bad data does not crash. It renders. A dashboard fed by a pipeline that stopped running last Tuesday looks exactly like a dashboard fed by one that is working, and it will go on looking that way until somebody who knows the business notices that a number is implausible. So the deliverable is not a pipeline, it is a warranty, and everything the role is known for follows from that. Loads are made idempotent because a retry must not double the rows. Incremental logic needs a deliberate answer for deletions, because most incremental strategies retain records the source has thrown away. Tests assert that a key is unique and a reference resolves, on every run rather than on the day someone complains. Alerting watches freshness, which catches the most common outage of all: nothing happening. The stack itself has converged further than job adverts suggest. A managed connector lands raw tables in a warehouse such as BigQuery, Snowflake or Redshift, dbt turns those into modelled tables the rest of the company is allowed to read, a scheduler triggers the sequence and retries it, and Python fills the gaps where no tool fits. Most of the transformation is written in SQL, not in Python, which surprises people arriving from a course.
Which stage breaks, and whose problem it is
| Stage | How it fails | Who owns the fix |
|---|---|---|
| Ingestion from source systems | A schema changes upstream, an API starts throttling, or an incremental sync keeps rows the source has deleted. | Data engineer, and realistically only if somebody warned them the source was moving. |
| Landing in the warehouse | A retried run duplicates yesterday's rows; timestamps arrive in whatever zone the source felt like using. | Data engineer. |
| Transformation into modelled tables | Two departments count revenue differently and both are technically correct. | An analytics engineer writes it, a business owner decides it. Leaving the definition unowned is the most expensive failure on this list. |
| Orchestration and scheduling | The run never happened and nothing announced it. | Data engineer or platform team. |
| Serving to reporting tools and models | Stale figures presented as current, because the report has no way of showing its own age. | Shared: the engineer publishes the freshness, the analyst has to put it on the screen. |
| Testing and monitoring | Nobody finds out for a week, then everybody finds out at once. | Nominally shared, in practice nobody, which is exactly why it should be assigned by name. |
