Skip to content
Getting Digital

Data & Analytics

The field is not a job title, it is a sequence. What happens between a source system and a decision, which stages of it employers are hiring for, and where to start in order to be one of those hires.

14 concepts · 4 courses · 26 certification exams · 4 related fields

Data work is sold as a subject and practised as a sequence. Records are produced by systems nobody in the data team controls, moved somewhere they can be queried, reconciled against each other, interrogated, drawn as a picture, and then made to happen again next week without anyone watching. Every real data job sits at one or two points on that sequence, and the points are not interchangeable. The uncomfortable thing about the field is that learners and employers have settled on different ends of it. Course catalogues are dense at the late stages, where the work photographs well: a model that predicts churn, a chart that finally changes a board's mind. Job adverts are dense at the early stages, where the work does not photograph at all. They ask whether you can pull years of CRM exports into a warehouse without dropping the deleted rows unnoticed, and whether you can explain why finance's revenue figure and marketing's revenue figure have never once agreed. That gap is not a failure of taste on anybody's part. It follows from where organisations feel pain. A company without trustworthy tables cannot use machine learning at all, so it hires for tables first, and it keeps hiring, because pipelines break in ways models do not: a source system renames a field, a time zone shifts, a vendor changes an export format, and every number downstream is wrong while still looking entirely plausible. Data visualization sits at the far end of that chain and inherits every error made before it.

What happens between a source system and a decision

  1. Production. The data is created by software the data team does not own: a CRM, a billing system, a checkout, a set of event trackers. None of it was designed for analysis, and that is not negligence, it is the point. Whoever built the checkout optimised for taking money. The job at this stage is not to change anything but to learn how the system behaves, because every quirk not learnt here reappears later disguised as an insight.
  2. Movement. Getting records out of those systems and into somewhere queryable is data engineering, and it is where a large share of the open roles live. It is connectors, schedules, retries, backfills, and a warehouse that has to stay affordable while it grows. Almost nobody enrols in a course about backfills. Almost every team that has tried to run analysis directly against a production database has learned, painfully, why the warehouse exists.
  3. Modelling in the warehouse. Raw tables are not answers. Someone has to turn them into a layer everyone agrees on, usually with dbt or an equivalent, and decide what an active customer is, when a subscription counts as churned, which refunds reverse which orders. This is a definitional job wearing an engineering costume, and it carries more leverage than any other stage, because every figure produced downstream inherits the decision taken here.
  4. Wrangling. Data wrangling is the reshaping that happens after you have the data and before you can ask anything of it: joins that do not line up, dates in four formats, nulls that mean three different things, and a spreadsheet somebody maintains by hand that is nonetheless the only record of which accounts are internal. This stage rarely appears in a syllabus and reliably appears in a Monday.
  5. Analysis. Data analysis is the part people picture: the question, the query, the comparison, the answer the analyst is prepared to defend when someone senior dislikes it. SQL does most of it. Python with pandas does the rest, mostly when the shape of the problem outgrows a single query. The scarce skill here is not technique, it is knowing which difference is real and which is noise.
  6. Prediction. Fitting a model and expecting it to hold on data it has not seen. This is the stage with the most courses, the most enthusiasm and the fewest junior openings, for an unremarkable reason: a company needs a working pipeline before it needs a forecast, and plenty of companies are still building the pipeline. Worth learning, after one can be trusted with a number.
  7. Presentation, and then again next week. A chart, a report, a dashboard in Power BI or Tableau that someone opens on a Monday morning and acts on. The presentation half is a writing skill more than a charting one. The next-week half is engineering again, and it decides whether an analysis becomes a standing input to decisions or a document nobody reopens.

Start where the enrolment isn't

Learn stages two and three first. Concretely: SQL until it is boring, a warehouse you have loaded yourself even if it is a free tier holding your own bank statements, and enough dbt to understand why a model layer exists at all. Then a spreadsheet at a level most people never reach, because business data still arrives as one and findings still leave as one. Then Python and pandas. Then a single dashboard tool, chosen by what your target employers already run rather than by which looks best, which in practice means Power BI inside Microsoft organisations and Tableau across much of the rest. What people who skip those stages get wrong is specific and recognisable. They can produce a number but cannot defend it, because they never saw where it came from. Asked why the figure moved, they re-run the notebook. Asked whether the join dropped rows, they do not know, and frequently do not know that this is a question at all. Their portfolio opens with a clean CSV, which is the single condition the job never meets. That last point deserves an argument rather than an assertion. Practitioners spend most of their time on data quality and courses spend almost none, and neither fact is anyone's fault. Errors enter at every stage before analysis and none of them announce themselves: a model trains perfectly well on wrong inputs and returns a confident answer, a chart renders a corrupted join without complaint. The only available defence is checking, and checking is most of the week. But a teaching dataset has to be clean to be teachable. You cannot set a reproducible exercise on a mess whose shape you do not control, so the exercise arrives tidy, the student rehearses the tidy part, and the ratio between what the course drills and what the job demands is inverted before anybody has done anything wrong. Knowing that in advance is worth more than another badge. If you want badges anyway, PL-300 certifies the reporting stage and DP-700 the engineering one, and neither of them proves you can find the bad join.

The field is mapped topic by topic, with its certifications and guides, at Data, Analytics and AI.

Concepts

Data & Analytics

Data Analysis

Data analysis is the practice of interrogating data to answer a specific question, and of establishing how much weight the answer can bear.

Data & Analytics

Data Wrangling

Data wrangling is the work of repairing and reshaping real records, and of deciding what their gaps and inconsistencies mean, until a table can be trusted with the question being asked of it.

Data & Analytics

Data Visualization

Data visualisation is the practice of encoding numbers as position, length and colour so that a comparison a reader would otherwise have to calculate becomes something they can simply see.

Data & Analytics

Data Engineering

Data engineering is the discipline of moving records from the systems that produce them into a store where they can be queried, and of keeping that supply correct while the sources underneath keep changing.

Data & Analytics

SQL

SQL is the declarative language for asking questions of relational databases: you state which rows and columns you want and how tables relate, and the database engine works out how to retrieve them.

Data & Analytics

ETL and ELT

ETL is the pattern for moving data from the systems that produce it into the one place it is analysed: extract the records from each source, transform them into a common, cleaned shape, and load them into a warehouse; ELT keeps the same three steps and runs the transformation inside the warehouse after loading.

Data & Analytics

Data Warehouse

A data warehouse is a database built for analysis rather than for running the business: it holds copies of records from many operational systems, integrated under one set of definitions, kept as history rather than overwritten, and arranged so that questions across years and departments run fast.

Data & Analytics

Data Lake

A data lake is a store that keeps data in its original form, files of any shape in cheap object storage, loaded before anyone has decided what questions will be asked of it, so that the structure is applied when the data is read rather than when it is written.

Data & Analytics

Business Intelligence (BI)

Business intelligence is the practice and the tooling for turning an organisation's data into reports and dashboards that people use to run it: connecting to the sources, modelling the data into measures everyone agrees on, and presenting it so that a manager sees the state of the business without asking an analyst.

Data & Analytics

Data Governance

Data governance is the set of decisions and responsibilities an organisation puts in place over its data: who owns each dataset, what each field means, who may read or change it, how long it is kept and how its quality is measured, recorded so that the answers do not depend on who is asked.

Data & Analytics

Data Quality

Data quality is the degree to which a dataset is fit for the purpose it is used for, assessed along named dimensions, whether the records are complete, unique, consistent, timely, valid and accurate, and measured against the requirements of the decision the data supports rather than against an absolute standard.

Data & Analytics

Data Modelling

Data modelling is the design of how data is structured for a given use: which entities are recorded, which attributes describe them, how the tables relate through keys, and, for analytical use, which table holds the measurements and which hold the context, decided before the data is loaded so that the questions the model must answer run correctly and fast.

Data & Analytics

Hypothesis Testing

Hypothesis testing is the statistical procedure for deciding whether an observed difference, between two groups, two periods or a sample and a claimed value, is larger than chance variation would produce if there were no real difference, by computing how probable a result at least that extreme would be under that assumption.

Programming & Web Development

Python

Python is a general-purpose programming language that became the standard interface for driving numerical and machine-learning code written in other languages.

Courses in this topic

Related fields

FAQ

I already know Python. Do I still start with SQL?
Yes, and the reason is not that SQL is the better language. It is that the warehouse is where the definitions live, and definitions are what you will be asked to defend. Python transfers instantly once you can read the tables properly; the reverse transfer is slower and more painful. If you know Python already, the fastest upgrade is not more Python, it is loading a real dataset into a warehouse yourself and modelling it.
Is a BI tool worth learning, or is it a commodity?
Both. Power BI and Tableau are similar enough that the second takes days once you know the first, so the tool itself is a commodity and not worth agonising over. What is not a commodity is deciding what belongs on the screen and what does not, then writing the sentence beside the chart that says what to do about it. Dashboards fail far more often from carrying too much than from running on the wrong software.
Analyst or data engineer?
Engineer if you would rather the system ran itself and nobody had to ask you anything. Analyst if you enjoy the argument about what the numbers mean and want to sit close to the decision. The boundary has moved, though: many jobs advertised as analyst roles are mostly warehouse modelling, and many engineering roles are mostly cleaning. Read the tasks in the advert rather than the title on it.

Last reviewed 26 September 2026 · Getting Digital