Match the question to the method
| What is being asked | The method it calls for | Usual failure |
|---|---|---|
| How much, where, compared with when? | Aggregate and compare against a prior period or a peer group. One query against the warehouse, no modelling. | A headline figure that conceals two opposite movements underneath it. |
| Is this difference real? | Quantify the uncertainty before interpreting the gap: an interval around each figure, or a test if the comparison was designed beforehand. | Watching a live experiment until it crosses the line you were hoping for, then stopping. |
| Why did the number move? | Decompose. Split the total by segment, channel or cohort until the movement localises in one place. | Reading a shift in the mix as a change in behaviour when only the weights moved. |
| Did the change cause it? | A controlled experiment, or failing that a comparison with a group the change never reached. | Treating before-and-after as causal on records collected for some other purpose. |
| What happens next? | A forecast or a fitted model, scored on periods deliberately held back from fitting. | Reporting how well it reproduces data it has already seen. |
Every row in that table begins with a question, which is the step most often skipped. Work that starts from a dataset and goes hunting for something interesting will always find something, because any sufficiently wide table contains a coincidence and the person hunting has no way to separate the coincidence from the finding. Fixing the question first sets the standard of proof in advance: you decide what would count as an answer, and what would count as not enough, before you know which way the evidence points. The tooling then follows the question rather than the reverse. Aggregation and comparison belong where the records already sit, because one query can do the work; pulling a million rows onto a laptop to group them is a habit worth losing early. Python becomes worth the trouble once one query can no longer express the job: reshaping between wide and long, joining against something the warehouse has never heard of, a simulation, an interval you want to compute yourself. What separates analysis from machine learning is the obligation to explain. A model may be a black box provided it predicts well on records it was never shown. An analysis may not, because somebody will act on it and somebody else will contest it. That obligation shapes the method: it is why the simpler estimate that can be checked by hand usually wins, why the working is kept rather than only the conclusion, and why the most valuable sentence in a report is often the one that narrows the claim to accounts opened after the migration and admits that the earlier ones cannot be spoken for.
