What the word deep is counting
Deep means layers, and nothing more
The depth here is a count of processing stages between input and output. It asserts nothing about insight, understanding or difficulty, and a deep model is not a profound one. Each stage transforms whatever the stage below it produced, and stacking them lets a network build up its description in instalments: crude edges and textures nearest the input, parts and arrangements in the middle, the answer at the far end. A model with a couple of hidden stages is shallow, one with many is deep, and the line between them is a matter of custom rather than definition. How a single layer works belongs to the neural network entry; this one is about what changes when you stack them.
Stacking matters because of what it removes from the job. Classical machine learning required somebody to decide, before any training began, which properties of the raw input the model would be allowed to see: which sound frequencies, which pixel statistics, which ratios between columns. That step held most of the expertise, and it was also where projects died, since a property nobody thought to compute is a property the model can never use. A deep network is instead trained from raw input all the way through to the final answer under a single error signal, so its intermediate descriptions are learned rather than specified, and they are fitted to the task at hand rather than to whatever struck an analyst as sensible beforehand. That is the trick in full, and it explains why the approach swept problems whose rules nobody could write down: vision, speech, language, protein structure, game play. It also explains the bill. Working the properties out instead of stating them takes far more examples and far more arithmetic than a model handed good ones, which is why the idea waited decades for parallel hardware and web-scale datasets, and why an experienced practitioner still reaches for something smaller when the data is a table of tidy, meaningful columns. Two further costs come with the territory. Interpretability suffers, because those learned intermediate descriptions are vectors nobody requested and nobody can read, which is a genuine obstacle in credit, medicine and anywhere a decision has to withstand a regulator's question. And they encode whatever the training data encoded, skews included, which is the mechanism behind much of what gets filed under responsible AI. The architecture families differ by the shape of the data rather than by philosophy: convolutions where position carries meaning, as in computer vision, and Transformers where any part of a sequence may bear on any other, the design sitting under current generative AI.
