Skip to content
Getting Digital

Deep Learning

Also: deep neural networks

Deep learning is machine learning built on neural networks of many stacked layers, which work out for themselves which properties of the raw data matter instead of being handed a list by a person.

Assessment. Depth takes the credit and end-to-end training deserves it: one error signal tunes the feature extractor and the predictor together, so nobody has to guess in advance what the model should be looking at. Many teams reaching for it on ordinary business tables are paying in compute and opacity for accuracy a boosted tree would have handed them on a laptop.

What the word deep is counting

Deep means layers, and nothing more

The depth here is a count of processing stages between input and output. It asserts nothing about insight, understanding or difficulty, and a deep model is not a profound one. Each stage transforms whatever the stage below it produced, and stacking them lets a network build up its description in instalments: crude edges and textures nearest the input, parts and arrangements in the middle, the answer at the far end. A model with a couple of hidden stages is shallow, one with many is deep, and the line between them is a matter of custom rather than definition. How a single layer works belongs to the neural network entry; this one is about what changes when you stack them.

Stacking matters because of what it removes from the job. Classical machine learning required somebody to decide, before any training began, which properties of the raw input the model would be allowed to see: which sound frequencies, which pixel statistics, which ratios between columns. That step held most of the expertise, and it was also where projects died, since a property nobody thought to compute is a property the model can never use. A deep network is instead trained from raw input all the way through to the final answer under a single error signal, so its intermediate descriptions are learned rather than specified, and they are fitted to the task at hand rather than to whatever struck an analyst as sensible beforehand. That is the trick in full, and it explains why the approach swept problems whose rules nobody could write down: vision, speech, language, protein structure, game play. It also explains the bill. Working the properties out instead of stating them takes far more examples and far more arithmetic than a model handed good ones, which is why the idea waited decades for parallel hardware and web-scale datasets, and why an experienced practitioner still reaches for something smaller when the data is a table of tidy, meaningful columns. Two further costs come with the territory. Interpretability suffers, because those learned intermediate descriptions are vectors nobody requested and nobody can read, which is a genuine obstacle in credit, medicine and anywhere a decision has to withstand a regulator's question. And they encode whatever the training data encoded, skews included, which is the mechanism behind much of what gets filed under responsible AI. The architecture families differ by the shape of the data rather than by philosophy: convolutions where position carries meaning, as in computer vision, and Transformers where any part of a sequence may bear on any other, the design sitting under current generative AI.

In practice

Two projects at one company; only one wants a deep network. The first predicts which subscribers will cancel next month from a table of tenure, plan, payment history and support contacts, columns a team curated over years. Boosted trees fit that in minutes on a laptop, arrive with a usable account of which column drove each prediction, and are hard to beat, because the properties are already good and there is no hidden structure inside a spreadsheet cell for extra stages to discover.

The second reads handwritten delivery notes photographed by drivers. Nobody can write down what makes a pen stroke a seven; there is no sensible column for the loop of a letter. Here the deep network is not the fashionable choice, it is the only choice that works, and the price of it, a large annotated corpus plus hardware to match, is simply what a problem whose rules cannot be stated happens to cost.

Often confused with

Machine Learning
Deep learning is a subset. Every deep model is machine learning, while most machine learning running in production is not deep, because scoring, ranking and forecasting over tables rarely need it.
Neural Networks
A neural network is the object; deep is an adjective describing how many layers it has. The two are not alternatives: the live question is how deep, never which of the pair.

Key takeaways

  • →Deep counts layers between input and output. The word describes shape, and claims nothing about understanding.
  • →The real gain is end-to-end training: the network learns what to look at, which is why it took over problems whose rules nobody could write down.
  • →You pay in data, compute and opacity, so tidy tabular problems are usually better served by simpler models.

Related concepts

  • Broader topicMachine Learning

    Deep learning is a sub-field of machine learning.

  • Neural networks are the building blocks deep learning is made of.

  • More specificGenerative AI

    Generative AI is built on deep-learning models.

  • More specificTransformers

    The Transformer is a deep-learning architecture.

  • More specificComputer Vision

    Computer vision is a flagship deep-learning application.

Courses that teach this

Where this concept sits in the field

Certifications that test this

Vendor exams whose syllabus covers this concept: facts, cost and a preparation path on each page.

FAQ

How many layers make a network deep?
There is no threshold and arguing towards one wastes an afternoon. The distinction that pays is whether the model works out its own intermediate representations or receives them from a person. A single hidden layer barely manages the former; a stack of them does.
Does it need a GPU?
For training anything substantial, effectively yes, because the same operation runs across enormous arrays and that is exactly what parallel hardware exists for. Running a model that is already trained is a different proposition: plenty are small enough for a phone or a browser tab, which is how on-device dictation and photo effects work with the network switched off.

Sources

The primary text this definition rests on. Read it before relying on this one.

Last reviewed 26 September 2026 · Getting Digital