Skip to content
Getting Digital

Transformers

Also: transformer architecture, attention model

A transformer is a neural-network design that takes a whole sequence in at once and, for every element in it, works out how strongly each of the other elements should influence that one.

Assessment. Learning the block diagram is wasted effort for anyone who will never train a model, and most introductions start there anyway. Learn what attention bought instead, because that single property explains the grasp of context, the absence of memory between requests, and why the bill looks the way it does.

The mechanism doing the work is called attention, and the clearest way to see what it buys is to look at what it displaced. The previous generation of sequence models read from left to right, carrying a running summary forward one position at a time. Two consequences followed. First, that summary is a bottleneck: information from the opening of a long passage has to survive being compressed and recompressed at every subsequent step to still be available at the end, and in practice much of it did not. Second, position ten could not be computed until position nine had finished, so the work refused to spread across a graphics card, which put a hard ceiling on how much text anyone could afford to train on. Attention dissolves both at once. Every position is compared directly with every other position, so a pronoun near the end of a paragraph consults the noun it refers to at the beginning without anything in between wearing the connection down. And because those comparisons do not depend on each other, they all run at the same time. The first change is why these models handle context as well as they do. The second is why they could be scaled at all, and scale is what produced the behaviour that made everyone pay attention in the first place.

The cost hiding inside the idea

Comparing every position with every other position means the work grows with the square of the input length, not in step with it. That one property explains a surprising amount of what you observe from outside: why a long document costs disproportionately more than a short one, why providers meter what you send as well as what you get back, and why so much research since has gone into approximating attention rather than improving it. When a long prompt feels expensive out of all proportion, this is the reason.

Almost every generative system you have used sits on this architecture: the assistants from the big American and Chinese labs, the open-weight models you can download and run on your own hardware, and increasingly the image, audio and video models too. It turned out to be indifferent to what the sequence contains. Cut a picture into patches, treat each patch as an element, and the same machinery performs computer vision; do it with slices of audio and it transcribes speech. That indifference is why a vision researcher and a language researcher now read the same papers, and why capability improvements in one area keep arriving in the others shortly afterwards. For anyone who is not going to train a model, the useful knowledge stops roughly there. Knowing that the T in GPT stands for transformer is trivia. Knowing that attention is a weighted comparison of everything against everything predicts the behaviour you will encounter: a strong grip on context inside a single request, no memory whatsoever between separate ones unless something outside the model puts it back, and a cost that climbs faster than your document does.

In practice

The standard demonstration is a pair of sentences that differ by one word, where a pronoun changes what it refers to and nothing nearby announces the switch. Resolving that requires comparing it against both candidate nouns and weighing which reading makes sense, and the two nouns may sit far apart. A left-to-right model had to be still carrying the right one in its running summary by the time it arrived; attention simply looks back at both and scores them.

  • The trophy would not fit into the suitcase because it was too large. Most readers attach the pronoun to the trophy without noticing they decided anything.
  • The trophy would not fit into the suitcase because it was too small. One adjective changed, and the pronoun now points at the suitcase instead.

Often confused with

Large Language Models
One is the trained artefact, the other is the blueprint it was built from. Plenty of things that are not language models are built from the same blueprint.
Neural Networks
This is a particular way of wiring a neural network for sequences, not a departure from the idea of one.
Generative AI
That term describes what comes out. This one describes how the machine producing it is arranged internally.

Key takeaways

  • →Attention replaced a step-by-step running summary with direct comparison between every pair of positions.
  • →That removed the long-range forgetting and, because the comparisons are independent, unlocked training at scale.
  • →The same all-against-all comparison is why cost rises with the square of input length rather than in step with it.

Related concepts

Courses that teach this

Where this concept sits in the field

FAQ

Why did this architecture matter so much?
Because it made the previous ceiling on training irrelevant. Once sequence work could be spread across modern hardware, the limit on how much text a model saw became a question of budget rather than of architecture, and the capabilities people now take for granted emerged from pushing exactly that.
Do I need to understand attention to use these models?
Not to use them, but it pays for itself quickly. It tells you why a model is excellent at connecting two ends of one document and completely blank about a conversation you had yesterday, and why splitting a long input into sensible pieces is usually cheaper and better than sending the whole thing.

Sources

The primary text this definition rests on. Read it before relying on this one.

Last reviewed 26 September 2026 · Getting Digital