The mechanism doing the work is called attention, and the clearest way to see what it buys is to look at what it displaced. The previous generation of sequence models read from left to right, carrying a running summary forward one position at a time. Two consequences followed. First, that summary is a bottleneck: information from the opening of a long passage has to survive being compressed and recompressed at every subsequent step to still be available at the end, and in practice much of it did not. Second, position ten could not be computed until position nine had finished, so the work refused to spread across a graphics card, which put a hard ceiling on how much text anyone could afford to train on. Attention dissolves both at once. Every position is compared directly with every other position, so a pronoun near the end of a paragraph consults the noun it refers to at the beginning without anything in between wearing the connection down. And because those comparisons do not depend on each other, they all run at the same time. The first change is why these models handle context as well as they do. The second is why they could be scaled at all, and scale is what produced the behaviour that made everyone pay attention in the first place.
The cost hiding inside the idea
Comparing every position with every other position means the work grows with the square of the input length, not in step with it. That one property explains a surprising amount of what you observe from outside: why a long document costs disproportionately more than a short one, why providers meter what you send as well as what you get back, and why so much research since has gone into approximating attention rather than improving it. When a long prompt feels expensive out of all proportion, this is the reason.
Almost every generative system you have used sits on this architecture: the assistants from the big American and Chinese labs, the open-weight models you can download and run on your own hardware, and increasingly the image, audio and video models too. It turned out to be indifferent to what the sequence contains. Cut a picture into patches, treat each patch as an element, and the same machinery performs computer vision; do it with slices of audio and it transcribes speech. That indifference is why a vision researcher and a language researcher now read the same papers, and why capability improvements in one area keep arriving in the others shortly afterwards. For anyone who is not going to train a model, the useful knowledge stops roughly there. Knowing that the T in GPT stands for transformer is trivia. Knowing that attention is a weighted comparison of everything against everything predicts the behaviour you will encounter: a strong grip on context inside a single request, no memory whatsoever between separate ones unless something outside the model puts it back, and a cost that climbs faster than your document does.
