Guess the next word, get a general writer
A large language model is a neural network whose entire training task is to continue a passage of text. Run that across a large share of what people have written and the network cannot succeed without absorbing spelling, grammar, register, document shapes, a rough map of who did what and when, and the moves an argument tends to make, since every one of those improves the guess. Nobody wrote a module for French or for Python. Both arrive because continuing text written in them demands it. That is the whole trick, and it accounts for the fact that one set of weights drafts an email, restates a contract clause in plain language and reads a stack trace without being rebuilt for each job. Almost all of them sit on the Transformer, the design that made training at this size practical, and what comes out is generative AI working in text.
What the objective buys you is fluency, which is not knowledge, and the distance separating them is where practical trouble lives. The model compresses what it read, so specifics return approximate and are delivered in the same even register as everything else. There is no separate store it can consult to check itself, which means it cannot report that your question falls outside what it absorbed. The remedy is not a cleverer model but retrieval: put the source in front of it and ask for an answer drawn only from that source. Steering the rest, the tone, the format, the refusals, is prompt engineering, and it resembles writing a specification far more than it resembles finding magic words.
- Fluent and wrong are indistinguishable. How well an answer reads carries no information about whether it is correct, so any review habit that rests on it reading fine will pass the worst output you produce.
- Private material has to arrive in the prompt. Your contracts, your ticket history and your price list were not in the training data, and faced with their absence the model improvises rather than declines.
- One trained artefact, many jobs. The same weights are adapted by instruction or by light fine-tuning, which is why the label foundation model stuck: the costly training happened once, upstream of you, and everybody downstream shares it.
- Recency is a supply question, not a capability. Training data stops somewhere. Anything after that reaches the model only because your system fetched it and passed it along.
