Skip to content
Getting Digital

Large Language Models

Also: LLM, foundation model, language model

Large language models are neural networks trained on a vast quantity of written text to predict the piece of text that comes next, an objective narrow enough to state in a line and broad enough to yield writing, translation, code and summary as side effects.

Assessment. Treat a large language model as an engine for transforming text you hand it, not as somewhere to look things up. Nearly every disappointment teams walk into comes from the second reading, and the repair is to change where the facts arrive from rather than to change whose model sits behind the call.

Guess the next word, get a general writer

A large language model is a neural network whose entire training task is to continue a passage of text. Run that across a large share of what people have written and the network cannot succeed without absorbing spelling, grammar, register, document shapes, a rough map of who did what and when, and the moves an argument tends to make, since every one of those improves the guess. Nobody wrote a module for French or for Python. Both arrive because continuing text written in them demands it. That is the whole trick, and it accounts for the fact that one set of weights drafts an email, restates a contract clause in plain language and reads a stack trace without being rebuilt for each job. Almost all of them sit on the Transformer, the design that made training at this size practical, and what comes out is generative AI working in text.

What the objective buys you is fluency, which is not knowledge, and the distance separating them is where practical trouble lives. The model compresses what it read, so specifics return approximate and are delivered in the same even register as everything else. There is no separate store it can consult to check itself, which means it cannot report that your question falls outside what it absorbed. The remedy is not a cleverer model but retrieval: put the source in front of it and ask for an answer drawn only from that source. Steering the rest, the tone, the format, the refusals, is prompt engineering, and it resembles writing a specification far more than it resembles finding magic words.

  • Fluent and wrong are indistinguishable. How well an answer reads carries no information about whether it is correct, so any review habit that rests on it reading fine will pass the worst output you produce.
  • Private material has to arrive in the prompt. Your contracts, your ticket history and your price list were not in the training data, and faced with their absence the model improvises rather than declines.
  • One trained artefact, many jobs. The same weights are adapted by instruction or by light fine-tuning, which is why the label foundation model stuck: the costly training happened once, upstream of you, and everybody downstream shares it.
  • Recency is a supply question, not a capability. Training data stops somewhere. Anything after that reaches the model only because your system fetched it and passed it along.

In practice

Take a question a freelancer asks: what notice period does my client contract require? Put it to a chat assistant as a bare question and you get a confident reply assembled from every agreement the model has ever read, which is a description of a common clause rather than a reading of yours. Paste the clause in and ask for the notice period, the date it runs from and anything that shifts if the client stops early, and the same model, on the same afternoon, becomes dependable. The job changed from recall to comprehension. The question was never too hard; the material was simply absent.

Often confused with

Transformers
The architecture is a design for wiring attention together, published once and reused everywhere. What you call by this name is a particular set of weights somebody spent months and a data centre producing from that design.
Generative AI
That label covers the entire category, images and audio and video alongside text. This is its language member, and the member that ends up writing instructions for the rest.
Retrieval-Augmented Generation
RAG is not a variety of model. It is the plumbing that finds the right passage and places it in front of whichever model you are already calling.

Key takeaways

  • →A single objective, continuing text at scale, produces broad language ability as a by-product; none of the individual skills were specified in advance.
  • →Fluency and accuracy vary independently. Supply the source material and the model is dependable; ask it to remember, and it improvises in the same steady voice.

Related concepts

Courses that teach this

Where this concept sits in the field

Certifications that test this

Vendor exams whose syllabus covers this concept: facts, cost and a preparation path on each page.

FAQ

Do these models understand what they are saying?
They behave as though they do across a wide spread of tasks, and nobody qualified to settle the philosophical argument has settled it. For work, use a behavioural test instead: does the output survive the awkward cases you already worry about, checked against answers you wrote down before you looked?
Where do the fabricated citations come from?
Because producing a plausible continuation is exactly the job it was trained for, and a citation-shaped string is plausible. Nothing inside it separates remembering from constructing. Hand it the documents and ask it to quote from them, and most of this failure goes away.
Fine-tune or just prompt?
Prompt first, and keep prompting for longer than feels sophisticated. Fine-tuning buys consistency of format, voice and refusal behaviour on a task you can already describe; it does not reliably install facts, and facts are usually what the team wanted. If the gap is knowledge, retrieval closes it.

Sources

The primary text this definition rests on. Read it before relying on this one.

Last reviewed 26 September 2026 · Getting Digital