Skip to content
Getting Digital

Prompt Engineering

Also: prompting, prompt design

Prompt engineering is the practice of composing what you send a generative model, the instruction, the supporting material and the worked examples, so that it returns something you can use.

Assessment. It is a genuine skill wearing a name that flatters the wrong half of it. The perishable half is phrasing, which providers keep making unnecessary with every release; the half that keeps paying is writing a brief precise enough that a stranger could mark the result right or wrong.

Two halves, ageing at different speeds

Everything filed under this heading belongs to one of two groups, and knowing which one you are practising is most of the skill. The first group is phrasing: the orderings, framings and forms of words that happen to work well on whichever model you are using this month. That knowledge is real and worth having, and it goes off quickly, because each provider trains its next release partly to remove the need for it. A trick that reliably unlocked better output last year is now either built into the default behaviour or broken, and the article that taught it to you has not been revised. The second group is specification, which is the same discipline behind a good bug report or a clear brief to a contractor. State the task. State what the output must contain and what it must never do. Supply the material the answer depends on instead of hoping it was memorised. Show one worked instance of an answer you would accept. None of that is tied to a particular model, it survives version changes, and it marks the line between a system that works and one that only demonstrates well. The word engineering is doing heavy lifting it has not earned, though, because it leaves out the step that would justify it. Writing a prompt is drafting. Holding a fixed set of inputs with their expected answers, re-running them after every edit and noticing when something got worse is what turns drafting into a controlled process. Without that, a team is not engineering anything. It is redecorating, one clever sentence at a time, with no way to tell an improvement from a lucky sample.

  • Write the acceptance criteria before the prompt. If you cannot say what would make an answer wrong, no amount of rewording will reliably produce a right one.
  • Supply the material rather than summoning it. Anything the answer depends on belongs in the request, which is why retrieval-augmented generation exists at all once that material grows past a paste.
  • One demonstrated answer does more than three paragraphs of description. Models imitate a shape you show far more dependably than they follow a shape you describe.
  • Pin the output format explicitly. Anything that downstream code has to parse should be specified as a structure, never requested politely.
  • Keep a regression file. A saved set of inputs with the answers you expect, re-run after each change, is the cheapest instrument available here and almost nobody builds one.

In practice

Consider a request teams send: summarise this support thread. It returns something plausible, different on every run, and nobody can argue that Tuesday's version improves on Monday's. The specified replacement names its reader, the engineer picking the ticket up tomorrow morning. It fixes the shape: what the customer wants, which remedies have been attempted, who the next action sits with. It forbids the thing that keeps going wrong, which is speculating about root cause. And it attaches one earlier summary that a human judged good. The model on both sides is identical. The second request can be marked right or wrong and the first cannot, and that distinction owes nothing at all to the model.

A test that costs nothing

Could a competent new colleague, handed your instruction and no other context, produce an answer you would sign off? If not, the ambiguity you are blaming on the model is yours, and rewording it more forcefully will not remove it.

Often confused with

Retrieval-Augmented Generation
Prompting is how you phrase the request. Retrieval is what fetches the facts the request will be answered from, and the two fail for entirely separate reasons.
Agentic AI
An agent is a loop that calls a model repeatedly and acts on each result. Your prompt is one component inside that loop, not the loop itself.

Key takeaways

  • →Split your own practice into phrasing and specification: the first expires with each model release, the second transfers across all of them.
  • →Nothing deserves the name engineering until there is measurement, and measurement here means saved cases with expected answers that you re-run.
  • →Most disappointing output comes from an instruction that never stated what a good answer would look like.

Related concepts

  • Prompt engineering is how you steer large language models.

  • RelatedAgentic AI

    An agent's planning quality is bounded by how its steps are prompted.

Courses that teach this

Where this concept sits in the field

Certifications that test this

Vendor exams whose syllabus covers this concept: facts, cost and a preparation path on each page.

FAQ

Does the skill disappear as models get better?
The phrasing half shrinks steadily, and that is the provider's explicit goal. The specification half does not, because the constraint it addresses is not the model's but yours: a task nobody has defined precisely cannot be completed correctly by any system, human or otherwise, and better models simply produce more fluent versions of the wrong thing.
How long does a specific technique stay useful?
Assume roughly the life of the model generation you discovered it on. Treat anything you read as provisional, verify it against your own saved cases, and be suspicious of advice that cannot explain why it works, since that is the kind most likely to have been a coincidence on somebody else's task.
Should prompts live in code or somewhere editable?
Somewhere a non-developer can edit them, with version history, and with the test cases sitting alongside. The people who know what a correct answer looks like are usually not the people who can deploy, and separating the two turns every wording change into a release.
What makes a prompt reliable rather than lucky?
Evidence across varied inputs. One impressive result tells you nothing, because the space of possible inputs is enormous and you sampled it once. Reliability means the same instruction holding up over the awkward cases, the empty ones and the ones in another language, which you only discover by collecting those cases deliberately.

Sources

The primary text this definition rests on. Read it before relying on this one.

Last reviewed 26 September 2026 · Getting Digital