Generative AI describes systems that answer with content instead of with a number or a category. Ask one for a summary and prose comes back; ask a classifier the same thing and you get one item from a fixed list. That sounds like a technicality and it governs almost everything about how a project runs, because a fixed list can be scored against a known right answer and an open output largely cannot. The phrase also groups products sharing nothing but that property: an assistant in a chat window, an image tool inside a design suite, a completion pane in a code editor, a summary pinned to a support ticket. Different buyers, different risks, frequently different models underneath, and a purchase order reading generative AI has described none of it. The machinery varies just as much. Text and code come from autoregressive large language models, emitting a token at a time, each conditioned on everything already written. Images and audio mostly come from diffusion models, which begin with noise and refine it towards something matching the request, a wholly different procedure with its own controls and its own failure modes. Older adversarial designs, where one network produces and a second judges, remain in service in places. What unites them is a sampling step, and sampling is why one request twice yields two results, a feature when you want options and a liability when you want a record. It also means the word accuracy does not carry over. There is no correct paragraph. The measurement problem therefore moves: instead of scoring against a key, somebody has to write down what an acceptable answer looks like and collect cases of both kinds until a change can be judged.
The category hides the questions that matter
Three things the label does not tell you
When a vendor says a product includes generative AI, the phrase carries no information about anything that should decide the purchase. Whose model is it, and does it run on infrastructure you control or leave the building with your customers' data attached to it? What does a single call cost, and what latency comes with it, since the pair of them sets the limits of what you can afford to build around it? And above all, who or what inspects the output before a human sees it? These models optimise for content resembling their training data, not for content that is true, so a product with no verification step is selling you the draft and keeping the review for you. Ask those three questions and the category dissolves back into ordinary procurement, which is where it always belonged.
