Skip to content
Getting Digital

Natural Language Processing

Also: NLP, computational linguistics, language AI

Natural language processing is the area of AI that makes human language readable and writable by machine: sorting text into categories, pulling structured facts out of it, translating it, generating it and measuring whether any of that came out right.

Assessment. The field has not been swallowed by general-purpose language models, but its centre has moved. Modelling is now largely a commodity you rent by the call, while the work that still takes judgement is deciding what a label means and proving a system applies it consistently, and none of that judgement lives inside the model.

Is it still a separate field?

Natural language processing is the branch of AI that gets computers working with text and speech: filing messages into categories, lifting out names, dates and amounts, translating, summarising, answering questions and judging what a piece of writing is doing. It is also the corner of AI most disturbed by the past few years, so the question a reader arrives with is sharper than any definition. If a general model handles all of this from an instruction, is there still a discipline here? For getting a first version running, mostly not, and pretending otherwise burns a quarter. A team that budgets a research programme for sentiment scoring or document classification is paying to solve something already solved elsewhere. Assembling a labelled corpus, selecting an architecture and training it has been replaced, for an enormous range of tasks, by writing careful instructions and inspecting the output. Deep learning began that collapse and the Transformer completed it by producing systems competent at work they were never aimed at. Reading that as a threat to the field misunderstands what the field was ever about, which was never the architecture of the season.

What stays distinctly NLP is everything the model declines to do for you. Somebody must decide what counts as a complaint rather than a question, and write it down precisely enough that two colleagues reading the same message reach the same verdict. Somebody must assemble the set of cases that establishes whether this week's version improves on last week's, because without it every change is an opinion delivered with confidence. Somebody must judge whether this task should touch a rented large language model at all, given how often it runs, how quickly it has to answer and whether the text is permitted to leave the building. Those decisions predate every current model and outlive each replacement of the one underneath. Speech, minority languages and regulated domains hold on to their own specialists for a related reason: a general model is fluent in whatever the internet writes most of, and thins out quickly past that boundary. They are also where projects die, far more reliably than at the modelling step. Prompt engineering covers the instruction-writing portion and it is genuine work, but a project that got its labels wrong cannot be rescued by better instructions, and tends to learn this only after the demo has been applauded. Ask any team whose classifier keeps disagreeing with its own reviewers: the fault sits almost never in the model and almost always in a definition nobody bothered to write down.

In practice

One comparison settles real budgets. A publisher holds an archive of articles and wants every mention of a named product tagged, then wants the archive re-tagged each time the product list changes. Route one sends every article to a hosted frontier model and pays per call. It works on day one with no labelled data and nobody to maintain it, and every re-run costs what the first run cost, multiplied by the size of the archive. Route two uses that same frontier model once, on a sample, to produce labels, trains a small encoder on them and runs it on machines you already rent. Route two spends engineer time up front and close to nothing per article afterwards, so re-tagging the archive stops being a budget conversation and becomes an overnight job. It also returns an answer quickly enough to tag at the moment of publication, keeps the text inside your own systems, and stays consistent from the first article to the last, because the model only changes when you change it. That last property is route one's quiet weakness: a supplier updates the hosted model mid-project, and articles tagged in spring were tagged by a different system than articles tagged in autumn, which is an awkward thing to explain to anyone auditing the result.

  • Route one suits work that runs occasionally, criteria still being argued about, an absence of labelled examples, and teams with nobody to look after a trained model.
  • Route two suits narrow tasks at continuous volume, answers needed immediately, text that must not leave the building, and outputs somebody may have to reproduce a year later.
  • Both together is what experienced teams settle on: the rented model manufactures training labels and adjudicates the hard cases, the small local model carries the traffic.

Often confused with

Large Language Models
A field with decades behind it should not be confused with the instrument currently dominating one step of it. Swap that instrument out and every question about labels and evaluation stays exactly where it was.
Generative AI
Generation is a single task here among many. Half this discipline is reading rather than writing: classifying, extracting, matching and measuring, none of which produce new text at all.
Computer Vision
The sibling discipline on the image side, now sharing most of its architecture. The methods converged; the annotation problems, evaluation traditions and legal constraints did not.

Key takeaways

  • →Getting a language task working for the first time is now mostly an exercise in instruction-writing, and treating it as a research project wastes a quarter.
  • →The durable skills sit either side of the model: defining labels well enough that people agree, and measuring output against cases fixed in advance.
  • →Rent a general model for work that runs occasionally; train a small specialised one for work that runs constantly, needs to be fast, or must stay in-house.

Related concepts

Courses that teach this

Where this concept sits in the field

FAQ

Is NLP simply large language models now?
The modelling step frequently is. The field around it is not: task definition, annotation, evaluation and the operational question of what runs where all survived intact, and they were always the parts that decided whether a system was any good.
Are the classical techniques worth learning?
The old pipeline of tokenisers, taggers and hand-built feature extraction is largely museum work for a new project. What repays study is evaluation, annotation practice, information retrieval and the statistics of measurement, because those hold whatever is generating the text this year.
What does the job involve in practice now?
Writing specifications, building evaluation sets, wiring up retrieval, choosing between a rented model and a small local one, and arguing about definitions with whoever owns the process being automated. Training something from scratch is a minority activity outside research groups.

Sources

The primary text this definition rests on. Read it before relying on this one.

Last reviewed 26 September 2026 · Getting Digital