Is it still a separate field?
Natural language processing is the branch of AI that gets computers working with text and speech: filing messages into categories, lifting out names, dates and amounts, translating, summarising, answering questions and judging what a piece of writing is doing. It is also the corner of AI most disturbed by the past few years, so the question a reader arrives with is sharper than any definition. If a general model handles all of this from an instruction, is there still a discipline here? For getting a first version running, mostly not, and pretending otherwise burns a quarter. A team that budgets a research programme for sentiment scoring or document classification is paying to solve something already solved elsewhere. Assembling a labelled corpus, selecting an architecture and training it has been replaced, for an enormous range of tasks, by writing careful instructions and inspecting the output. Deep learning began that collapse and the Transformer completed it by producing systems competent at work they were never aimed at. Reading that as a threat to the field misunderstands what the field was ever about, which was never the architecture of the season.
What stays distinctly NLP is everything the model declines to do for you. Somebody must decide what counts as a complaint rather than a question, and write it down precisely enough that two colleagues reading the same message reach the same verdict. Somebody must assemble the set of cases that establishes whether this week's version improves on last week's, because without it every change is an opinion delivered with confidence. Somebody must judge whether this task should touch a rented large language model at all, given how often it runs, how quickly it has to answer and whether the text is permitted to leave the building. Those decisions predate every current model and outlive each replacement of the one underneath. Speech, minority languages and regulated domains hold on to their own specialists for a related reason: a general model is fluent in whatever the internet writes most of, and thins out quickly past that boundary. They are also where projects die, far more reliably than at the modelling step. Prompt engineering covers the instruction-writing portion and it is genuine work, but a project that got its labels wrong cannot be rescued by better instructions, and tends to learn this only after the demo has been applauded. Ask any team whose classifier keeps disagreeing with its own reviewers: the fault sits almost never in the model and almost always in a definition nobody bothered to write down.
