Computer vision is how software gets from a grid of pixel values to a decision. A photograph arrives as numbers describing colour and brightness at each position, with nothing in it marking where one object ends and the next begins, and the job is to recover that structure. Almost all of the work now runs on deep learning: first through convolutional networks, built on the assumption that a pattern means the same thing wherever in the frame it turns up, and more recently through Transformers borrowed wholesale from language research. The tasks hiding behind the one label are more separate than the label suggests, and what separates them most is what labelling the training data costs you.
- Classification attaches one label to a whole image. Cheapest to annotate, and rarely enough on its own.
- Detection draws a box around each object and names it. This is what most industrial, retail and security work requires.
- Segmentation decides which pixels belong to which object, at an annotation cost that makes budget holders flinch.
- Tracking keeps one identity attached to a thing across frames, the point at which video stops being a pile of photographs.
- Reading, meaning optical character recognition and its relatives, lifts text off packaging, forms, number plates and screens.
What it is still bad at
The standing weakness is distribution shift, which means a system works on the pictures it was trained on and degrades on everything else. Change the mounting angle of a camera, swap a fluorescent tube for an LED panel, move from summer light to winter light, and something that passed acceptance testing begins missing cases without announcing that anything has changed. Occlusion is the second: half an object hidden behind another object is far harder than any tidy dataset prepares a model for. Rare events are the third and the cruellest, because the defect a factory most wants caught is by definition the one it has fewest photographs of. Past recognition, the gaps widen further. Counting a crowd, judging whether one object is about to hit another, reading a diagram as an argument rather than as shapes, telling a real face from a printed picture of one: these ask for a model of the world, and what these systems have is a model of appearance. The practical answer is almost never a newer architecture. It is controlling the input, keeping a person in the path for low-confidence cases, and testing on imagery from the site that will run the thing.
