Skip to content
Getting Digital

Computer Vision

Also: CV, machine vision, image AI

Computer vision is the branch of artificial intelligence that turns images and video into structured answers: what is in the frame, where it sits, which pixels belong to it, and where it moves next.

Assessment. Vision demos travel much worse than anyone expects, and the accuracy quoted in a pilot usually describes the camera as much as the model. A project that controls its lighting and its lens is a better use of money than one that shops for a better architecture.

Computer vision is how software gets from a grid of pixel values to a decision. A photograph arrives as numbers describing colour and brightness at each position, with nothing in it marking where one object ends and the next begins, and the job is to recover that structure. Almost all of the work now runs on deep learning: first through convolutional networks, built on the assumption that a pattern means the same thing wherever in the frame it turns up, and more recently through Transformers borrowed wholesale from language research. The tasks hiding behind the one label are more separate than the label suggests, and what separates them most is what labelling the training data costs you.

  • Classification attaches one label to a whole image. Cheapest to annotate, and rarely enough on its own.
  • Detection draws a box around each object and names it. This is what most industrial, retail and security work requires.
  • Segmentation decides which pixels belong to which object, at an annotation cost that makes budget holders flinch.
  • Tracking keeps one identity attached to a thing across frames, the point at which video stops being a pile of photographs.
  • Reading, meaning optical character recognition and its relatives, lifts text off packaging, forms, number plates and screens.

What it is still bad at

The standing weakness is distribution shift, which means a system works on the pictures it was trained on and degrades on everything else. Change the mounting angle of a camera, swap a fluorescent tube for an LED panel, move from summer light to winter light, and something that passed acceptance testing begins missing cases without announcing that anything has changed. Occlusion is the second: half an object hidden behind another object is far harder than any tidy dataset prepares a model for. Rare events are the third and the cruellest, because the defect a factory most wants caught is by definition the one it has fewest photographs of. Past recognition, the gaps widen further. Counting a crowd, judging whether one object is about to hit another, reading a diagram as an argument rather than as shapes, telling a real face from a printed picture of one: these ask for a model of the world, and what these systems have is a model of appearance. The practical answer is almost never a newer architecture. It is controlling the input, keeping a person in the path for low-confidence cases, and testing on imagery from the site that will run the thing.

In practice

A utility asks customers to photograph their own meters instead of sending an engineer round. The model is trained on the archive the company already holds: engineer photographs, square on, in working light, of the two meter types fitted in most homes. It reads those nearly perfectly, so the pilot sails through.

Then it meets the public. Meters in cupboards shot at an angle, night-time attempts where the flash blows out the digits, an older dial meter whose pointers rotate in alternate directions, a cracked cover, a cat. What fixed it was not a bigger model. It was an on-screen outline the customer has to line the meter up with, a rejection message when the image is too dark, a confidence threshold below which the reading goes to a person, and retraining on customer photographs rather than engineers'. The lesson generalises: the capture pipeline is part of the model, whether or not anyone budgeted for it.

Often confused with

Deep Learning
Vision is a problem domain; deep learning is the technique currently used on most of it. The field existed for decades on hand-written edge and corner detectors before that technique arrived.
Generative AI
One reads images, the other writes them. Shared architectures now sit on both sides, which blurs a distinction worth keeping: recognition can be scored against ground truth, and generation largely cannot.
Natural Language Processing
The sibling field, for text. They converged on similar architectures but not on similar costs: anyone fluent can annotate a sentence, while a segmentation mask needs a trained annotator and real time per image.

Key takeaways

  • →Vision is several different tasks wearing one name, and they differ mainly in what annotating them costs.
  • →Distribution shift, occlusion and rare events are its persistent weaknesses, and none of the three is cured by a larger model.
  • →Controlling the camera, the lighting and the framing buys more accuracy per pound than changing the architecture.

Related concepts

  • Broader topicDeep Learning

    Computer vision is a flagship deep-learning application.

  • Transformers have spread from language into vision.

Courses that teach this

Where this concept sits in the field

FAQ

Do I have to train a model from scratch?
Almost never. Pre-trained detectors already handle common objects, and adapting one to your own images is a far shorter road than starting from nothing. Training from scratch makes sense only when your imagery looks nothing like ordinary photographs, as with medical scans, satellite bands or microscopy.
Why did accuracy collapse after the camera was moved? Why?
Because the model learned the old viewpoint along with the objects. Angle, distance, lens and lighting are all part of what it was fitted to, so moving the mount presents it with a subtly different world. Treat any physical change to the rig as a trigger for re-testing, and photograph a fixed validation set from the new position before trusting the numbers again.
Can it run in real time on video?
Yes, and this is routine on modest hardware near the camera. You pay for it in accuracy, since the models small enough to keep up with a frame rate are less capable than the ones that run offline. Tracking adds its own cost, because holding an identity across frames is a separate problem from recognising an object in one.
Is this area regulated?
Anything that identifies people is. Face recognition and other biometric uses are treated separately from ordinary vision work in European law and in a growing list of national and municipal rules, and the penalties attach to deployment rather than research. If your system distinguishes individuals rather than categories, make it a legal question before it becomes a technical one.

Sources

The primary text this definition rests on. Read it before relying on this one.

Last reviewed 26 September 2026 · Getting Digital