In this lesson
Computer Vision in Practice
CNNs gave machines the ability to see. In this lesson you will see how that ability is applied to some of the most consequential problems in the world: diagnosing disease, guiding autonomous vehicles, and finding objects in images in real time.
What Is Computer Vision?
Computer vision is the field of AI concerned with enabling machines to interpret and understand visual information from the world: photographs, video, medical scans, satellite imagery, and more. Where a language model processes sequences of tokens, a computer vision model processes grids of pixel values and learns to extract meaning from spatial patterns.
Humans do this effortlessly. You glance at a street scene and instantly recognise pedestrians, road signs, distances, and potential hazards. For decades, this was considered uniquely biological. The breakthrough of deep learning, specifically convolutional neural networks (covered in Lesson 4.3), changed that assumption.
Today, computer vision systems match or exceed human performance on many narrow tasks. They can spot tumours in chest X-rays that radiologists miss on first review, count vehicles across thousands of traffic cameras simultaneously, and detect manufacturing defects at a speed no human inspector could match.
This lesson applies the CNN architecture from Lesson 4.3 and the transfer learning techniques from Lesson 4.5 to real problems. You do not need to train models from scratch. The pre-trained weights that researchers spent weeks and thousands of GPU hours computing are available to you in a few lines of code.
Core Computer Vision Tasks
Computer vision is not a single problem. It encompasses a family of related tasks, each with its own output format and level of difficulty.
Image Classification
Assign a single label to the whole image. Input: one image. Output: one class label with a confidence score. Example: "this chest X-ray shows pneumonia (94% confidence)."
Object Detection
Find where objects are in an image and what they are. Output: bounding boxes with class labels. Example: draw boxes around every car and pedestrian in a dashcam frame.
Image Segmentation
Assign a class label to every pixel. Far more precise than bounding boxes. Example: colour each pixel in a road scene as road, car, sky, pavement, or pedestrian.
Pose Estimation
Detect the positions of key joints in a human body or object. Output: a set of keypoints (e.g., shoulders, elbows, wrists, hips, knees, ankles). Used in sports analytics, physical therapy, and action recognition.
Video Understanding
Extend image tasks to sequences of frames. Adds the challenge of modelling motion and temporal context. Used in activity recognition, surveillance, and sports analysis.
Optical Character Recognition (OCR)
Extract text from images. Used in digitising documents, reading number plates, and processing forms. Modern systems use a combination of CV and language model techniques.
Each task increases in complexity and output specificity. A classification model needs to understand the image globally; a segmentation model must make a precise decision for every one of millions of pixels. The right task depends on what your application actually requires.
Object Detection and YOLO
Object detection is one of the most commercially important CV tasks. Autonomous vehicles, retail analytics, security cameras, and medical imaging all rely on it. The core challenge is that a single image may contain dozens of objects at different scales and positions, and the model must find them all at once.
How Bounding Boxes Work
A bounding box is defined by four numbers: the x and y coordinates of the top-left corner, plus the width and height of the box. A detection model predicts these four numbers along with a class probability distribution for each detected object. At inference time, a confidence threshold filters out low-probability detections, and a technique called Non-Maximum Suppression (NMS) removes duplicate boxes that refer to the same object.
YOLO: You Only Look Once
Early detection models like R-CNN (Region-based CNN, Girshick et al. 2014) used a two-stage approach: first propose candidate regions, then classify each one. This was accurate but slow, making real-time video processing impractical.
YOLO (Redmon et al. 2016) reframed detection as a single regression problem. Instead of proposing regions separately, YOLO divides the image into a grid and predicts bounding boxes and class probabilities for each grid cell in a single forward pass through the network. This made real-time detection possible on standard hardware.
The YOLO Family
| Version | Year | Key Advance | Notable Use |
|---|---|---|---|
| YOLOv1 | 2016 | Single-pass detection (Redmon et al.) | Proved real-time detection was possible |
| YOLOv3 | 2018 | Multi-scale detection, Darknet-53 backbone | Widely adopted in industry and research |
| YOLOv5 | 2020 | PyTorch rewrite, easier fine-tuning (Ultralytics) | Most deployed version in production systems |
| YOLOv8 | 2023 | Unified detection + segmentation + pose | Modern default for custom CV pipelines |
| YOLOv11 | 2024 | Improved efficiency and accuracy (Ultralytics) | Edge deployment on IoT devices |
YOLO models prioritise speed. For applications requiring the highest possible accuracy on still images (e.g., batch medical image analysis), two-stage detectors like Faster R-CNN may outperform YOLO. For real-time video at 30+ frames per second, YOLO is typically the better choice.
Image Segmentation
Object detection tells you where something is with a rectangle. Segmentation tells you exactly which pixels belong to it. This precision matters enormously in surgery planning, autonomous driving, and satellite land-cover mapping.
Semantic Segmentation
Assigns a class to every pixel, but does not distinguish between individual instances. All cars in a scene are labelled "car" without differentiating car #1 from car #2. Models: FCN (Fully Convolutional Network, Long et al. 2015), DeepLab (Chen et al., Google 2014).
Instance Segmentation
Labels each pixel AND tracks individual object instances. Car #1 and car #2 get different colour masks. Harder than semantic segmentation. Standard model: Mask R-CNN (He et al., FAIR 2017), which extends Faster R-CNN with a mask prediction branch.
Panoptic Segmentation
Combines both: every pixel labelled (semantic) and every countable object individually tracked (instance). Introduced by Kirillov et al. (FAIR, 2019). Used in the most demanding scene understanding tasks.
The U-Net Architecture
U-Net (Ronneberger et al. 2015) was originally designed for biomedical image segmentation and has become one of the most widely used CV architectures in medicine. Its distinctive feature is a symmetric encoder-decoder structure with skip connections that pass high-resolution spatial information directly from the encoder to the corresponding decoder layer, enabling precise pixel-level localisation.
OpenCV in Practice
OpenCV (Open Source Computer Vision Library) is the foundational toolkit for CV in Python. Released in 2000 and maintained by the OpenCV Foundation, it provides hundreds of image processing functions that sit below the neural network level: loading and resizing images, colour space conversion, edge detection, contour finding, and camera access.
Most CV pipelines start with OpenCV for preprocessing, feed processed images into a deep learning model, and then use OpenCV again to overlay results on the original image for visualisation.
Basic Image Operations
import cv2 import numpy as np # Load an image (BGR colour order by default in OpenCV) img = cv2.imread("street.jpg") print(img.shape) # (height, width, channels) e.g. (720, 1280, 3) # Convert to RGB for display in matplotlib img_rgb = cv2.cvtColor(img, cv2.COLOR_BGR2RGB) # Convert to greyscale grey = cv2.cvtColor(img, cv2.COLOR_BGR2GRAY) # Resize to 224x224 for a model that expects that input size resized = cv2.resize(img, (224, 224)) # Apply Gaussian blur to reduce noise before edge detection blurred = cv2.GaussianBlur(grey, (5, 5), 0) # Canny edge detection edges = cv2.Canny(blurred, threshold1=50, threshold2=150) # Normalise pixel values to [0, 1] for model input img_norm = resized.astype(np.float32) / 255.0
Drawing Detection Results
# After a model returns detections, draw them back onto the original image # detections: list of (label, confidence, x, y, w, h) for label, confidence, x, y, w, h in detections: # Draw bounding box rectangle cv2.rectangle(img, (x, y), (x + w, y + h), (0, 255, 0), 2) # Compose label string text = f"{label}: {confidence*100:.1f}%" # Draw filled background for the text so it is legible (tw, th), _ = cv2.getTextSize(text, cv2.FONT_HERSHEY_SIMPLEX, 0.6, 1) cv2.rectangle(img, (x, y - th - 8), (x + tw, y), (0, 255, 0), -1) cv2.putText(img, text, (x, y - 4), cv2.FONT_HERSHEY_SIMPLEX, 0.6, (0, 0, 0), 1) # Save the annotated image cv2.imwrite("annotated.jpg", img)
A Typical CV Pipeline
Using Pre-Trained Models
You rarely need to train a CV model from scratch. Pre-trained models that have already learned rich visual representations from millions of images are available through torchvision, Hugging Face, and the TensorFlow Hub.
Object Detection with Hugging Face
The Hugging Face Transformers library provides a high-level pipeline for object detection using modern transformer-based detectors like DETR (Detection Transformer, Carion et al., FAIR 2020) and DETA, which bring the self-attention mechanism from Lesson 4.4 into the vision domain.
from transformers import pipeline from PIL import Image # Load a pre-trained DETR model (Facebook Research, fine-tuned on COCO) detector = pipeline( "object-detection", model="facebook/detr-resnet-50", threshold=0.7 # only return detections with confidence > 70% ) # Load image from disk image = Image.open("street.jpg") # Run detection results = detector(image) # Each result is a dict: {'score', 'label', 'box': {'xmin','ymin','xmax','ymax'}} for r in results: box = r["box"] print(f"{r['label']:15s} score={r['score']:.2f} " f"box=({box['xmin']},{box['ymin']},{box['xmax']},{box['ymax']})")
Image Classification with torchvision
import torch from torchvision import models, transforms from PIL import Image import json, urllib.request # Load ResNet-50 pre-trained on ImageNet-1k (1000 classes) model = models.resnet50(weights=models.ResNet50_Weights.IMAGENET1K_V2) model.eval() # switch off dropout and batchnorm training behaviour # Standard ImageNet preprocessing preprocess = transforms.Compose([ transforms.Resize(256), transforms.CenterCrop(224), transforms.ToTensor(), transforms.Normalize(mean=[0.485, 0.456, 0.406], std=[0.229, 0.224, 0.225]), ]) # Load and preprocess the image img = Image.open("cat.jpg") batch = preprocess(img).unsqueeze(0) # add batch dimension: (1, 3, 224, 224) # Run inference (no gradient computation needed) with torch.no_grad(): logits = model(batch) # shape: (1, 1000) probs = torch.softmax(logits, dim=1) # Get the top 3 predictions top3 = probs.topk(3) top3_indices = top3.indices[0].tolist() top3_probs = top3.values[0].tolist() # Map indices to human-readable labels using ImageNet class list labels_url = "https://raw.githubusercontent.com/anishathalye/imagenet-simple-labels/master/imagenet-simple-labels.json" labels = json.loads(urllib.request.urlopen(labels_url).read()) for idx, prob in zip(top3_indices, top3_probs): print(f"{labels[idx]:30s} {prob*100:.2f}%")
Most pre-trained detection models are trained on COCO (Common Objects in Context, Lin et al. 2014): 118,000 training images with 80 object categories. Classification models are typically trained on ImageNet (Deng et al. 2009): 1.28 million training images across 1,000 categories. When you load a pre-trained model, the class labels correspond to these datasets.
Medical Imaging AI
Healthcare is one of the most consequential domains for computer vision. Medical images, including X-rays, CT scans, MRI, histopathology slides, and retinal fundus photographs, contain patterns that are difficult, time-consuming, and expensive for human experts to interpret at scale. AI is beginning to change that.
Key Applications
| Task | Modality | Notable Result |
|---|---|---|
| Diabetic retinopathy screening | Retinal fundus photos | Google's ADAS-M system matched specialist performance at moderate to severe grades (Gulshan et al., JAMA 2016) |
| Breast cancer detection | Mammography | McKinney et al. (Nature 2020) showed a model outperformed radiologists on UK and US datasets with 11% fewer false positives |
| Chest X-ray pathology | Chest radiographs | CheXNet (Stanford, Rajpurkar et al. 2017) detected pneumonia better than radiologists as measured by F1 score |
| Skin lesion classification | Dermoscopy / photos | Esteva et al. (Nature 2017) matched board-certified dermatologists on 129,450 skin lesion images |
| Histopathology cancer grading | Whole-slide images | PathAI and Paige AI offer FDA-cleared tools for prostate cancer and cervical cancer screening |
| Retinal disease (multi-class) | OCT scans | DeepMind / Moorfields (De Fauw et al., Nature Medicine 2018) recommended correct treatment urgency in 94% of cases |
What Makes Medical Imaging Different
Medical CV has unique challenges compared to general CV:
Data Scarcity
Annotating medical images requires expert clinicians, making large labelled datasets expensive and rare. Transfer learning from ImageNet or similar large datasets is standard practice, even though natural images and medical scans look very different.
Class Imbalance
Disease prevalence is often low. In a typical screening programme, 95% of scans may be normal. A model that predicts "normal" for every scan would achieve 95% accuracy but would be clinically useless. Metrics like AUC-ROC, sensitivity, and specificity matter more than raw accuracy.
Distribution Shift
A model trained on images from one scanner manufacturer, one hospital's patient population, or one geographic region may fail on data from a different source. This is a major barrier to clinical deployment and a frequent finding in independent validation studies.
Regulatory Requirements
In most countries, clinical AI tools must obtain regulatory clearance. In the United States, the FDA reviews AI software as a medical device (SaMD). In the EU, CE marking under the Medical Device Regulation (MDR) is required. Regulatory pathways reward interpretability and robustness over peak benchmark accuracy.
DICOM and Medical Image Formats
Medical images are typically stored in DICOM (Digital Imaging and Communications in Medicine) format, which bundles the pixel data with rich metadata: patient information, scanner parameters, acquisition settings, and standardised anatomical labels. The pydicom library is the standard Python tool for reading DICOM files. Many CV pipelines convert DICOM images to PNG or NumPy arrays before feeding them into models.
import pydicom import numpy as np # Load a DICOM chest X-ray file dcm = pydicom.dcmread("chest_xray.dcm") # Access metadata print(dcm.PatientAge) # e.g. '045Y' print(dcm.Modality) # 'CR' (computed radiography) print(dcm.StudyDescription) # 'CHEST AP' # Extract pixel array and apply stored windowing for display pixel_array = dcm.pixel_array.astype(np.float32) print(pixel_array.shape) # (2048, 2048) for a high-resolution chest X-ray # Normalise to [0, 255] for model input or OpenCV display pixel_min, pixel_max = pixel_array.min(), pixel_array.max() img_8bit = ((pixel_array - pixel_min) / (pixel_max - pixel_min) * 255).astype(np.uint8)
Real-World Applications
Computer vision is embedded in systems that billions of people encounter every day. Here is a survey of key domains.
Autonomous Vehicles
Real-time object detection (pedestrians, cyclists, vehicles, signs), lane detection, depth estimation from stereo cameras or LiDAR, and semantic segmentation of driveable area. Tesla, Waymo, and Mobileye all rely on multi-camera CV pipelines.
Retail and Logistics
Amazon Go uses CV to track what customers place in their baskets. Warehouse systems count inventory, identify damaged packages, and route items through sorting facilities without human intervention.
Satellite Imagery
Land cover classification, deforestation monitoring, building footprint extraction, crop yield prediction, and disaster response mapping. Planet Labs captures images of the entire Earth's land surface daily.
Manufacturing
Automated visual inspection identifies surface defects, dimensional errors, and assembly mistakes on production lines at speeds far exceeding human inspection. Used widely in semiconductor manufacturing, automotive, and electronics.
Sports Analytics
Player tracking, ball trajectory analysis, offside detection in football, shot selection in basketball, and injury risk assessment. Broadcast networks overlay live statistics derived from CV inference.
Agriculture
Drone imagery analysis for crop health assessment, disease detection, and yield mapping. Root-zone analysis from satellite multispectral imaging guides irrigation and fertiliser decisions at scale.
Roboflow is a popular platform for annotating custom image datasets, augmenting training data, and exporting in formats compatible with YOLO, Detectron2, and other frameworks. It dramatically reduces the time needed to prepare a custom object detection dataset from scratch. Their Universe hub contains over 200,000 publicly contributed datasets across domains from medical to wildlife to industrial inspection.
Ethics in Computer Vision
Computer vision raises some of the most pressing ethical questions in AI, partly because it interacts with the physical world and involves the identification of individual people in ways that other AI applications do not.
Facial Recognition
Facial recognition is one of the highest-stakes CV applications. Systems matching faces against watchlists are deployed in airports, stadiums, retail stores, and public spaces, often without the knowledge of the people being scanned. Several studies have documented significant accuracy disparities across demographic groups. Joy Buolamwini and Timnit Gebru's Gender Shades study (referenced in Lesson 5.1) showed error rates for dark-skinned women were up to 34 percentage points higher than for light-skinned men in commercial facial analysis systems.
In response to accuracy and civil liberties concerns, IBM, Amazon, and Microsoft all paused or restricted sales of facial recognition technology to law enforcement between 2020 and 2021. Several cities in the United States, including San Francisco and Boston, have banned government use of the technology outright.
Training Data Representation
A model trained predominantly on images from one demographic group will perform worse on under-represented groups. Audit your training data for demographic balance before deploying any CV system that interacts with people.
Surveillance and Privacy
CV systems enable persistent, automated surveillance at a scale that was previously impossible. This changes the power dynamics between organisations that deploy cameras and the individuals those cameras observe. Consider whether your application would be acceptable if people knew exactly what was happening to their image data.
Consent and Transparency
People have a reasonable expectation of knowing when they are being identified rather than merely observed. Transparency about CV system use, combined with clear mechanisms for opting out or challenging an incorrect identification, is both ethically necessary and increasingly required by law under frameworks like the EU AI Act.
High-Stakes Deployment
For medical imaging, autonomous vehicles, and criminal justice applications, model errors have severe consequences. Treat benchmark accuracy with scepticism and invest in understanding failure modes, class-conditional error rates, and performance across subgroups before deployment.
A model achieving 99% accuracy on a benchmark dataset may perform dramatically worse in deployment due to distribution shift: different lighting, different camera angles, different patient demographics, or weather conditions not present in training data. Prospective clinical trials and real-world performance monitoring are essential for any high-stakes CV deployment. Benchmark accuracy is a starting point, not a deployment guarantee.
Key Takeaways
- Computer vision covers a family of tasks: classification (what?), detection (what and where?), segmentation (which pixels?), and pose estimation.
- YOLO (Redmon et al. 2016) reformulated object detection as a single regression problem, enabling real-time detection in one forward pass. The family has continued to evolve through YOLOv8 and YOLOv11.
- Segmentation assigns class labels to individual pixels. U-Net's encoder-decoder architecture with skip connections is the standard for medical image segmentation tasks.
- OpenCV is the foundational toolkit for image pre-processing, colour conversion, and result overlay. Nearly all CV pipelines combine OpenCV with a deep learning framework.
- Pre-trained models on COCO (detection) and ImageNet (classification) are available through Hugging Face and torchvision. For most custom tasks, fine-tuning these is far more practical than training from scratch.
- Medical imaging AI must contend with data scarcity, class imbalance, distribution shift between hospitals, and strict regulatory requirements. Benchmark accuracy alone is an insufficient measure of clinical readiness.
- Facial recognition, surveillance, and biometric identification applications require careful attention to demographic fairness, consent, and transparency. Accuracy disparities across demographic groups are well documented and must be addressed before deployment.
In the final lesson of the course, you will look ahead: multimodal models that combine vision, language, and audio; AI agents that can take actions in the world; the landscape of AI safety research; and the economic and societal changes that the current generation of AI systems is already beginning to drive.
Going Deeper