Knowledge Check

Phase 4 — Neural Network Architectures

Four questions on neurons, convolutional networks, attention, and how modern AI systems are built on top of existing ones.

Question 1 of 4
Question 1 of 4

The basic computational unit of a neural network, loosely inspired by the brain, is called what?

AA layer
BA neuron (also called a node or unit)
CA tensor
DA gradient
Explanation
A neuron (or node) receives numerical inputs, multiplies each by a learned weight, adds a bias value, and passes the result through an activation function. Its output then feeds into neurons in the next layer. This simple operation, stacked in layers and trained with gradient descent across billions of examples, produces the sophisticated behaviour of modern AI models. A tensor is the multi-dimensional array structure used to store data; a gradient is the direction signal used during training.
Question 2 of 4

Why are Convolutional Neural Networks (CNNs) especially well suited to image tasks?

AThey process all pixel values in a single parallel operation
BThey convert images into text before any processing occurs
CThey scan images with small learnable filters that detect local patterns like edges and textures
DThey use attention to compare every pixel to every other pixel simultaneously
Explanation
CNNs use small learnable filters (kernels) that slide across an image detecting local features. Early layers pick up edges and corners; deeper layers detect textures, shapes, and eventually complex objects like faces. This works because spatial proximity matters in images — nearby pixels tend to be related. It is far more efficient than connecting every pixel to every neuron. Option D describes Vision Transformers (ViTs), a newer architecture that does use attention across image patches.
Question 3 of 4

The key innovation introduced in the Transformer architecture was what?

AUsing more layers than any model that came before it
BTraining on images rather than text sequences
CA self-attention mechanism that lets every position in the input attend to every other position
DReplacing gradient descent with a faster optimisation method
Explanation
Before Transformers, models processed sequences step by step, making it hard to relate words far apart in a sentence. Self-attention lets every position in the input look at every other position simultaneously and decide how much to "attend" to each one. This is how a model knows that "it" in "The animal did not cross the street because it was too tired" refers to "animal" and not "street." That 2017 paper, "Attention Is All You Need," became one of the most cited in AI history and underpins GPT, BERT, and virtually every major language model today.
Question 4 of 4

Transfer learning is best described as which of the following?

AMoving a trained model from one server to another without retraining
BUsing labelled data from one domain to automatically label data in another
CRetraining a model from scratch whenever the task changes
DTaking a model pre-trained on a large task and adapting it for a specific, smaller one
Explanation
Training a model from scratch requires enormous amounts of data and compute. Transfer learning sidesteps this by starting from a model that has already learned rich, general representations from a large dataset (like ImageNet for vision, or a web-scale text corpus for language). You then fine-tune it on your specific task with far less data and compute. This is how most practical AI systems are built today. A medical imaging startup does not train from scratch — it starts from a model already trained on millions of general images and adapts from there.
🎉
Phase 4 Complete
4
out of 4

Excellent. You can explain the architectures that power most of the AI you encounter — a rare and genuinely useful foundation.

← Back to course