Reference

AI Glossary

Every key term used across the 25 lessons, with plain English definitions and real-world examples. Click any term to expand it.

🔍 50 terms
A
AccuracyPhase 3
The percentage of predictions a model gets correct out of all predictions made. Accuracy = (Correct predictions) / (Total predictions). Simple and intuitive, but misleading when classes are imbalanced.
If a spam filter correctly labels 90 of 100 emails, its accuracy is 90%. But if 95 of those were ham, always guessing "not spam" would score 95% accuracy without learning anything.
Activation FunctionPhase 4
A mathematical function applied to the output of a neuron to introduce non-linearity. Without activation functions, a neural network would behave like a single linear equation no matter how many layers it has. Common choices: ReLU, sigmoid, tanh, softmax.
ReLU simply returns 0 for negative values and the number itself for positive values: f(x) = max(0, x). This lets networks learn complex, curved decision boundaries.
Attention MechanismPhase 4
A technique that lets a model focus on the most relevant parts of its input when producing each output. Attention computes a weighted sum of "values" based on the similarity between a "query" and a set of "keys." It is the core innovation behind the Transformer architecture.
When translating "The cat sat on the mat" to French, the model attends strongly to "cat" when generating "le chat" and to "mat" when generating "le tapis."
Artificial General Intelligence (AGI)Phase 1
A hypothetical AI system that can perform any intellectual task a human can, including learning new skills from scratch, applying reasoning across domains, and understanding context without additional training. AGI does not currently exist.
A human who can learn chess, write poetry, diagnose illness, and cook dinner demonstrates general intelligence. Current AI systems can only do one of these per model.
B
BackpropagationPhase 4
The algorithm that trains neural networks by computing how much each weight contributed to the error, then adjusting all weights in the direction that reduces that error. It works by applying the chain rule of calculus from the output layer backwards to the input layer.
Think of it like tracing blame for a wrong answer. If the network predicted "dog" but the answer was "cat," backprop figures out which neurons fired too strongly and dials them back.
Bias (model)Phase 3
In machine learning, "bias" has two meanings. (1) Statistical bias: the error from wrong assumptions in the model, causing it to miss relevant relationships in the data (underfitting). (2) AI fairness bias: systematic unfairness baked into a model because of skewed training data or labels.
A hiring algorithm trained on historical data where engineers were mostly male may learn to score male candidates higher, not because of skill, but because of historical bias in the data it learned from.
Batch SizePhase 4
The number of training examples the model sees before updating its weights. Larger batches give more stable gradient estimates but require more memory. Smaller batches are noisier but often generalise better and train faster per epoch.
A batch size of 32 means the model looks at 32 images, computes the average error, then adjusts its weights once before moving to the next 32 images.
C
ClassificationPhase 3
A type of supervised learning where the model predicts which category (class) an input belongs to. The output is a discrete label. Binary classification has two classes; multi-class has three or more.
Spam detection (spam / not spam) is binary classification. Handwriting recognition (0 through 9) is 10-class classification.
ClusteringPhase 3
An unsupervised learning technique that groups data points so that items within a group are more similar to each other than to items in other groups. No labels are required. Common algorithms include K-Means and DBSCAN.
A supermarket uses clustering on purchase data to discover that customers naturally fall into segments: families, students, and elderly shoppers. No one told the algorithm these categories exist.
Convolutional Neural Network (CNN)Phase 4
A neural network architecture designed for grid-structured data like images. It uses convolutional filters (small sliding windows) that learn to detect local patterns like edges, textures, and shapes. Pooling layers then downsample to reduce computation.
A CNN trained on faces first learns to detect edges, then eyes and noses, then whole faces by the deeper layers. Each layer builds on patterns found in the previous one.
Cross-ValidationPhase 3
A technique for evaluating model performance by splitting data into K folds, training K times (each time using a different fold as the validation set), and averaging the results. Gives a more reliable estimate than a single train/test split.
5-fold cross-validation trains five separate models. Each model sees 80% of the data for training and 20% for validation, but no two models use the same 20% for validation.
D
Deep LearningPhase 1
A subset of Machine Learning that uses neural networks with many layers (hence "deep") to learn hierarchical representations of data. Each layer transforms the data into a more abstract representation. Deep learning excels at unstructured data like images, audio, and text.
A deep learning model for speech recognition processes raw audio waveforms, extracts frequency patterns, identifies phonemes, and assembles words -- each level of abstraction handled by a different set of layers.
Decision TreePhase 3
A machine learning model that makes predictions by learning a flowchart of yes/no questions about the input features. Easy to interpret and explain. Prone to overfitting on their own, but powerful when combined into ensembles (like Random Forests).
A credit approval decision tree might ask: salary above 40k? Yes: approved. No: employed over 2 years? Yes: approved. No: rejected. Every path through the tree is a decision rule.
DropoutPhase 4
A regularisation technique where a random fraction of neurons are set to zero during each training step. This prevents neurons from co-adapting too closely, reducing overfitting and making the network more robust.
With dropout rate 0.3, each neuron has a 30% chance of being switched off for that batch. The network learns to not rely on any single neuron, which forces redundant representations.
E
EmbeddingPhase 5
A dense vector (list of numbers) that represents a word, sentence, image, or other object in a continuous space. Items that are semantically similar end up with vectors that are close together. Embeddings allow mathematical operations like finding meaning through subtraction and addition.
The classic example: vector("king") - vector("man") + vector("woman") is approximately equal to vector("queen"). The model has learned gender and royalty as directions in space.
EpochPhase 4
One complete pass through the entire training dataset. A model typically trains for multiple epochs so it can refine its weights repeatedly. Too few epochs leads to underfitting; too many can cause overfitting.
Training for 10 epochs on 50,000 images means the model processes each of those 50,000 images 10 times, updating its weights after every batch throughout all 10 passes.
F
F1 ScorePhase 3
The harmonic mean of precision and recall. F1 = 2 * (Precision * Recall) / (Precision + Recall). A single metric that balances both: it is high only when both precision and recall are high. Useful for imbalanced datasets where accuracy is misleading.
A disease screening test with precision 0.9 and recall 0.6 has F1 = 0.72. A perfect test would have F1 = 1.0. The score penalises large gaps between precision and recall.
Feature EngineeringPhase 2
The process of using domain knowledge to create, transform, or select input variables (features) that help a machine learning model learn better. Deep learning reduces but does not eliminate the need for feature engineering.
For a house price model, raw data might include a purchase date. Feature engineering might extract "age of house in years" and "was it built during a housing boom" as more informative features.
Fine-TuningPhase 4
Continuing to train a pretrained model on a new, smaller dataset so it adapts to a specific task. The model starts with knowledge from its original training and refines it rather than learning from scratch. Usually requires far less data and compute than training from scratch.
GPT-4 was pretrained on internet text, then fine-tuned on human feedback to become a helpful assistant. The fine-tuning data was much smaller but changed the model's behaviour dramatically.
G
Gradient DescentPhase 4
The optimisation algorithm used to train most machine learning models. It repeatedly adjusts the model's parameters in the direction that reduces the loss function, using the gradient (slope) of the loss to know which direction is "downhill."
Imagine you are blindfolded on a hilly landscape and want to reach the lowest point. You feel the slope under your feet and take a step downhill. Repeat thousands of times. That is gradient descent.
GeneralisationPhase 3
A model's ability to perform well on new, unseen data it was not trained on. Good generalisation means the model learned underlying patterns, not just memorised training examples. The gap between training performance and test performance measures generalisation.
A model that scores 99% on its training set but 62% on new examples has poor generalisation. It memorised the training data instead of learning the underlying concept.
H
HyperparameterPhase 3
A configuration value set before training that controls the learning process itself, rather than being learned from data. Examples: learning rate, batch size, number of layers, number of epochs, regularisation strength. Choosing good hyperparameters is often done by trial-and-error or grid search.
The number of trees in a Random Forest is a hyperparameter. You pick it before training. The split thresholds within each tree are parameters the model learns during training.
K
K-MeansPhase 3
An unsupervised clustering algorithm that partitions data into K clusters by iteratively assigning each point to its nearest cluster centre (centroid) and then moving each centroid to the average position of its assigned points. Sensitive to the initial centroid positions and the number K.
K-Means with K=3 on customer data might discover three natural customer segments. You must decide K in advance, which is why the elbow method and silhouette score are used to pick the right value.
L
Large Language Model (LLM)Phase 5
A very large neural network (billions of parameters) trained on vast amounts of text data to predict and generate language. LLMs learn grammar, facts, reasoning patterns, and style from the statistics of language. Examples: GPT-4, Claude, Gemini, LLaMA.
An LLM does not "know" facts the way a human does. It predicts what tokens are likely to follow based on patterns it saw during training. This is why they can sound confident but still be wrong.
Learning RatePhase 4
A hyperparameter that controls how large a step the model takes in the direction of the gradient during each weight update. Too large and the model overshoots the minimum and diverges. Too small and training takes extremely long or gets stuck.
A learning rate of 0.001 is common for Adam optimizer. Think of it as the size of each step you take when descending a hill in the dark. Too big and you step right over the valley. Too small and you are still walking after a week.
Loss FunctionPhase 4
A mathematical function that measures how wrong the model's predictions are. Training aims to minimise the loss. Different tasks use different loss functions: Mean Squared Error for regression, Cross-Entropy for classification, and many others.
If a model predicts a house costs 200,000 but the true price is 250,000, the MSE loss for that prediction is (200,000 - 250,000)^2 = 2,500,000,000. The model then adjusts weights to reduce this value.
M
Machine LearningPhase 1
A subset of AI where systems learn from data to improve performance on a task, without being explicitly programmed with rules for that task. The three main types are supervised learning, unsupervised learning, and reinforcement learning.
A spam filter built with rules requires a human to write every rule. An ML spam filter learns what spam looks like by studying thousands of examples, finding patterns a human never specified.
ModelPhase 3
In machine learning, a model is a mathematical function with learned parameters that maps inputs to outputs. The "model" is what is left after training is complete and it is ready to make predictions on new data.
After training on 10,000 chest X-rays, the resulting neural network weights form the "model." You can then give it a new X-ray and it will output a probability of pneumonia.
N
Narrow AIPhase 1
An AI system designed and trained to do one specific task or a small set of related tasks. All current commercially deployed AI systems are narrow AI. They cannot transfer knowledge to unrelated domains without retraining.
AlphaGo plays Go better than any human but cannot play chess, hold a conversation, or recognise a face. It is extremely narrow, extremely good, and nothing else.
Neural NetworkPhase 4
A computational model loosely inspired by the brain, consisting of layers of interconnected nodes (neurons). Each connection has a weight. The network learns by adjusting these weights so its outputs match the training labels. The foundation of deep learning.
A neural network with one input layer (pixel values), two hidden layers (learned patterns), and one output layer (class probabilities) can classify handwritten digits with over 99% accuracy.
NormalisationPhase 2
Scaling numerical features so they fall within a similar range (e.g. 0 to 1, or with mean 0 and standard deviation 1). Prevents features with large values from dominating the learning process and helps gradient descent converge faster.
If one feature is "age" (20-80) and another is "salary" (20,000-200,000), without normalisation the salary feature dominates. After scaling both to 0-1, the model treats them equally.
O
OverfittingPhase 3
When a model learns the training data too well, including its noise and random fluctuations, and therefore performs poorly on new data. An overfit model has memorised rather than generalised. Solutions include more data, regularisation, dropout, and early stopping.
A student who memorises every past exam question word-for-word will fail when the actual exam uses different wording for the same concept. They overfit to the training set (past papers).
P
ParametersPhase 4
The internal values of a model that are learned during training -- primarily the weights and biases of the neurons. "Number of parameters" is commonly used to describe the size of a model. GPT-4 is estimated to have over 1 trillion parameters.
A simple linear regression has 2 parameters: slope and intercept. ResNet-50 has 25 million. GPT-3 has 175 billion. More parameters generally means more capacity, but also more data and compute required.
PrecisionPhase 3
Of all the times the model predicted "positive," how many were actually positive? Precision = True Positives / (True Positives + False Positives). High precision means when the model says yes, it is usually right. Optimise for precision when false alarms are costly.
For cancer screening, high precision means few healthy people are incorrectly told they might have cancer. This reduces unnecessary panic and follow-up procedures.
Prompt EngineeringPhase 5
The practice of crafting inputs (prompts) to language models in ways that consistently produce better, more accurate, or more useful outputs. Techniques include few-shot examples, chain-of-thought instructions, role assignment, and output format constraints.
Instead of asking "summarise this article," a well-engineered prompt might say: "You are an editor at a science magazine. Summarise this article in 3 bullet points for a non-specialist reader aged 16. Use plain English and no jargon."
R
RAG (Retrieval-Augmented Generation)Phase 5
A technique that improves LLM answers by first retrieving relevant documents from a knowledge base, then giving the retrieved text to the model as context before generating a response. Reduces hallucination and allows up-to-date knowledge without retraining.
A customer support chatbot using RAG searches a company's FAQ database for the three most relevant articles, then asks the LLM to answer the question using only those articles. The model cannot make up an answer that contradicts company policy.
RecallPhase 3
Of all the actual positive cases, how many did the model correctly identify? Recall = True Positives / (True Positives + False Negatives). High recall means the model misses very few real positive cases. Optimise for recall when missing a positive case is costly.
For cancer screening, high recall means the model catches nearly every real cancer case, even if it flags some healthy patients for further testing. Missing a cancer is far worse than a false alarm.
RegressionPhase 3
A type of supervised learning where the model predicts a continuous numerical value rather than a category. The output is a number on a scale. Common algorithms include Linear Regression and Ridge Regression.
Predicting tomorrow's temperature in degrees Celsius is regression. Predicting whether tomorrow will be hot, warm, or cold is classification. The difference is continuous output vs discrete labels.
RegularisationPhase 3
Techniques that reduce overfitting by penalising model complexity. L1 regularisation (Lasso) pushes many weights to exactly zero. L2 regularisation (Ridge) shrinks all weights towards zero. Dropout is a regularisation technique specific to neural networks.
Ridge regression adds a penalty equal to the sum of squared weights to the loss function. This discourages the model from relying too heavily on any single feature and improves generalisation to new data.
Reinforcement LearningPhase 3
A type of machine learning where an agent learns to make decisions by taking actions in an environment and receiving rewards or penalties. The agent learns a policy that maximises cumulative reward over time. Used in game-playing AI, robotics, and recommendation systems.
AlphaGo learned to play Go by playing millions of games against itself. Each winning move reinforced good strategies; each losing move discouraged them. No human ever told it what a good move looks like.
S
Supervised LearningPhase 3
A type of machine learning where the model learns from labelled training examples: each input comes paired with the correct output. The model learns to map inputs to outputs by minimising the difference between its predictions and the true labels.
Training an email spam classifier on 10,000 emails labelled "spam" or "not spam" is supervised learning. The labels supervise the learning, telling the model when it is wrong.
SoftmaxPhase 4
An activation function used in the output layer of multi-class classifiers that converts raw scores (logits) into a probability distribution that sums to 1. The class with the highest probability is the model's prediction.
A 3-class model with raw outputs [2.0, 1.0, 0.1] is converted by softmax to approximately [0.67, 0.24, 0.09]. These are probabilities: 67% confidence in class 1, 24% in class 2, 9% in class 3.
T
TokenPhase 5
The basic unit of text that a language model processes. Tokens are not always whole words: they can be syllables, punctuation, or word fragments. GPT models split text into subword tokens. "Unbelievable" might be split into ["Un", "believ", "able"].
The sentence "I love AI!" might be tokenised into ["I", " love", " AI", "!"]. Four tokens, not three words. LLMs have a context window measured in tokens, which limits how much text they can process at once.
Transfer LearningPhase 4
Reusing a model trained on one task as the starting point for a different but related task. The pretrained model has already learned useful general features (edges, textures, language patterns) that do not need to be relearned from scratch on the new task.
A model trained on 1 million general images already knows how to detect edges, textures, and shapes. Fine-tuning it on 500 medical X-rays is far more effective than training from scratch on just those 500 images.
TransformerPhase 4
A neural network architecture introduced in 2017 ("Attention Is All You Need") built entirely on attention mechanisms. It processes all tokens in parallel (unlike RNNs) and uses self-attention to build context-aware representations. The foundation of GPT, BERT, and all modern LLMs.
Unlike an RNN that reads "The cat sat on the mat" word by word, a Transformer looks at all words simultaneously and computes how much each word should influence the meaning of every other word.
Training DataPhase 2
The labelled (or unlabelled) examples the model learns from during training. The quality, size, and diversity of training data is often the single biggest factor in model performance. Garbage in, garbage out.
GPT models were trained on hundreds of billions of words from the internet, books, and code. The breadth and quality of this training data is why they can write code, translate languages, and answer questions about history.
V
Validation SetPhase 3
A held-out portion of data (separate from training and test sets) used to tune hyperparameters and monitor for overfitting during training. It gives an honest signal of how the model is learning without contaminating the final test evaluation.
With 10,000 examples: 7,000 for training, 1,500 for validation (tune hyperparameters), 1,500 for test (final honest evaluation). Never use the test set until training is completely finished.
VectorPhase 2
An ordered list of numbers representing a point or direction in multi-dimensional space. In machine learning, features are represented as vectors; model weights are vectors; embeddings are vectors. Most of deep learning is matrix and vector arithmetic.
An RGB image pixel is a 3-dimensional vector [255, 128, 0] (red, green, blue). A word embedding might be a 768-dimensional vector. Cosine similarity between two embedding vectors measures how semantically similar two words or sentences are.
W
WeightPhase 4
A numerical value on a connection between neurons that determines how strongly one neuron influences the next. During training, weights are continuously adjusted to reduce prediction error. The "knowledge" of a neural network is encoded in its weights.
A weight of 0.8 on the connection from a "fur" detector neuron to a "cat" classifier means detecting fur is strong evidence for "cat." A weight near 0 means that connection has learned to be irrelevant.
No terms match ""