In this lesson
Phase 5 · Lesson 5.2

Large Language Models: How They Actually Work

What is actually happening inside ChatGPT, Claude, or Gemini? What is a token? Why do these systems hallucinate? And how does training on next-word prediction produce something that can write code, explain science, and hold a conversation?

🕑 35 min read 📊 3 visualisations 💻 Python demos

What is a Large Language Model?

A Large Language Model (LLM) is a Transformer-based neural network trained on enormous quantities of text to predict the next token in a sequence. That is the entire training objective. From this deceptively simple task, these models develop a remarkable range of capabilities: they write, summarise, translate, reason, code, and converse.

"Large" refers to two things simultaneously: the number of parameters in the model (ranging from billions to trillions) and the volume of training data (ranging from hundreds of billions to trillions of tokens of text). Both scale dimensions matter significantly to performance, and the relationship between scale and capability has been studied carefully.

What makes LLMs distinct from earlier language models is the combination of Transformer architecture, massive scale, and a training approach that requires no hand-labelled data at all. The supervision signal comes entirely from the natural structure of text: given the previous words in a sentence, the next word is the label. This means the entire written record of humanity, once digitised, becomes a training dataset.

These Are Probabilistic Systems

LLMs do not "think" in the way humans do, and they do not "look up" answers in a database. At each step, they compute a probability distribution over the entire vocabulary and sample from it. The output you see is a sequence of such samples. This is why the same prompt can produce different outputs on different runs, and it is also why understanding probability and uncertainty is essential for working with these systems responsibly.

Tokens: The Language of LLMs

LLMs do not process text character by character or word by word. They process tokens: subword units produced by a tokenisation algorithm. A token is roughly 4 characters of English text on average, though this varies significantly by language and content type.

The dominant tokenisation method is Byte-Pair Encoding (BPE), introduced by Sennrich, Haddow, and Birch (2016) for machine translation. BPE works by starting with individual characters and iteratively merging the most frequent adjacent pairs until reaching the desired vocabulary size. Common words become single tokens. Rare words are split into recognisable subword pieces. This allows models to handle any text including invented words or misspellings, since any string can be decomposed into the byte-level characters that form the base of the vocabulary.

Input text: "Convolutional neural networks learn hierarchically."
Tokens (each colour = one token):
Conv olution al neural networks learn hier arch ically .
10 tokens for 50 characters. "Convolutional" splits into 3 tokens because it appears rarely as a whole word. "neural" and "networks" are common enough to each be one token.
Input: "ChatGPT" → Tokens:
Chat G PT
3 tokens. The model did not know "ChatGPT" as a whole word during its vocabulary training (it is a product name), so it splits into familiar substrings.
Input: "1 + 1 = ?" → Tokens:
1 + 1 = ?
5 tokens. Numbers and symbols are often their own tokens. This means arithmetic is not natural for LLMs — they process digits as discrete symbols, not as numerical values.

Tokenisation also explains why LLMs can struggle with certain tasks that seem easy to humans. Counting letters in a word is hard because the model sees token-level units, not characters. Arithmetic is hard because numbers are symbol sequences, not magnitudes. Rhyming is hard because the model does not hear sounds, it sees token identities.

Vocabulary Size

GPT-2 has a vocabulary of 50,257 tokens. GPT-3 and GPT-4 use a similar vocabulary size. BERT uses a WordPiece vocabulary of approximately 30,000 tokens. A larger vocabulary means fewer tokens per sentence (more efficient) but a larger embedding matrix. The vocabulary is fixed after tokenisation training and does not change when the language model itself is trained.

Pre-Training at Scale: The Next-Token Prediction Game

The pre-training objective is called causal language modelling: given all previous tokens in a sequence, predict the next one. This is done for every position in every sequence in the training data, simultaneously, using the Transformer decoder architecture with causal masking (as covered in Lesson 4.4).

LLM Pre-Training: Predicting the Next Token "The" "cat" "sat" "on" Input context (all previous tokens) Transformer Decoder-only LLM Probability over 50,000+ tokens "the": 0.42 "the" Model predicts next token True next token Loss = cross-entropy between prediction and truth. Backprop updates all parameters. Repeat for trillions of tokens across all of training data.

At each training step, the model sees a context of previous tokens and predicts the next one. The loss is computed against the actual next token. This process runs across trillions of examples with no human labelling required.

GPT-3 was trained on approximately 300 billion tokens drawn from Common Crawl (web text), WebText2 (curated web), Books1, Books2, and Wikipedia. GPT-4's training data composition was not publicly disclosed. The key insight of pre-training is that predicting natural language is an enormously information-rich task. To predict the next word in a sentence accurately, you need some model of syntax, semantics, world knowledge, common sense, cause and effect, and much more. The training signal pushes the model to encode all of this implicitly in its parameters.

Training Cost

Training GPT-3 (175 billion parameters) was estimated to cost approximately $4.6 million in compute at 2020 cloud prices (Strubell et al. 2020 methodology applied to GPT-3). GPT-4 scale training is estimated to have cost over $100 million. These numbers explain why only a small number of well-resourced organisations can train frontier LLMs from scratch, and why fine-tuning and transfer learning are so important for everyone else.

Scaling Laws: Why Bigger (Often) Means Better

In 2020, a team at OpenAI (Kaplan et al.) published a landmark paper showing that LLM performance on language modelling tasks follows remarkably clean power law scaling with model size (parameters), dataset size (tokens), and compute budget. Doubling the model size on the same data improves performance in a predictable, measurable way. This allowed researchers to forecast, before training, how well a model would perform.

In 2022, DeepMind researchers (Hoffmann et al., "Training Compute-Optimal Large Language Models," commonly called the Chinchilla paper) refined these laws. They found that many existing large models, including GPT-3, were significantly undertrained: they had many parameters but not enough training tokens for those parameters. The optimal strategy, they argued, is to scale parameters and training tokens roughly equally. Chinchilla (70 billion parameters, 1.4 trillion tokens) outperformed GPT-3 (175 billion parameters, 300 billion tokens) on many benchmarks, despite being four times smaller.

Emergent Abilities (Wei et al., 2022)

Certain capabilities are absent in smaller models and appear suddenly at larger scales. Few-shot learning, multi-step reasoning, and code generation emerged at GPT-3 scale. These phase transitions were studied by Jason Wei et al. at Google Research in 2022. The emergence is not fully understood theoretically.

The Chinchilla Lesson

More parameters are not always better. A smaller model trained on more data can outperform a larger model trained on less data with the same compute budget. This has driven a shift toward training smaller, more thoroughly trained models: LLaMA-3 8B with 15 trillion tokens is more capable than GPT-3 with 175B parameters.

RLHF: From Raw Base Model to Helpful Assistant

A pre-trained language model is extraordinarily knowledgeable, but it is also unruly. Ask it a question and it may respond by generating more questions, writing fiction, continuing the sentence in unexpected directions, or outputting harmful content. It has learned to model all of human text, including text that is unhelpful, dishonest, or dangerous.

Converting a base model into a useful, safe assistant requires additional training. The dominant technique is Reinforcement Learning from Human Feedback (RLHF), applied at scale by the InstructGPT paper (Ouyang et al., OpenAI, 2022) and the technology that powers ChatGPT.

The RLHF Pipeline: Base Model to ChatGPT Step 1: SFT Supervised Fine-Tuning Human-written ideal responses as labels Step 2: Reward Model Humans rank model outputs Train RM to predict human preference scores Step 3: RL Fine-Tuning PPO optimises LLM to maximise reward model score + KL penalty Helpful, Harmless, Honest Assistant ~13K high-quality prompt-response pairs ~33K comparisons by human raters KL penalty prevents drifting from base model InstructGPT (Ouyang et al., OpenAI, 2022) — the paper behind ChatGPT

RLHF has three stages. SFT fine-tunes the base model on human-written demonstrations. A reward model is trained on human preference comparisons. Finally, RL (specifically Proximal Policy Optimisation) updates the LLM to generate outputs the reward model scores highly, while a KL divergence penalty keeps it from drifting too far from the base model.

A simpler alternative to RLHF called Direct Preference Optimisation (DPO) was proposed by Rafailov et al. (2023). DPO reframes the RLHF objective as a supervised classification problem, eliminating the need for a separate reward model training stage and the complexity of RL. Many open-source models released since 2023 have used DPO or variants of it.

Anthropic developed a complementary approach called Constitutional AI (Bai et al., 2022), which uses a set of principles (the "constitution") to guide the model to critique and revise its own outputs. This uses AI feedback (RLAIF) in place of some human feedback, making the process more scalable.

Hallucination: Why LLMs Confidently Say False Things

LLMs sometimes state things that are completely fabricated, with the same tone and apparent confidence they use for accurate information. They cite papers that do not exist, give wrong dates for real events, invent biographical details for real people, and produce code that looks correct but does not run. This phenomenon is called hallucination.

Understanding why it happens requires revisiting what LLMs actually do. At every step, the model computes a probability distribution over the vocabulary and samples a token. It does not "consult" a database of facts. It generates text that is consistent with the patterns in its training data. When asked about something it learned well, those patterns produce accurate-looking and accurate text. When asked about something rare, absent, or contradictory in its training data, those same patterns produce fluent but potentially incorrect text.

Factual Hallucination

Invented facts stated as true: wrong dates, nonexistent citations, incorrect statistics, invented names. The model is not lying. It is generating plausible-sounding text, and that text happens to be factually wrong.

Reasoning Errors

The model produces a confident chain of reasoning that contains a logical error. Each step looks plausible given the previous one, but the conclusion is wrong. This is especially dangerous because it is harder to spot than a simple factual error.

Sycophantic Hallucination

When a user suggests an incorrect answer, models fine-tuned with RLHF sometimes agree because agreement was rewarded during training. The model prioritises pleasing the human over accuracy. Studies have documented this pattern across multiple major LLMs.

Knowledge Cutoff Errors

Models have a training cutoff date. Events, papers, people, or products that appeared after the cutoff are unknown. When asked about them, the model either states uncertainty (good) or generates plausible-sounding but fabricated information (bad).

Mitigation: Retrieval-Augmented Generation (RAG)

The most widely deployed technical solution is Retrieval-Augmented Generation (RAG), introduced by Lewis et al. (Facebook AI Research, 2020). Instead of generating answers purely from the model's parametric memory, RAG first retrieves relevant documents from an external knowledge base (a database, the web, a document collection) and includes them in the model's context. The model then generates its answer grounded in the retrieved evidence.

This does not eliminate hallucination entirely, but it gives the model accurate, up-to-date source material to work from, and it makes verification easier because the sources can be cited. The search systems used by ChatGPT, Bing Chat (Copilot), and Perplexity AI all use versions of this approach.

The Practical Rule

Never rely on an LLM as the sole source for factual claims, particularly for consequential decisions in medicine, law, finance, or safety-critical applications. Treat LLM output as a knowledgeable first draft that requires verification. The fluency and confidence of the output is not correlated with its accuracy.

Context Windows and "Memory"

A context window is the maximum number of tokens the model can process in a single forward pass, including both the input (your prompt, any retrieved documents, conversation history) and the output (the response). Anything outside the context window is simply not seen by the model. LLMs have no persistent memory between separate conversations unless it is explicitly engineered.

ModelContext WindowApproximate pages of text
GPT-3 (2020)4,096 tokens~3 pages
GPT-4 (2023)8,192 / 32,768 / 128,000 tokens~6 / 25 / 96 pages
Claude 3 / 3.5 (2024)200,000 tokens~150 pages
Gemini 1.5 Pro (2024)1,000,000 tokens~750 pages or ~1 hour of video
GPT-4o (2024)128,000 tokens~96 pages

The computational cost of the Transformer's attention mechanism scales quadratically with sequence length (O(n²)), which is why very long context windows are technically challenging and expensive. Research into more efficient attention mechanisms (sparse attention, linear attention, state-space models like Mamba) is active precisely because context length is an important capability bottleneck.

"Lost in the Middle" Problem

A 2023 study by Liu et al. found that LLMs tend to perform best when relevant information appears at the beginning or end of a long context, and significantly worse when it appears in the middle. Simply having a long context window does not mean the model uses all of it equally. When building RAG or document Q&A systems, placing the most relevant retrieved content near the end of the prompt (just before the question) tends to improve results.

The LLM Landscape

As of mid-2025, the LLM landscape consists of a small number of frontier closed-source models and a growing ecosystem of high-quality open-weight models.

OrganisationNotable ModelsAccessKey Notes
OpenAI GPT-3 (2020), GPT-4 (2023), GPT-4o (2024) API / ChatGPT GPT-3 established LLMs. GPT-4 is multimodal (text + images). GPT-4o is faster and cheaper.
Anthropic Claude 1 (2023), Claude 2 (2023), Claude 3 / 3.5 (2024) API / Claude.ai Emphasis on safety and Constitutional AI. Claude 3 Opus matched or exceeded GPT-4 on many benchmarks. 200K context window.
Google DeepMind PaLM (2022), Gemini 1.0/1.5 (2023-2024) API / Bard / Gemini app Gemini 1.5 Pro introduced 1M token context window. Gemini Ultra 1.0 first model claimed to exceed human expert scores on MMLU benchmark.
Meta LLaMA (2023), LLaMA 2 (2023), LLaMA 3 (2024) Open weights (free for research and commercial use) The open-weight models that democratised LLM development. LLaMA 3 70B is competitive with many closed-source models.
Mistral AI Mistral 7B (2023), Mixtral 8x7B (2023) Open weights French startup. Mistral 7B punched well above its weight class. Mixtral introduced sparse Mixture-of-Experts architecture for efficient large-scale inference.
The Open vs Closed Model Debate

Open-weight models (where the model weights are publicly released) enable researchers, students, and small companies to run and fine-tune LLMs without paying API costs or sharing data with a provider. They enable inspection, auditing, and customisation. The trade-off is that open weights can also be fine-tuned to remove safety guardrails. Meta, Mistral, and others have argued that the research and economic benefits of openness outweigh the risks. OpenAI and Anthropic have argued the opposite for their most capable models. This debate is ongoing and consequential.

Key Takeaways

Coming Up: Prompt Engineering

In Lesson 5.3, you will learn how to communicate effectively with LLMs. Zero-shot prompting, few-shot examples, chain-of-thought reasoning, system prompts, and temperature control are all practical skills that can dramatically change the quality of outputs you get from these models.

Practice Notebook
Run this lesson's code in Google Colab
RAG from scratch, embeddings, GPT-2 generation · Free GPU · No setup
Open In Colab

Going Deeper

Want to understand LLMs at a deeper level? Start here.

Progress
Done with this lesson?
Mark it complete to track your progress.