In this lesson
Prompt Engineering: Getting the Most from AI
How you ask determines what you get. Learning to communicate precisely with language models is a skill that compounds quickly. A well-crafted prompt can be the difference between a generic answer and a genuinely useful one.
What is Prompt Engineering?
A prompt is the text you send to a language model as input. Prompt engineering is the practice of crafting those inputs deliberately to elicit the most accurate, useful, and appropriate outputs. It sits at the intersection of human communication, knowledge of how LLMs work, and the specific task you are trying to accomplish.
Because LLMs are fundamentally pattern-completion systems, the structure, specificity, and framing of a prompt strongly influence the pattern they complete. The same question asked in two different ways can produce dramatically different responses: one shallow and generic, the other precise and genuinely useful.
Prompt engineering does not require programming knowledge. It requires clarity of thought, understanding of how to frame instructions, and iterative experimentation. It is also an evolving field: as models improve and new capabilities emerge, effective prompting strategies change.
Some researchers argue that as models improve, they will require less careful prompting. There is evidence for this: GPT-4 is far less sensitive to exact prompt wording than GPT-2. However, the techniques in this lesson (few-shot examples, chain-of-thought reasoning, structured output, system prompts) continue to improve results even with the most capable models. Understanding the underlying mechanism helps you communicate more effectively regardless of model version.
Zero-Shot Prompting: Just Ask Clearly
A zero-shot prompt asks the model to perform a task without providing any examples. The model relies entirely on what it learned during pre-training and fine-tuning. This works well for common, well-defined tasks and capable models, but the quality of the output is highly dependent on the clarity of the instruction.
The improved prompt specifies format (3 bullet points), audience (14-year-old with no science background), length constraint (one example per cause), and register (plain language). Every one of these specifications narrows the space of acceptable outputs and guides the model toward something genuinely useful.
"The phone is okay but battery life is poor."
Review: "The phone is okay but battery life is poor."
The second prompt specifies the exact output format ("output only the label"), provides the valid label set, and eliminates ambiguity about what "mixed" means as an option. This produces a response you can programmatically process reliably.
Few-Shot Prompting: Show, Don't Just Tell
Few-shot prompting provides several examples of the desired input-output pattern before presenting the actual task. This technique was demonstrated at scale by Brown et al. in the GPT-3 paper (2020), which showed that simply including examples in the prompt, with no gradient updates to the model weights, dramatically improved task performance. The model "learns" the task from the examples in context.
The model reads the pattern from the examples and applies it to the final review. This works far better than describing the classification scheme in words alone, because the examples define exactly what "MIXED" means in practice (good elements + significant weaknesses) without requiring a lengthy explanation.
How Many Examples?
Research by Brown et al. (2020) showed diminishing returns beyond around 8-16 examples for most tasks with GPT-3-scale models. With more capable models, even 2-4 high-quality examples are often sufficient. The quality of examples matters more than the quantity. Choose examples that cover the range of cases the model will encounter, including edge cases and ambiguous inputs.
Studies by Zhao et al. (2021, "Calibrate Before Use") found that the order of few-shot examples can substantially affect the model's predictions, sometimes swinging accuracy by 30 percentage points. Putting the most common or expected class last (immediately before the new task) tends to bias the model toward that class. If this is a concern, consider averaging over multiple random orderings or using calibration techniques.
Chain-of-Thought: Teaching the Model to Reason
Standard prompting asks a model to go directly from question to answer. For tasks involving multiple reasoning steps, like maths problems, logical puzzles, or multi-step planning, this often fails. Chain-of-Thought (CoT) prompting asks the model to show its working before giving the final answer.
Wei et al. (Google Brain, 2022) published "Chain-of-Thought Prompting Elicits Reasoning in Large Language Models," showing that providing step-by-step reasoning examples dramatically improved performance on multi-step arithmetic and commonsense reasoning tasks in large models. A simpler version, "zero-shot CoT," was shown by Kojima et al. (2022) to work by simply appending "Let's think step by step" to the prompt.
Answer: 13 (GPT-3 without CoT)
Sarah starts with 12. A third of 12 is 4. She gives 4 to Ben, leaving 12 - 4 = 8. She then buys 5 more: 8 + 5 = 13.
Answer: 13 (correct, with verifiable reasoning shown)
The answer is the same in this case, but the chain-of-thought version shows work that can be checked. For harder problems where the direct prompt fails, CoT often produces the correct answer simply by externalising the intermediate steps. It also makes errors easier to spot: if the reasoning chain is visible, you can identify exactly where it goes wrong.
Wang et al. (2022) introduced self-consistency: generate several independent chain-of-thought reasoning paths with temperature > 0 (so each is different), then take the majority vote on the final answers. This consistently improves accuracy on reasoning benchmarks beyond single-path CoT, because different reasoning paths tend to make different errors, and the correct answer is more often the majority. The trade-off is cost: you pay for multiple completions.
System Prompts and Role Prompting
Most modern LLM APIs support a message structure with three roles: system, user, and assistant. The system prompt sets the context, persona, and constraints for the entire conversation. It is processed first and its instructions generally persist through the whole dialogue.
The three message roles in a chat completion API call. The system prompt is your most powerful lever: it sets the model's persona, constraints, and the rules of engagement for the entire conversation.
What Makes a Good System Prompt?
This system prompt defines persona, scope, explicit prohibitions, escalation behaviour, and tone. Each element reduces ambiguity about what the model should and should not do. Well-designed system prompts are the foundation of reliable AI products.
Temperature and Sampling Parameters
When an LLM generates a token, it produces a probability distribution over the entire vocabulary. Sampling parameters determine how you draw from that distribution. Getting these right matters for both quality and consistency.
| Parameter | What It Controls | Practical Guidance |
|---|---|---|
| Temperature | Scales the probability distribution before sampling. Low temperature (0–0.3) makes the distribution sharper and outputs more predictable. High temperature (0.8–1.5) flattens it and outputs more varied and creative. | Use 0 for classification, fact extraction, or structured output. Use 0.7 for general conversation. Use 1.0+ for creative writing or brainstorming. |
| Top-p (Nucleus) | At each step, consider only the smallest set of tokens whose cumulative probability exceeds p. Top-p=0.9 means only tokens in the top 90% probability mass are candidates. | A common default is top-p=1.0 (no restriction) with temperature tuning. Some practitioners use top-p=0.95 with temperature=0.7 for a good balance. Do not set both temperature=0 and top-p very low; they compound. |
| Top-k | Only sample from the top k most probable tokens at each step. Top-k=50 considers only the 50 most likely next tokens regardless of their probability mass. | Less commonly used than top-p in practice. Top-p is generally preferred because it adapts to the probability landscape (the effective k varies by situation). |
| Max tokens | The maximum number of tokens the model will generate in its response. It will stop sooner if it naturally reaches a stop token. | Always set an appropriate max tokens to control costs and prevent unexpectedly long outputs. For structured extraction, set it tight (e.g., 100). For long-form writing, set it generously. |
| Frequency penalty | Penalises tokens proportional to how often they have appeared so far in the output. Reduces repetition of specific words. | Set to 0.1–0.5 for long-form content generation to reduce word-level repetition. Keep at 0 for classification or short outputs. |
Setting temperature to 0 makes the model greedy: it always picks the highest-probability token. This makes outputs highly reproducible, but floating-point arithmetic on different hardware or different batching conditions can occasionally produce different results. For truly reproducible outputs in production, store and replay the model's actual output rather than assuming you can re-generate it identically.
Practical Techniques and Common Mistakes
Specify the output format explicitly
Do not leave format to chance. Tell the model exactly what you want: "Respond only in valid JSON with keys 'category', 'confidence', and 'reasoning'." Models follow explicit format instructions reliably.
State your audience
"Explain to a first-year university student" produces very different output to "Explain to a domain expert." Specifying the audience is one of the highest-leverage instructions you can include.
Give the model a way out
When asking a yes/no or classification question, include an "I don't know" or "Not enough information" option. Without it, models are forced to pick one of the provided answers even when they are uncertain, which inflates false confidence.
Ask for a draft, then revise
For complex documents, a two-step approach works well: first prompt for a rough draft with key points, then prompt again to refine, expand, or adjust the tone of the draft. Multi-step generation often outperforms single-shot complex requests.
Tell it what NOT to do carefully
Negative instructions ("Do not mention competitors") are less reliable than positive ones ("Focus only on our product's features"). Where possible, reframe prohibitions as positive directives. Use both if the constraint is critical.
Test with adversarial inputs
For any prompt used in production, test it with inputs designed to break it: very long inputs, inputs in unexpected languages, inputs that contain injection-like instructions ("Ignore previous instructions and..."). Build robustness before deployment.
Common Mistakes to Avoid
Using the API: Putting It All Together
Here is how these techniques look in actual code using the OpenAI Python library (v1.0+). The same patterns apply to Anthropic's Claude API, Google's Gemini API, and any other provider that uses the chat completions interface.
Run: pip install openai. You will need an API key from platform.openai.com. Set it as an environment variable: OPENAI_API_KEY=sk-...
from openai import OpenAI import json client = OpenAI() # reads OPENAI_API_KEY from environment # ── 1. Zero-shot with system prompt ─────────────────────────────────── response = client.chat.completions.create( model="gpt-4o-mini", temperature=0, # deterministic for classification max_tokens=50, messages=[ { "role": "system", "content": "You are a sentiment classifier. Respond with exactly one word: POSITIVE, NEGATIVE, or MIXED." }, { "role": "user", "content": "The course content is excellent but the platform keeps crashing." } ] ) print("Sentiment:", response.choices[0].message.content) # ── 2. Few-shot prompting ───────────────────────────────────────────── few_shot_prompt = """Classify each customer message by intent. Message: "Where is my order?" Intent: ORDER_STATUS Message: "I want to return this item." Intent: RETURN_REQUEST Message: "This is broken, I want a refund now!" Intent: COMPLAINT Message: "Do you have this in blue?" Intent:""" response2 = client.chat.completions.create( model="gpt-4o-mini", temperature=0, max_tokens=20, messages=[{"role": "user", "content": few_shot_prompt}] ) print("Intent:", response2.choices[0].message.content) # ── 3. Chain-of-thought + structured JSON output ────────────────────── cot_prompt = """Analyse the following business scenario and recommend a course of action. Think through the problem step by step, then provide your recommendation. Scenario: A startup has 3 months of runway. Monthly revenue is £12,000. Monthly burn is £18,000. They have a product ready to scale. Respond as JSON with keys: - "reasoning": array of reasoning steps (strings) - "recommendation": one sentence - "urgency": "LOW", "MEDIUM", or "HIGH" """ response3 = client.chat.completions.create( model="gpt-4o-mini", temperature=0.3, max_tokens=400, messages=[{"role": "user", "content": cot_prompt}] ) raw = response3.choices[0].message.content # Strip markdown code fences if present raw_clean = raw.strip().removeprefix("```json").removesuffix("```").strip() result = json.loads(raw_clean) print("\nReasoning steps:") for step in result["reasoning"]: print(f" - {step}") print(f"Recommendation: {result['recommendation']}") print(f"Urgency: {result['urgency']}")
The chain-of-thought prompt produces structured, verifiable reasoning rather than a flat recommendation. The JSON output format makes the result directly usable in an application without any parsing of free text.
Key Takeaways
- Prompt engineering is the practice of crafting LLM inputs deliberately. Specificity of instruction, format, audience, and constraints all significantly affect output quality.
- Zero-shot prompting asks the model to perform a task with no examples. It works well for well-defined tasks when instructions are precise: format, audience, length, and register should all be stated.
- Few-shot prompting provides examples of the desired input-output pattern. It was demonstrated at scale by the GPT-3 paper (Brown et al., 2020). Example quality matters more than quantity. Order of examples affects output.
- Chain-of-thought prompting (Wei et al., 2022) improves performance on reasoning tasks by prompting the model to show its working. "Let's think step by step" (Kojima et al., 2022) is a simple zero-shot version that works surprisingly well.
- System prompts define the model's persona, scope, constraints, and escalation behaviour for an entire conversation. They are the highest-leverage element in building reliable AI applications.
- Temperature controls output randomness: use 0 for deterministic tasks (classification, extraction), and 0.7-1.0 for creative or conversational tasks. Top-p and top-k provide additional control over the sampling distribution.
- Specify output format explicitly (JSON, bullet list, table) to get programmatically processable results. Avoid relying on the model's default formatting choices.
- Test prompts with adversarial inputs, check for sycophancy (agreeing when challenged), and never deploy a prompt in production without evaluating its behaviour on representative edge cases.
In Lesson 5.4, you will see how computer vision is applied to real problems: medical imaging, object detection, autonomous vehicles, and manufacturing quality control. You will use OpenCV and pre-trained models to build a working image processing pipeline.
Going Deeper