In this lesson
AI Ethics: Bias, Fairness and Accountability
The most technically impressive AI system can still cause serious harm. Understanding where things go wrong and why is now a core responsibility for every AI practitioner.
Why AI Ethics Matters Now
For much of its early history, AI was a research discipline with limited real-world reach. That has changed dramatically. AI systems now decide who receives a loan, who is flagged as a flight risk in court, whose job application advances, and which patients are prioritised for additional medical care. When these systems fail or discriminate, the consequences are not abstract. They are experienced by real people.
The field of AI ethics is not about making AI "nice." It is about ensuring that the power of these systems is exercised responsibly, that harms are identified before deployment rather than discovered after, and that the people most affected by AI decisions have meaningful recourse when things go wrong.
This is not a soft skill or an optional extra. It is increasingly a legal requirement, a professional expectation, and a practical necessity for building systems that people will actually trust and use.
Many people assume that because an AI system is trained on data rather than programmed with human prejudices, it must be objective. This is wrong. Data is generated by human societies, and human societies carry historical inequalities. A model trained on biased data learns those biases. Automation does not neutralise discrimination. In some cases, it amplifies it and makes it harder to challenge.
Where Bias Comes From
Bias in AI systems does not usually arise from malicious intent. It enters the pipeline in several distinct, often invisible ways. Understanding where it enters is the first step to addressing it.
Bias can enter at every stage of the AI pipeline. Addressing it at deployment alone is too late. Rigorous checks are needed at each step, with particular attention to the feedback loop that can compound errors over time.
Historical Bias
The world has not always been fair. Training on historical records encodes historical inequalities. A hiring model trained on past hiring decisions inherits the biases of past hiring managers.
Representation Bias
When training data underrepresents certain groups (women in STEM datasets, darker skin tones in medical imaging), the model performs worse for those groups.
Measurement Bias
The variables used to measure a concept may not measure it equally across groups. Using healthcare cost as a proxy for healthcare need systematically underestimates the needs of groups who historically received less care.
Feedback Loop Bias
A predictive policing model directs more patrols to certain neighbourhoods, leading to more arrests there, which generates more data confirming the model's predictions. The bias becomes self-reinforcing.
Aggregation Bias
Training a single model on a diverse population can obscure important subgroup differences. A medical model trained on combined data may perform well on average but fail for specific demographic groups.
Deployment Bias
A model is used in a context it was not designed for. A risk score developed for one country's legal system is applied in another with very different social conditions and base rates.
Real Case Studies
Examining specific documented cases is essential. Abstract discussions of bias are less instructive than concrete examples of how it has caused real harm.
COMPAS Recidivism Risk Tool
COMPAS (Correctional Offender Management Profiling for Alternative Sanctions) is a commercial tool used by courts across the United States to predict the likelihood that a defendant will re-offend. In 2016, investigative journalists at ProPublica published an analysis of 7,000 defendants in Broward County, Florida.
Their analysis found that the tool was significantly more likely to incorrectly label Black defendants as high risk when they did not go on to re-offend, and more likely to incorrectly label white defendants as low risk when they did. The tool's overall accuracy was similar across racial groups, but the types of errors were distributed very differently.
Northpointe, the company behind COMPAS, responded that the tool satisfied a different statistical definition of fairness: that its risk scores meant the same thing regardless of race (that is, a score of 7 corresponded to the same re-offending rate for Black and white defendants). Both claims were true simultaneously. This case launched a now-famous debate in the research community about which mathematical definition of fairness is most appropriate, and whether they can all be satisfied at once when re-offending rates differ between groups.
Amazon's Automated Recruiting Tool
Amazon developed an AI recruiting tool that scored job applicants on a scale of one to five stars. The tool was trained on resumes submitted to Amazon over a ten-year period, a period during which the tech industry was overwhelmingly male. The model learned to penalise resumes containing the word "women's" (as in "women's chess club") and downgraded graduates of all-women's colleges.
Amazon's team discovered these issues in 2015, made corrections, but found they could not guarantee the tool would not find other problematic proxies for gender. The project was scrapped in 2018. Reuters reported on the case publicly that year.
Racial Bias in a Medical Care Algorithm
A 2019 study published in Science by Obermeyer et al. analysed a widely deployed commercial algorithm used by US health systems to identify patients who needed extra care management. The algorithm used predicted healthcare costs as a proxy for healthcare need. The researchers found that for the same healthcare cost level, Black patients were significantly sicker than white patients.
The root cause was a measurement bias: Black patients historically received less care for the same conditions, so their costs were lower even though their underlying needs were the same. The algorithm interpreted lower costs as lower need, and as a result, at any given risk score threshold, Black patients who were enrolled in care management programmes were considerably sicker than their white counterparts. The authors estimated this reduced the fraction of Black patients receiving additional care by more than half compared to what an unbiased algorithm would have produced.
Gender Shades: Disparities in Face Analysis
Joy Buolamwini (MIT Media Lab) and Timnit Gebru (Microsoft Research at the time) published "Gender Shades: Intersectional Accuracy Disparities in Commercial Gender Classification" at the 2018 ACM Conference on Fairness, Accountability, and Transparency. The study evaluated three major commercial gender classification systems from IBM, Microsoft, and Face++.
Across all three systems, overall accuracy ranged from 93.7% to 99.1%. But when broken down by skin tone and gender together, darker-skinned female subjects were classified correctly only 65.3% to 79.9% of the time. The disparity was intersectional: darker-skinned women fared far worse than lighter-skinned men or either group considered alone. The training datasets used by these systems contained disproportionately more lighter-skinned subjects.
After the study's publication, all three companies updated their systems and reported substantially improved performance across demographic groups. The study directly motivated Amazon, IBM, and Microsoft to later pause or limit the sale of facial recognition technology to police departments while accuracy and fairness standards were being debated.
Fairness is Not One Thing
The COMPAS case highlighted a genuine mathematical tension. Researchers Chouldechova (2017) and Kleinberg, Mullainathan, and Raghavan (2016) independently proved that several commonly used fairness criteria are mathematically incompatible with each other whenever the positive outcome rate differs between groups. This is sometimes called the impossibility theorem of fairness.
There is no single correct definition of "fair" that a model can simultaneously satisfy. Different definitions reflect different values, and the choice between them has significant ethical implications.
| Fairness Definition | What It Requires | When It Prioritises |
|---|---|---|
| Demographic Parity (Statistical Parity) |
Equal positive prediction rates across all groups. If 30% of Group A receives a loan, 30% of Group B should too. | Equal representation of outcomes across groups, regardless of underlying differences in the predictor variables. |
| Equalized Odds | Equal true positive rates AND equal false positive rates across groups. The model should be equally accurate for everyone. | Equal quality of predictions. If you are incorrectly flagged as high risk, that probability should not depend on your group membership. |
| Calibration (Predictive Parity) |
A predicted probability of 70% should correspond to a 70% actual outcome rate for all groups. Scores mean the same thing across groups. | Consistency of interpretation across groups. This is what Northpointe argued COMPAS satisfied. |
| Individual Fairness | Similar individuals should receive similar predictions, regardless of group membership. | Treating each person based on their own features rather than their demographic group average. |
| Counterfactual Fairness | Would the prediction change if the individual's protected attribute changed, all else being equal? | Causal reasoning about how the model uses protected attributes or their proxies. |
When the base rates of an outcome differ between groups (for example, if recidivism rates differ between demographic groups for historical reasons), you cannot simultaneously satisfy calibration and equalized false positive rates. You must choose. This is not a technical problem that better algorithms will solve. It is a societal question about which type of error is more acceptable, and who bears the cost of being wrong. The answer should come from democratic deliberation and input from affected communities, not from engineers optimising a loss function.
Explainability and Accountability
When an AI system makes a decision that affects someone's life, that person has a legitimate interest in understanding why. This is not just a philosophical point. The EU's General Data Protection Regulation (GDPR) includes provisions granting individuals rights related to automated decision-making, requiring that decisions be explainable in meaningful terms.
Interpretable Models vs Post-Hoc Explanations
Some models are inherently interpretable. A logistic regression or a shallow decision tree can be read directly by a human expert. The model's reasoning is transparent by design. Cynthia Rudin (Duke University) has argued forcefully, particularly in a 2019 paper in Nature Machine Intelligence, that in high-stakes domains (criminal justice, medical diagnosis, credit lending), we should prefer inherently interpretable models over complex black boxes with post-hoc explanations, because post-hoc explanations are approximations that may not accurately represent what the model actually computed.
Deep neural networks are not inherently interpretable. Their reasoning is distributed across millions of parameters. To explain them, researchers have developed post-hoc explanation methods:
SHAP (2017)
SHapley Additive exPlanations, developed by Lundberg and Lee. Uses game theory (Shapley values from cooperative game theory) to assign each input feature a contribution score for a specific prediction. Theoretically grounded and consistent.
LIME (2016)
Local Interpretable Model-agnostic Explanations, by Ribeiro, Singh, and Guestrin. Approximates the complex model locally around a specific prediction using a simple linear model. Intuitive but the local approximation may not always be faithful.
Grad-CAM (2017)
Gradient-weighted Class Activation Mapping, by Selvaraju et al. For image classification CNNs, produces a heatmap showing which regions of the image most influenced the prediction. Widely used in medical imaging to verify model reasoning.
Attention Visualisation
For Transformer models, plotting attention weights shows which tokens a model focused on when generating an output. Useful but research has shown attention weights do not always correspond to what actually drives predictions.
Post-hoc explanations are approximations, not ground truth. A SHAP explanation tells you how a simplified model accounts for the prediction, not necessarily how the actual model computed it. Using these tools uncritically, or communicating them to users as if they were definitive, can create a false sense of transparency. Explanations are a starting point for scrutiny, not a substitute for it.
Regulation and Governance
Governments and international bodies have begun to respond to the risks of AI with regulatory frameworks. Understanding the landscape is becoming an important part of AI practice, particularly for anyone building systems that will be deployed commercially.
EU AI Act (2024)
The world's first comprehensive AI law, entering into force August 2024. It takes a risk-based approach: systems are classified by risk level (unacceptable, high, limited, minimal). High-risk systems (credit scoring, recruitment, critical infrastructure) face strict obligations including conformity assessments, transparency requirements, and human oversight mandates. AI systems that pose unacceptable risks (social scoring by governments, real-time remote biometric surveillance in public) are banned.
EU GDPR (2018)
While not AI-specific, the General Data Protection Regulation contains provisions directly relevant to AI: Article 22 gives individuals the right not to be subject to solely automated decisions with significant effects, and the right to a meaningful explanation. It also requires data minimisation, purpose limitation, and privacy by design. GDPR applies to any organisation processing data of EU residents.
UNESCO Recommendation (2021)
The first global standard on AI ethics, adopted by all 193 UNESCO member states. It covers human rights, transparency, accountability, fairness, environmental sustainability, and gender equality in AI. It is non-binding but represents the broadest international consensus on AI ethics principles.
US Executive Order on AI (2023)
Signed October 2023, it directed federal agencies to establish standards for AI safety and security, protect privacy, promote equity and civil rights, and support workers affected by AI. It required developers of powerful AI systems to share safety test results with the government. The Biden-era order was partially rescinded in early 2025, reflecting ongoing policy debates about AI governance in the US.
Beyond government regulation, voluntary frameworks have emerged from the research community. Model Cards (proposed by Timnit Gebru, Margaret Mitchell et al. in 2018) are structured documentation forms for ML models that disclose performance characteristics across demographic subgroups. Datasheets for Datasets (Gebru et al., 2021) provide an analogous framework for documenting training data. These tools have been widely adopted by major AI labs as part of responsible release practices.
What You Can Do as an AI Practitioner
Ethics is not only for policymakers and researchers. Every person who builds, deploys, or recommends AI systems has a responsibility to consider these issues. Here are concrete steps you can take at each stage of an AI project.
Audit your training data
Before training, examine the demographic composition of your dataset. Ask: who is represented? Who is not? What are the base rates of the outcome variable across subgroups? Document your findings.
Disaggregate your evaluation metrics
Never report only overall accuracy. Compute precision, recall, and F1 separately for each demographic subgroup relevant to your deployment context. Disparities that aggregate metrics hide become visible this way.
Question your proxy variables
Ask whether the variable you are using to measure the concept you care about actually measures it equally across groups. If it does not, consider whether an alternative measurement is available.
Involve affected communities
The people most likely to be impacted by your system should have input into its design, evaluation, and deployment conditions. This is not just ethical. It is also practically useful for identifying failure modes.
Document your models
Use a structured format like a Model Card to record what your model was trained on, what it was evaluated on, its intended use cases, its known limitations, and its performance across demographic groups.
Design for human oversight
For high-stakes decisions, build in meaningful human review rather than full automation. Ensure that humans in the loop have sufficient information, time, and authority to actually override the system when needed.
Further Reading
Key Takeaways
- AI systems trained on historical data inherit historical inequalities. Automation does not eliminate bias. In some cases it amplifies and obscures it.
- Bias enters the pipeline at multiple points: data collection, labelling, model training, evaluation, and deployment. It must be addressed at every stage, not just at the end.
- The COMPAS, Amazon hiring, healthcare cost proxy, and Gender Shades cases are documented, peer-reviewed examples of algorithmic harm. They are not edge cases. They are illustrations of systematic risks in routine AI deployment.
- There is no single definition of fairness. Demographic parity, equalized odds, and calibration are mathematically incompatible when base rates differ. Choosing between them is a value judgment, not a technical optimisation problem.
- Post-hoc explainability tools (SHAP, LIME, Grad-CAM) are useful but imperfect approximations. In high-stakes settings, inherently interpretable models may be more appropriate than black boxes with explanations.
- The EU AI Act (2024) is the world's first comprehensive AI law. It classifies AI systems by risk level and imposes significant obligations on high-risk applications. Regulatory knowledge is becoming a professional requirement.
- Practical steps for ethical AI practice include disaggregating evaluation metrics by subgroup, questioning proxy variables, involving affected communities in design, documenting models with structured formats like Model Cards, and designing meaningful human oversight into high-stakes systems.
In Lesson 5.2, you will go inside ChatGPT, Claude, and similar systems. How are they actually trained? What are tokens? What does it mean for a model to "hallucinate"? Why do these systems sometimes confidently say things that are completely false? And how do they scale to hundreds of billions of parameters?
Going Deeper
Want to go further on AI Ethics? Start here.
Reflect
Before you move on
No right answers here. These questions are for you.
A facial recognition system achieves 95% accuracy overall but only 73% accuracy on darker-skinned women. The company says the system is "highly accurate." What is wrong with that claim?
Aggregate accuracy hides unequal harm. A system that works well for the majority group and poorly for a minority can still report impressive headline numbers. The people most likely to be wrongly identified are already among the most vulnerable. Ethical evaluation requires disaggregating performance by demographic group, not just reporting a single score.
A hospital AI recommends denying a patient's insurance claim. The AI is a black box. Who is responsible if the decision turns out to be wrong?
Responsibility does not disappear because a machine made the recommendation. The hospital that deployed the system, the company that built it, and the clinician who accepted the output without scrutiny all bear some degree of accountability. AI does not create a responsibility vacuum; it redistributes it in ways institutions must think carefully about before deployment.
A highly accurate medical AI is also completely uninterpretable. A less accurate model can explain every decision in plain language. Which would you deploy in a hospital, and why?
There is no single right answer, and that is the point. A black-box model may save more lives on average, but its errors are invisible and impossible to contest. An interpretable model allows clinicians to catch mistakes, builds trust, and supports patient rights. Many jurisdictions require explainability for high-stakes decisions. The tradeoff is real, and the right choice depends on context, stakes, and who bears the consequences of error.