When AI is unfair
Who is missing from the examples?
A green mango that is ripe
A fruit machine has learned mostly from yellow Dasheri mangoes. A Langra mango stays green even when it is ripe. What will the machine probably say about a ripe Langra?
What this lesson covers
The idea
Bias in AI often comes from the data it learned from: whoever is missing from the examples is served worse.
A green mango that is ripe
A fruit machine has learned mostly from yellow Dasheri mangoes. A Langra mango stays green even when it is ripe. What will the machine probably say about a ripe Langra?
- Ripe, because the machine will check how soft it feels
- Not ripe, because every green mango it saw was unripe
- It will refuse to answer
Teach the machine fairly
A fruit stall wants a very simple machine that says if a mango is ripe. The crate has mostly Dasheri mangoes and a few Langra.
Pick 6 mangoes to teach it. Then see how it does on new mangoes of each kind.
Who was missing?
Nobody told the machine to be unfair. It only follows patterns in its examples. Without any ripe Langra among them, it learns that green means unripe and gets Langra wrong. Your Langra examples fixed that.
Real tools can have the same problem. A camera or voice tool trained mostly on some people's faces or voices may work worse for other people.
Bias in AI often comes from the data it learned from: whoever is missing from the examples is served worse.
Results that serve some groups better than others are called bias. It can also come from a system's design or from human choices.
Notes
Bias in AI often comes from the data it learned from: whoever is missing from the examples is served worse.
Before you trust a tool, ask: who is missing from its examples?
Check yourself
A tool sorts forms for a science camp. It learned from ten years of choices, when almost no village school children were picked. Now it gives low scores to village forms. What is the most likely reason?
Meera's plant app learned from 200 photos of tomato leaves and only 5 of chilli leaves. It is often wrong about chilli leaves. What is the best fix?
A team's app is right 95 times out of every 100 overall. Is that enough to say it works well for everyone?
- A machine has no opinions of its own, so its scores cannot be unfair. A machine has no opinions, but it copies the patterns in its examples. If past choices left children out, the scores can too.
- It learned from past choices that left out village children — correct. Yes. Few village children were among the chosen examples. The tool copied that pattern, so it serves them worse.
- The village children's project ideas must have been weaker. Nothing here says that. A tool with few examples of village children cannot score their ideas fairly, however good they are.
- Add many more tomato photos to the examples. More of the same does not help the missing group. The app needs to see more chilli leaves.
- Stop testing it on chilli leaves, since it fails there. Not testing hides the problem. It does not fix it, and farmers who grow chillies still get poor answers.
- Add many more chilli leaf photos, then test again — correct. Yes. Add examples of the group that is missing, then test that group again to check the fix.
- No, a small group could still be missed, so test each group — correct. Yes. A high average can hide a small group that is often wrong. Testing each group shows who is missed.
- Yes, because 95 out of 100 is a very high score. It is a good score overall. But if most examples came from one group, a few others could still be wrong most of the time.
- No, an app is only fair if it scores 100 out of 100. No app is perfect. The aim is fair results for each group, not an impossible score.