‹ Class 8 · Ch 3
Data and Fairness in AI · Principle 1 of 3

Where bias hides in data

Three places where unfairness can slip in.

Think

Where does unfairness start?

An app learns from old records to suggest after-school clubs. Its suggestions turn out unfair to some students. Where did the unfairness most likely come from?

What this lesson covers

The idea

Unfairness can enter an AI system through who is missing from the data, how the labels were chosen, and how the results are used.

Where does unfairness start?

An app learns from old records to suggest after-school clubs. Its suggestions turn out unfair to some students. Where did the unfairness most likely come from?

  • A mistake in the computer's sums
  • The old records it learned from
  • The students who asked for suggestions

Find where it slipped in

Each card describes an app that is unfair to some people. Tap the place where the unfairness slipped in.

A few cards are tricky. Tap a group. Read the reason, then press Next card.

Three ways in

None of those apps made an arithmetic error. The unfairness came from people's choices: who was left out, what counted as right, and how results were used.

So fixing it starts with questions about the data and its use, not with a faster computer.

Unfairness can enter AI through who is missing, how labels were chosen and how results are used.

An unfair pattern that a machine learns from data is called bias. Careless use can add unfairness even when the data is good.

Notes

Unfairness can enter AI through who is missing, how labels were chosen and how results are used.

When something looks unfair, ask: Who is missing? Who chose the answers? How is the result used?

Check yourself

Anil's class trains a mood app. One student labels every photo “happy” or “sad”, and he marks all serious faces “sad”. Where did the bias enter?

A canteen app gives fair meal suggestions in tests. Then a school uses its scores to give out trip seats, with no human review. What is the problem?

Priya receives a dataset of 5,000 student photos. Which question would help her find hidden bias fastest?

  • Who is missing from the data. The photos may show many different people. The trouble here is one student's idea of “sad”, which became the answer for every photo.
  • How the labels were chosen — correct. Yes. One person's view decided every answer, so the machine learned that serious faces are sad. Labels from several people would be fairer.
  • How the result is used. We do not know yet how the app will be used. The problem is already inside the examples it learned from.
  • The app's sums must be wrong, since the result is unfair. The sums can be right and the app can still cause harm. Here the app is being used for a job it was not made for.
  • Nothing is wrong, because the app passed all its tests. Tests only show that it works for the job it was tested on. A new use needs new checking.
  • It is used for a bigger decision than it was made for — correct. Yes. Unfairness can enter at the end, when a result is used carelessly. A person should review any decision that affects students.
  • Which groups appear in it, and which seem missing? — correct. Yes. Counting who is in the data shows thin or missing groups. Those are the first places bias hides.
  • How many photos does it have, in total?. A big dataset can still leave out whole groups. 5,000 photos of one kind of student is still one-sided.
  • Were the photos taken recently, or long ago?. New photos are useful, but a new dataset can still miss some people.
Hold to talk

Subscription Status