‹ Class 8 · Ch 3
Data and Fairness in AI · Principle 2 of 3

Balance the data

Fix thin groups with new examples, not copies.

Think

Whose voices are in the data?

A voice helper will serve a whole town. It learns from the voice clips below. Which group will it probably get wrong most often?

🧑 Age 18 to 59 100
In your data: 71%In real life: 50%
🧒 Under 18 30
In your data: 21%In real life: 30%
🧓 Age 60 and over 10
In your data: 7%In real life: 20%

Age 18 to 59 has too much of the data: 71% against 50% in real life.

The thin line on each bar marks the share in real life. More clips of a missing group must be newly collected, not copies of the ones you already have.

What this lesson covers

The idea

To make data fairer, collect new examples from the thin groups instead of copying, and remember that balanced counts alone do not guarantee fairness.

Whose voices are in the data?

A voice helper will serve a whole town. It learns from the voice clips below. Which group will it probably get wrong most often?

  • The 18 to 59 group
  • The under 18 group
  • The 60 and over group

Make the data match real life

Each + adds 10 voice clips from new people in that age group. Each − removes 10. Match every group to its mark for real life.

The shares are from a made-up town. You may collect up to 250 clips. Removing is allowed, but it throws data away.

New voices, not copies

Adding clips from new people moves a thin group toward its mark. Copying the same ten voices would raise the count, but teach the helper nothing new.

If you only removed clips, the shares match but little data is left. Balance is only a start too. One quiet room for every new child would still fail on a noisy bus.

To make data fairer, collect new examples from thin groups, not copies. Balanced counts help, but variety and quality matter too.

Data that reflects everyone a system will serve is called diverse and representative.

Notes

To make data fairer, collect new examples from thin groups, not copies. Balanced counts help, but variety and quality matter too.

Record people only with their permission, and tell them how their voice clips will be used.

Check yourself

In a made-up town, 25% of the people who will use a helper are children. The team plans 200 voice clips in total. How many should come from children?

Answer: 50 clips

25% is 25 out of every 100. There are two hundreds in 200, so 2 × 25 = 50 clips from children.

A crop-photo app has 200 photos from the plains and only 20 from hill farms. What is the best way to fix this?

A team makes each age group's share match real life, but records everyone in one quiet room in one city. What can still go wrong?

Before recording children's voices for the dataset, what should the team do first?

  • Copy the 20 hill photos until there are 200 of them. The count is equal now, but it is the same 20 photos again and again. The app still learns about only those 20 hill farms.
  • Delete 180 plains photos so both groups have 20. The counts match, but you threw away useful photos and left very little data for everyone. It is better to add than to throw away.
  • Take new photos on many hill farms, in different seasons — correct. Yes. New photos bring new examples of hill farms. Different seasons, places and weather add the variety the app needs.
  • It is fair now, since every age group has its right share. Shares are only one part. People in other places, with other accents or in noisy surroundings could still be missing.
  • It may still fail in noisy villages or for other accents — correct. Yes. Balanced counts do not guarantee fairness. The recordings also need variety, from many places, voices and conditions.
  • A bigger total number of clips will fix any gap that is left. More clips from the same quiet room add count, not variety. A big dataset can still leave out whole places and voices.
  • Ask the children and their parents, and explain the plan — correct. Yes. Consent means people choose with full information. Tell them how the clips will be used. Fair data should never be collected by ignoring the people in it.
  • Record them quietly in public places, since voices are not private. A voice can identify a person, so it is personal data. People should know about the recording and agree to it.
  • Copy clips from websites, because that is much quicker. Taking people's voices without asking is not fair to them, even if it is faster. Permission comes first.
Hold to talk

Subscription Status