AIExplainer

What is counterfactual fairness?

A fairness metric in AI that ensures decisions are fair by comparing actual outcomes with hypothetical outcomes where a sensitive attribute is different

Counterfactual fairness is a way to measure fairness in AI systems by asking 'what would have happened if this person's sensitive attribute, like their race or gender, was different?' It helps identify if the AI system is making biased decisions based on these attributes

Imagine a judge making a decision about someone's sentence. Counterfactual fairness is like asking 'would the judge have given the same sentence if the person was a different race or gender?' It helps ensure the decision is fair and not biased

For example, a bank uses an AI system to approve or reject loan applications. Counterfactual fairness can be used to check if the AI system is rejecting applications from people of a certain race or gender more often than others, even if they have the same credit score and income

Counterfactual fairness is used in AI systems to detect and prevent bias in decision-making, such as in hiring, lending, or law enforcement

One misconception is that counterfactual fairness is the same as other fairness metrics, like demographic parity or equalized odds. However, counterfactual fairness is a more nuanced metric that takes into account the specific circumstances of each individual

The concept of counterfactual fairness was first introduced in 2017 by researchers at the University of California, Berkeley, as a way to address the limitations of other fairness metrics

counterfactual fairness metric fairness through counterfactuals

Three products for different needs — explore what’s relevant to you.