AIExplainer
Machine Learning Intermediate 2 min read

What is class-imbalanced dataset?

A dataset where the number of samples in one class significantly exceeds the number of samples in another class

In machine learning, a class-imbalanced dataset occurs when one category or class has a much larger number of examples than others, which can affect the performance of models trained on that data

Imagine trying to learn the difference between cars and bicycles by looking at 1000 pictures of cars and only 10 pictures of bicycles - it's hard to get a good sense of what a bicycle looks like

A credit card company's dataset of transactions, where the vast majority are legitimate and only a small fraction are fraudulent, is an example of a class-imbalanced dataset

Class-imbalanced datasets are often encountered in real-world problems, such as detecting rare diseases or predicting fraudulent transactions, and require special techniques to handle the imbalance

One common misconception is that class-imbalanced datasets can be handled by simply oversampling the minority class or undersampling the majority class, but this can lead to overfitting or loss of important information

The problem of class-imbalanced datasets has been recognized since the early days of machine learning, and various techniques have been developed to address it, such as SMOTE, ADASYN, and cost-sensitive learning

imbalanced dataset skewed dataset uneven dataset

Three products for different needs — explore what’s relevant to you.