What is Data Augmentation?
Creating additional training examples by slightly modifying existing ones — flipping, cropping, or rephrasing — to increase variety without new data collection.
Data Augmentation explained in plain English
Data augmentation creates additional training examples by slightly modifying existing ones — flipping, cropping, or rephrasing — so the system sees more variety without collecting entirely new data.
It helps models generalise when real-world data is limited or expensive.
Analogy
Data augmentation is like a musician practising a piece in different keys and tempos. The song is the same, but the varied conditions build flexibility and resilience.
Example
A small photo dataset of defects on a factory line can be augmented with rotations and brightness changes to train a more robust inspector model.
How is Data Augmentation used?
Self-driving car systems train on rotated and shifted images of roads. Language models benefit from paraphrased sentences. Medical AI uses augmented scans when real patient data is limited.
Common misconceptions about Data Augmentation
Augmentation must stay realistic — extreme transformations can introduce noise that hurts rather than helps.
People also read
- A/B testing
A method of comparing two versions of a product or service to determine which one performs better
- ablation
A technique used to remove or disable parts of a machine learning model to understand their importance
- accuracy
The degree to which a model's predictions match the actual outcomes
- activation function
A mathematical function that introduces non-linearity into a neural network model
- active learning
A machine learning approach where the model actively selects the most informative data to learn from
- adaptation
The process of adjusting to new or changing conditions
- agglomerative clustering
A type of hierarchical clustering that groups similar data points together
- anomaly detection
The process of identifying data points that do not conform to expected patterns or behaviors
- area under the PR curve
A measure of a model's performance in classification tasks
- area under the ROC curve
A measure of a model's ability to distinguish between positive and negative classes