HomeLearnCoursesHackathonsAccount
Unsupervised Learning & Clustering
Dimensionality Reduction and Anomaly Detection · 1/2

Compressing features with PCA

Clustering isn't the only kind of unsupervised learning. Dimensionality reduction tackles a different problem: datasets often have dozens or hundreds of features, far too many to visualize or reason about directly, and many of those features overlap or move together. Principal Component Analysis, PCA, compresses a large number of features down into a much smaller number of new combined features, called components, that still capture most of the meaningful variation in the original data.

The practical payoff is twofold. You can plot high-dimensional data in two or three dimensions after reducing it, letting you visually spot patterns or clusters that would be invisible in the original feature space. And feeding a smaller number of components into another model, supervised or unsupervised, often trains faster and sometimes generalizes better, since PCA tends to strip out redundant, correlated noise along the way.