Principal component analysis (PCA) is a dimensionality reduction technique that transforms a dataset with many correlated variables into a smaller set of uncorrelated components. To plot PCA in Python, you use scikit-learn for the computation and matplotlib or seaborn for visualization. In R, you use the built-in prcomp() function and ggplot2 for plotting. The core output is identical in both languages: a scatter plot of samples projected onto the first two or three principal components.
What Does a PCA Plot Actually Show?
A PCA plot shows how individual samples relate to each other based on all measured variables at once. Each point represents one sample. The position of that point is determined by its scores on the principal components.
Principal components are new axes. They are calculated so that the first component (PC1) captures the largest possible amount of variance in the data. The second component (PC2) captures the largest remaining variance, and it is mathematically uncorrelated with PC1. This continues for as many components as there are original variables.
When you plot samples on PC1 versus PC2, you are looking at a two-dimensional summary of a potentially high-dimensional dataset. Samples that cluster together tend to have similar profiles across the original variables. Samples that separate along PC1 differ most in whatever the dominant source of variation is.
A critical point that is often misunderstood: PCA does not know about group labels. It is an unsupervised method. If you color points by a group variable after the fact, any separation you see is a result of the data structure, not something PCA was told to find. This is different from supervised methods like linear discriminant analysis.
How Do You Plot PCA in Python Step by Step?
Python’s scikit-learn library provides a straightforward PCA implementation. The workflow has four steps: standardize, fit, transform, and plot.
Step 1: Standardize the data
PCA is sensitive to scale. If one variable is measured in thousands and another in decimals, the larger-scale variable will dominate the components. Standardizing each variable to mean zero and standard deviation one prevents this. In scikit-learn, you use StandardScaler before PCA.
Step 2: Fit PCA
Create a PCA object, specifying the number of components you want. Fitting the model calculates the principal component directions (eigenvectors) and the amount of variance each component explains.
Step 3: Transform the data
The transform step projects your original samples onto the new component axes. The result is a matrix where each row is a sample and each column is a component score.
Step 4: Plot with matplotlib or seaborn
Create a scatter plot with PC1 on the x-axis and PC2 on the y-axis. If you have group labels, pass them to the color argument in seaborn’s scatterplot function or loop through groups in matplotlib. Add axis labels that include the percentage of variance explained by each component.
Here is where a non-obvious detail matters. The variance explained by each component is available in the explained_variance_ratio_ attribute after fitting. Always report these percentages on your axis labels. A PC1 that explains 15% of variance tells a very different story than one explaining 60%.
How Do You Plot PCA in R Step by Step?
R’s base prcomp() function handles PCA natively. The plotting workflow is similarly direct.
Step 1: Standardize and run prcomp
Set the scale argument to TRUE inside prcomp(). This standardizes variables automatically, which is the same step Python requires separately. The function returns an object containing the rotation matrix (loadings), the sample scores, and the standard deviations of each component.
Step 2: Extract scores
The sample coordinates are stored in the x element of the prcomp output. Each column corresponds to a principal component. You can convert this to a data frame and add your grouping variable as an extra column.
Step 3: Calculate variance explained
Square the standard deviations from the prcomp output, divide by the sum of all squared standard deviations, and multiply by 100. This gives the percentage of variance for each component. Alternatively, use the summary() function on the prcomp object, which reports this directly.
Step 4: Plot with ggplot2
Use ggplot() with geom_point(), mapping PC1 to the x aesthetic and PC2 to the y aesthetic. Map your grouping variable to the color aesthetic. Add stat_ellipse() if you want to draw confidence ellipses around groups. Label axes with the variance percentages.
One practical difference between the languages: R’s prcomp uses the singular value decomposition approach by default, while scikit-learn also uses SVD. The numerical results should be nearly identical when both are applied to the same standardized data. Small differences can appear due to sign conventions, but the relative positions of samples are the same.
Python vs R for PCA Plotting: Which Should You Use?
Both languages produce equivalent PCA plots. The choice depends on your existing workflow and what you plan to do next.
| Feature | Python (scikit-learn) | R (prcomp + ggplot2) |
|---|---|---|
| PCA function | sklearn.decomposition.PCA | prcomp() or princomp() |
| Standardization | Separate StandardScaler step | Built-in scale=TRUE argument |
| Plotting library | matplotlib or seaborn | ggplot2 or base plot |
| Variance explained | explained_variance_ratio_ attribute | summary() or manual calculation |
| 3D plotting | matplotlib mplot3d toolkit | plotly or scatterplot3d package |
| Best for | Machine learning pipelines | Statistical analysis and publication figures |
If your downstream work involves machine learning models, Python keeps everything in one ecosystem. If your work is primarily statistical analysis or academic publication, R’s ggplot2 produces publication-quality figures with less customization effort.
What Are the Most Common Mistakes When Plotting PCA?
The most common mistake is skipping standardization. When variables are on different scales, PCA results become misleading. A variable measured in millions will dominate the components even if it is not biologically or clinically meaningful.
Another frequent error is interpreting PCA plots as showing group differences when the separation is driven by a single outlier. Always check whether removing one or two extreme points eliminates the apparent clustering. If it does, the pattern is not reliable.
Plotting only PC1 and PC2 without checking how much variance they capture is also problematic. If PC1 and PC2 together explain only 25% of the total variance, the two-dimensional plot is missing most of the story. In that case, consider whether PCA is the right visualization for your data at all.
A less obvious issue: PCA assumes linear relationships among variables. If your data has strong nonlinear structure, PCA may not capture it well. Methods like t-SNE or UMAP are sometimes used for visualization in those cases, though they have their own limitations and should not be treated as replacements for PCA in all contexts.
How Do You Interpret a PCA Plot Correctly?
Interpretation starts with the variance explained. Check the percentage on each axis. Components that explain very little variance produce plots where apparent patterns may be noise.
Next, look at the loadings. Loadings tell you which original variables contribute most to each component. In Python, these are in the components_ attribute. In R, they are in the rotation element. A high absolute loading means that variable strongly influences the component direction.
If samples separate along PC1, examine which variables have the highest loadings on PC1. That tells you what is driving the separation. Without checking loadings, you know that samples differ but not why.
Clustering in a PCA plot indicates similarity, but it does not prove that groups are statistically different. Formal testing requires additional methods. PCA is an exploratory tool. It generates hypotheses, not conclusions.
Frequently Asked Questions
Do I need to standardize data before running PCA?
Yes, if your variables are on different scales. Without standardization, variables with larger numerical ranges dominate the principal components. Both Python and R provide straightforward ways to standardize before or during PCA.
How many principal components should I plot?
Plotting the first two components is standard for a 2D scatter plot. If PC1 and PC2 together explain less than about 50% of variance, consider a 3D plot or a scree plot to see how many components are meaningful.
Can I use PCA for classification?
No, PCA is an unsupervised method and does not use class labels. It reduces dimensions and reveals structure. For classification, you would use supervised methods like logistic regression or random forests, potentially on PCA-reduced data.
Are Python and R PCA results identical?
The numerical results are nearly identical when both use SVD on standardized data. Small differences in sign or floating-point precision can occur, but the relative positions of samples in the plot remain the same.

