Before you run a single statistical test on two variables, you need to know one thing: do they even move together? A scatter plot answers that question at a glance, long before you touch a formula. It is the starting point of almost every study in economics, psychology, agriculture, and business that looks at how two numerical variables relate to each other.
Table of Contents
- What is a scatter plot in bivariate data analysis?
- Reading the pattern: positive, negative, or no relationship
- Positive relationship
- Negative relationship
- No relationship
- Linear or curvilinear? Why the shape matters
- From visual pattern to a number: the correlation coefficient
- What the values actually mean
- Seeing it in action: rainfall and rice yield
- The catch: correlation isn’t causation
- Putting it all together
What is a scatter plot in bivariate data analysis?
When you collect data on two numerical variables for the same set of subjects, you get bivariate data – pairs of values like (hours studied, marks scored) or (advertising spend, sales revenue). A scatter plot takes each of these pairs and plots them as a single dot on a graph.
The independent variable, usually labelled x, goes on the horizontal axis. The dependent variable, y, goes on the vertical axis. Each dot represents one observation. Once you have plotted dozens or hundreds of these points, a pattern often starts to appear, and that pattern is exactly what analysts are hunting for. Scatterplots are used to examine relationships between variables, spot outliers, and check whether a regression model actually fits the data before you commit to deeper analysis.
Reading the pattern: positive, negative, or no relationship
Once the dots are on the graph, the shape they form tells you almost everything you need to know about the relationship between the two variables.
Positive relationship
If the dots generally climb from the bottom-left to the top-right of the graph, you have a positive relationship. As one variable increases, the other tends to increase too. Hours spent studying versus exam marks is a classic example – more study time is usually paired with higher scores.
Negative relationship
If the dots slope downward from top-left to bottom-right, that is a negative relationship. As one variable rises, the other tends to fall. Think of the price of a vegetable and the quantity people buy – as price goes up, demand typically drops.
No relationship
Sometimes the dots are scattered all over the graph with no visible slope in either direction. This indicates no relationship between the variables – knowing the value of x tells you nothing useful about y. Shoe size and exam marks would likely fall into this category.
A positive linear relationship exists when the response variable tends to increase as the explanatory variable increases, while a negative linear relationship shows the opposite pattern, and the strength of either relationship depends on how tightly the points cluster around an imaginary line running through them.
Linear or curvilinear? Why the shape matters
Not every relationship is a straight line. Some variables move together in a curve – productivity might rise with hours worked up to a point, then fall as fatigue sets in. A scatter plot lets you spot this curvilinear shape immediately, which a single summary number cannot always reveal on its own.
This visual check matters because it decides which statistical tool you reach for next. If the pattern looks like it hugs a straight line, you can move on to calculating a linear correlation coefficient. If it curves, a straight-line correlation measure would actually mislead you, since a trend line fitted to a scatter plot only makes sense when the underlying relationship is reasonably linear in the first place.
From visual pattern to a number: the correlation coefficient
A scatter plot gives you a qualitative sense of the relationship, but researchers usually want a precise, comparable number. That number is the Pearson product moment correlation coefficient, denoted by r.
Pearson’s r is the most widely used measure of a linear correlation between two variables, and it condenses everything a scatter plot shows visually into a single value between โ1 and +1.
What the values actually mean
A value of +1 or โ1 indicates a perfect linear relationship, while a value of 0 means there is no linear relationship at all. Values in between tell you how closely the points hug that imaginary straight line:
- Close to +1: Strong positive relationship – the dots form a tight upward line.
- Close to โ1: Strong negative relationship – the dots form a tight downward line.
- Close to 0: Weak or no linear relationship – the dots are widely scattered.
The greater the absolute value of r, the stronger the relationship, with values near the extremes of โ1 and +1 indicating that a change in one variable is accompanied by a highly consistent, predictable change in the other. This is exactly why glancing at the scatter plot first is so useful – a tight cluster of points along a line is a visual preview of a high absolute r value, while a loose cloud of points warns you that r will be close to zero.
Seeing it in action: rainfall and rice yield
Bivariate analysis is not just a classroom exercise. Agricultural researchers routinely use scatter plots and correlation coefficients to understand how weather affects crop output, which has enormous practical stakes for a country where farming still supports a huge share of the population.
In a study of Bankura district in West Bengal, researchers plotted annual rainfall against rice production and found a moderate positive correlation of r = 0.40, indicating that rainfall variation has a noticeable but not overwhelming influence on rice yield in the region. This is a good reminder that real-world correlations are rarely perfect – a moderate r still tells a meaningful story, even if other factors like soil quality, irrigation, and fertiliser use are also shaping the outcome.
The catch: correlation isn’t causation
A scatter plot showing a clear upward or downward trend can be tempting to over-interpret. Seeing two variables move together does not prove that one causes the other. An observed relationship might actually be driven by a third variable affecting both plotted variables, or the pattern could simply be coincidental – larger cities, for instance, tend to have both more parks and more crime simply because they have more people, not because parks cause crime.
This is why careful researchers treat a scatter plot and its correlation coefficient as a starting point for investigation, not a final verdict. If a genuine cause-and-effect link needs to be established, additional analysis is required to rule out other explanations and control for confounding factors.
Putting it all together
A scatter plot is one of the simplest tools in a data analyst’s kit, yet it does an enormous amount of work. It shows you the direction of a relationship, hints at its strength, flags whether the pattern is linear or curved, and even helps you spot outliers before you calculate anything. Once that visual check is done, the Pearson correlation coefficient gives you a precise number to describe what your eyes already suspected.
Think about it this way: every time you see a claim like “students who sleep more score higher marks” or “cities with more rainfall grow more rice,” somewhere behind that claim is very likely a scatter plot and a correlation coefficient doing the heavy lifting.
What do you think? Pick two variables from your own field of study – what pattern would you expect if you plotted them against each other, positive, negative, or none at all? And can you think of a case where two variables might look correlated on a scatter plot purely because of a hidden third factor?
References
- https://statisticsbyjim.com/graphs/scatterplots/
- https://openstax.org/books/contemporary-mathematics/pages/8-8-scatter-plots-correlation-and-regression-lines
- https://www.atlassian.com/data/charts/what-is-a-scatter-plot
- https://www.scribbr.com/statistics/pearson-correlation-coefficient/
- https://www.britannica.com/topic/Pearsons-correlation-coefficient
- https://statisticsbyjim.com/basics/correlations/
- https://link.springer.com/article/10.1007/s43621-025-01392-6
Leave a Reply