AP Statistics
7 topics to cover in this unit
AI-generated review video covering all topics
Watch NowFollow-along note packet with fill-in-the-blank
Start Notes20 AP-style questions to test your understanding
Start QuizAlright, statisticians! We're moving from looking at just ONE variable to seeing how TWO quantitative variables might be related. This is where we learn how to make a scatterplot and, crucially, how to describe the pattern, direction, strength, and any unusual features we see. Think of it like comparing two friends' heights and shoe sizes – is there a relationship? Let's find out!
Okay, so we can describe a scatterplot visually, but how do we put a NUMBER on that relationship? Enter the correlation coefficient, 'r'! This little hero tells us the strength and direction of a *linear* relationship between two quantitative variables. But heads up: 'r' is powerful, but it comes with a major caveat!
If we see a linear pattern, we can draw a line through it! Not just any line, though – we're talking about the Least-Squares Regression Line (LSRL). This line is the best fit for our data, allowing us to model the relationship and even make predictions. But remember, with great power comes great responsibility... and some rules!
How good is our 'best-fit' line? That's where residuals come in! A residual is simply the difference between the actual observed value and the value our regression line *predicted*. We use these 'leftovers' to create a residual plot, which is like a diagnostic tool to tell us if our linear model is actually a good fit. No pattern? Good to go! A pattern? Uh oh, linear might not be the best choice!
Let's get down to brass tacks: how do we actually *calculate* that magical Least-Squares Regression Line? We'll dive into the formulas for the slope and y-intercept, and introduce two more critical measures of how well our line fits: the standard deviation of the residuals (s) and the coefficient of determination (r-squared). These numbers give us even more insight into the quality of our linear model!
Sometimes, a single data point can really mess with our regression line. We're talking about outliers, high-leverage points, and influential points! These are the 'troublemakers' of our data set. We need to identify them and understand how they can pull our LSRL around, potentially giving us a misleading model. It's all about checking for unusual data!
What if our scatterplot clearly isn't linear, but still shows a strong curve? Don't despair! Sometimes, we can 'straighten out' curved data by transforming one or both of our variables using logarithms or powers. This lets us apply a linear model to the transformed data, which can then be used to make predictions. It's like putting on special glasses to see the linearity hidden within the curve!