Skip to main content
AP exam review verified for 2027

AP Statistics Unit 1 Review: Exploring One-Variable Data and Collecting Data

Review AP Statistics Unit 1 to build the foundation for the entire course: describing and comparing distributions of one variable, calculating summary statistics, and understanding how data are collected through sampling and experimental design. Every inference procedure in later units depends on the vocabulary and reasoning skills introduced here.

Use this page to review all 13 topics, from identifying variables and reading graphs to evaluating sampling methods and experimental designs.

What is AP Statistics unit 1?

AP Statistics Unit 1 asks two big questions: What does the data show, and where did the data come from? Answering the first question requires graphical and numerical tools for describing one-variable distributions. Answering the second requires understanding how sampling and experimental design affect what conclusions are valid.

Unit 1 covers classifying variables, graphing and describing one-variable distributions, calculating and interpreting summary statistics, using z-scores to compare relative positions, and evaluating data collection methods including random sampling, bias, and experimental design.

Exploratory data analysis

Topics 1.1 through 1.9 build the skill of describing data. You classify variables as categorical or quantitative, represent them with appropriate graphs (bar charts, pie charts, histograms, dotplots, stem-and-leaf plots, boxplots), describe distributions using shape, center, variability, and unusual features, and calculate statistics like mean, median, IQR, standard deviation, and z-scores.

Choosing the right statistic

A key judgment in Unit 1 is knowing when to use resistant measures (median and IQR) versus nonresistant ones (mean, range, standard deviation). Outliers pull the mean and standard deviation but barely affect the median and IQR. Skewed distributions call for median and IQR; roughly symmetric distributions can be summarized with mean and standard deviation.

Data collection and conclusions

Topics 1.10 through 1.13 shift to how data are produced. Random sampling allows generalization to a population. Random assignment of treatments in a well-designed experiment allows cause-and-effect conclusions. Observational studies, even with random samples, cannot establish causation because confounding variables remain possible.

The statistical investigation cycle

Every topic in Unit 1 connects to a single cycle: pose an investigative question, collect data using a sound method, describe and analyze the data, and draw conclusions that are appropriate for the study design. A question about a population requires a representative sample. A cause-and-effect claim requires random assignment. Describing data requires matching the graph and statistic to the variable type and distribution shape. Getting this cycle right is what makes statistical reasoning valid.

AP Statistics unit 1 topics

1.1

Introducing Statistics: What Can We Learn from Data?

Identify the components of a statistical study: population, sample, investigative question, and the role of variation and uncertainty in drawing conclusions from data.

open guide
1.2

Variables

Classify variables as categorical or quantitative, and quantitative variables as discrete or continuous. Distinguish parameters (population) from statistics (sample).

open guide
1.3

Tabular Representation and Summary Statistics for One Categorical Variable

Build frequency and relative frequency tables for categorical data and use counts and proportions to make claims in context.

open guide
1.4

Graphical Representations for One Categorical Variable

Construct and interpret bar charts and pie charts; use them to compare two or more groups on the same categorical variable.

open guide
1.5

Graphical Representations for One Quantitative Variable

Create histograms, dotplots, and stem-and-leaf plots to display the distribution of a quantitative variable and see its shape and spread.

open guide
1.6

Descriptions for One Quantitative Variable Distributions

Describe a quantitative distribution by addressing shape (symmetric, skewed, unimodal, bimodal), center, variability, and unusual features such as outliers, gaps, or clusters in context.

open guide
1.7

Summary Statistics for One Quantitative Variable

Calculate mean, median, quartiles, range, IQR, and standard deviation; apply the 1.5×IQR and 2-standard-deviation outlier rules; choose resistant vs. nonresistant measures appropriately.

open guide
1.8

Graphical Representations of Summary Statistics for One Quantitative Variable

Construct boxplots from the five-number summary, identify outliers on a boxplot, and use the relative positions of the mean and median to describe skew.

open guide
1.9

Comparisons of the Distributions for One Quantitative Variable

Compare two or more distributions using parallel boxplots, back-to-back stem-and-leaf plots, or dotplots; calculate and interpret z-scores to compare relative positions within or across distributions.

open guide
1.10

The Investigative Question Revisited and Data Collection

Distinguish experiments, observational studies, surveys, and censuses; determine when random selection justifies generalization and when random assignment justifies causal conclusions.

open guide
1.11

Random Sampling

Identify and compare simple random samples, stratified samples, cluster samples, and systematic samples; explain which method is most appropriate for a given population and question.

open guide
1.12

Potential Problems with Sampling

Recognize voluntary response bias, undercoverage bias, nonresponse bias, and response bias; explain how each type of bias affects the validity of conclusions.

open guide
1.13

Experimental Design

Identify the four elements of a well-designed experiment (comparison, random assignment, replication, control); distinguish completely randomized, randomized block, and matched pairs designs; explain why random assignment supports cause-and-effect conclusions.

open guide
guide

Unit 1 Overview: Exploring One-Variable Data and Collecting Data

Open this guide for a closer review of the topic.

open guide

Unit 1 review notes

1.1

Statistical studies, variables, and key vocabulary

A statistical study collects data from a sample to answer an investigative question about a population. Every study has observational units (the individuals measured), variables (characteristics that vary across units), parameters (numerical summaries of the population), and statistics (numerical summaries of the sample). Variables are either categorical (group labels, like political party) or quantitative (numerical measurements or counts, like height in cm). Quantitative variables are further classified as discrete (countable values, like number of siblings) or continuous (any value in an interval, like weight).

  • Population (N) vs. sample (n): The population includes all individuals of interest; the sample is the subset actually studied. N denotes population size, n denotes sample size.
  • Parameter vs. statistic: A parameter describes the population (often unknown); a statistic describes the sample and is used to estimate the parameter.
  • Categorical variable: Takes category names or group labels as values; summarized with counts and proportions, not means.
  • Discrete vs. continuous quantitative: Discrete variables take countable values (whole numbers); continuous variables can take any value within an interval.
Given a study description, can you identify the population, sample, observational unit, variable type, and whether a number reported is a parameter or a statistic?
ConceptPopulationSample
Size symbolNn
Numerical summaryParameterStatistic
Example (mean)μ (mu)x̄ (x-bar)
Example (std dev)σ (sigma)s
1.3

Representing categorical data with tables and graphs

Categorical data are organized into frequency tables (counts per category) or relative frequency tables (proportions per category). Bar charts display these counts or proportions with bar height representing the frequency or relative frequency of each category. Pie charts show each category as a slice whose area equals its relative frequency, with all slices summing to 1. Both graph types can be used to compare two or more groups on the same categorical variable.

  • Frequency table: Shows the count of observational units in each category.
  • Relative frequency table: Shows the proportion (or percentage) of observational units in each category; all proportions sum to 1.
  • Bar chart: Bars represent categories; height equals frequency or relative frequency. Bars do not need to touch.
  • Pie chart: Each slice area equals the category's relative frequency; the whole pie represents 100% of the data.
Can you convert a frequency table to a relative frequency table and read a bar chart to justify a claim about the data in context?
DisplayShows counts?Shows proportions?Good for comparing groups?
Frequency tableYesNoYes, side by side
Relative frequency tableNoYesYes, standardizes group sizes
Bar chartEitherEitherYes
Pie chartEitherYesLimited
1.5

Graphing and describing quantitative distributions

Quantitative distributions are displayed with histograms (values grouped into bins, bar height shows frequency or relative frequency), dotplots (one dot per value), and stem-and-leaf plots (retains original values, useful for small data sets). To describe any quantitative distribution, address shape, center, variability, and unusual features in context. Shape terms include symmetric, skewed right (long right tail), skewed left (long left tail), unimodal, bimodal, and uniform. Unusual features include outliers, gaps, and clusters.

  • Skewed right: The right tail is longer; most values cluster at the lower end. Income distributions are a classic example.
  • Skewed left: The left tail is longer; most values cluster at the upper end. Exam scores near a ceiling often skew left.
  • Unimodal vs. bimodal: Unimodal distributions have one main peak; bimodal distributions have two prominent peaks.
  • Outlier: A value that falls far from the bulk of the data; always note outliers when describing a distribution.
  • Context requirement: Every description must reference the variable name and units, not just abstract shape terms.
Given a histogram or dotplot, can you write a complete description that addresses shape, center, variability, and any unusual features in context?
1.7

Summary statistics and boxplots

Measures of center are the mean (x̄ = (1/n)Σxᵢ) and median (middle value when ordered). Measures of variability are the range (max minus min), IQR (Q3 minus Q1), and standard deviation (s = √[(1/(n−1))Σ(xᵢ − x̄)²]). The five-number summary (minimum, Q1, median, Q3, maximum) is displayed as a boxplot, where the box spans the middle 50% of the data. Outliers by the 1.5×IQR rule fall below Q1 − 1.5×IQR or above Q3 + 1.5×IQR; on a boxplot, whiskers extend to the most extreme non-outlier values and outliers are plotted individually. Changing units (e.g., inches to centimeters) scales all statistics by the same factor.

  • Resistant vs. nonresistant: Median and IQR are resistant: outliers barely affect them. Mean, range, and standard deviation are nonresistant: outliers can shift them substantially.
  • Mean vs. median and skew: In a right-skewed distribution, the mean is typically greater than the median. In a left-skewed distribution, the mean is typically less than the median.
  • 1.5×IQR rule: A value is a potential outlier if it is more than 1.5×IQR below Q1 or above Q3.
  • Five-number summary: Minimum, Q1, median, Q3, maximum; the basis for a boxplot.
  • Unit changes: Multiplying all values by a constant multiplies the mean, median, IQR, and standard deviation by that constant.
Can you calculate the five-number summary, apply the 1.5×IQR rule to identify outliers, and explain why the median and IQR are preferred over the mean and standard deviation for a skewed distribution?
StatisticMeasuresResistant to outliers?
MeanCenterNo
MedianCenterYes
Standard deviationSpreadNo
IQRSpreadYes
RangeSpreadNo
1.9

Comparing distributions and z-scores

When comparing two or more distributions of the same quantitative variable, address center, variability, shape, and unusual features for each group and make explicit comparative statements (e.g., 'Group A has a higher median than Group B'). Back-to-back stem-and-leaf plots, side-by-side dotplots, and parallel boxplots are common comparison displays. A z-score standardizes a value: z = (xᵢ − μ) / σ, where μ is the population mean and σ is the population standard deviation. A positive z-score means the value is above the mean; a negative z-score means it is below. Z-scores allow comparison of values from distributions with different means and standard deviations.

  • Z-score formula: z = (xᵢ − μ) / σ; measures how many standard deviations a value is from the mean.
  • Comparing z-scores: A student who scores 2 standard deviations above the mean on one test is in a relatively stronger position than one who scores 1 standard deviation above the mean on another, regardless of the raw scores.
  • Comparative language: Use words like 'greater than,' 'less than,' or 'similar to' when comparing distributions; avoid vague statements.
Can you calculate and interpret a z-score, and use it to compare the relative positions of two values from different distributions?
DisplayBest forWhat to compare
Parallel boxplotsComparing medians and spread across groupsMedian, IQR, outliers
Back-to-back stemplotsTwo small data setsShape, center, unusual values
Side-by-side dotplotsSmall-to-moderate data setsShape, clusters, gaps
z-scoresValues from different distributionsRelative position in standard deviations
1.10

Study design: experiments, observational studies, and random sampling

An experiment imposes treatments on experimental units and measures a response variable; random assignment of treatments allows cause-and-effect conclusions. An observational study records variables without imposing treatments; even with a random sample, confounding variables prevent causal conclusions. A census collects data from every member of the population. Random sampling methods include simple random sample (SRS, every sample of size n equally likely), stratified sampling (divide into strata, randomly sample within each), cluster sampling (randomly select entire clusters), and systematic sampling (select every kth unit after a random start). Random selection from a population justifies generalizing results to that population.

  • Experiment vs. observational study: Experiments impose treatments and can support causation; observational studies observe without intervention and cannot.
  • Confounding variable: A variable related to both the explanatory and response variables that provides an alternative explanation for an observed association.
  • Simple random sample (SRS): Every possible sample of size n has an equal chance of selection; the gold standard for unbiased sampling.
  • Stratified sampling: Population divided into homogeneous strata; random samples taken from each stratum. Reduces variability when strata differ from each other.
  • Cluster sampling: Population divided into clusters; entire clusters are randomly selected. Practical when a population is geographically spread out.
Given a study description, can you identify whether it is an experiment or observational study, name the sampling method used, and state what conclusions are justified?
Study typeImposes treatment?Supports causation?Supports generalization?
Experiment with random assignmentYesYesOnly if units randomly selected
Observational study with random sampleNoNoYes, to the sampled population
Observational study without random sampleNoNoOnly to similar individuals
1.12

Bias in sampling and elements of experimental design

Bias is a systematic error that causes a statistic to consistently over- or underestimate a parameter. Voluntary response bias occurs when only self-selected volunteers respond (typically those with strong opinions). Undercoverage bias occurs when part of the population has no chance of being selected. Nonresponse bias occurs when selected individuals do not respond and differ systematically from those who do. Response bias occurs when question wording, interviewer presence, or social desirability causes inaccurate answers. A well-designed experiment requires comparison of at least two treatment groups, random assignment of treatments, replication, and direct control of extraneous variables. Blinding (single or double) prevents knowledge of treatment assignment from influencing responses or assessments. A completely randomized design assigns all treatments at random across all units. A randomized block design groups similar units into blocks first, then randomly assigns treatments within each block. A matched pairs design is a special block design pairing each unit with itself or a very similar unit.

  • Voluntary response bias: Sample consists entirely of volunteers; people with strong opinions are overrepresented.
  • Undercoverage bias: Part of the population is excluded from the sampling frame, so it cannot be selected.
  • Nonresponse bias: Selected individuals who do not respond may differ from those who do, distorting results.
  • Random assignment: Treatments are assigned to experimental units by chance, reducing the influence of confounding variables and supporting causal conclusions.
  • Matched pairs design: Each experimental unit receives both treatments (in random order) or is paired with a very similar unit; reduces variability due to individual differences.
Can you identify the type of bias in a sampling scenario and explain which element of experimental design (comparison, random assignment, replication, or control) is missing from a flawed experiment?
Bias typeCauseDirection of error
Voluntary responseOnly self-selected volunteers respondTypically overstates extreme opinions
UndercoveragePart of population excluded from frameUnderrepresents excluded group
NonresponseSelected individuals do not respondDepends on who declines
ResponseQuestion wording or social pressureToward socially acceptable answers

Key terms

TermDefinition
ParameterA numerical summary of a population characteristic, such as the population mean μ or population standard deviation σ; usually unknown and estimated using a sample statistic.
Categorical VariableA variable that takes category names or group labels as values; summarized with counts and proportions, not means or standard deviations.
HistogramA graph for quantitative data that groups values into bins; bar height shows the frequency or relative frequency of values in each bin.
SkewnessThe asymmetry of a distribution; right-skewed distributions have a longer right tail and the mean is typically greater than the median, while left-skewed distributions have a longer left tail and the mean is typically less than the median.
IQRThe interquartile range, calculated as Q3 minus Q1; measures the spread of the middle 50% of the data and is resistant to outliers.
1.5×IQR ruleA value is a potential outlier if it falls below Q1 minus 1.5×IQR or above Q3 plus 1.5×IQR.
Box PlotA graphical display of the five-number summary (minimum, Q1, median, Q3, maximum); the box spans the middle 50% of the data and whiskers extend to the most extreme non-outlier values.
Sensitivity to extreme valuesThe degree to which a statistic is affected by outliers; the mean and standard deviation are sensitive (nonresistant), while the median and IQR are not (resistant).
Z-ScoreA standardized value calculated as z = (xᵢ minus μ) divided by σ; measures how many standard deviations a data value is above (positive) or below (negative) the mean.
Simple Random SampleA sample of size n selected so that every possible sample of that size has an equal chance of being chosen; the basis for most random sampling methods.
Cluster SamplingA sampling method that divides the population into groups (clusters), randomly selects some clusters, and includes all individuals in the selected clusters.
strataSubgroups of a population that share a common characteristic; used in stratified random sampling to ensure each subgroup is represented in the sample.
Response VariableThe outcome measured on each experimental unit after a treatment is applied; the variable whose behavior the experiment is designed to study.
Experimental UnitThe smallest unit to which a treatment is assigned in an experiment; when experimental units are people, they are called subjects or participants.
Matched Pairs DesignAn experimental design in which each unit receives both treatments in random order, or similar units are paired and each treatment is randomly assigned within each pair; reduces variability due to individual differences.

Common unit 1 mistakes

Claiming causation from an observational study

Even when a random sample is used, an observational study cannot establish cause and effect because confounding variables are not controlled. Only random assignment of treatments in an experiment supports causal conclusions.

Using mean and standard deviation for skewed data

When a distribution is skewed or has outliers, the mean and standard deviation are pulled toward the extreme values and misrepresent the typical value. Use median and IQR instead, and explain why they are more appropriate.

Forgetting context in distribution descriptions

Writing 'the distribution is skewed right with a median of 45' is incomplete. Always name the variable and units: 'the distribution of exam scores (in points) is skewed right with a median of 45 points.'

Confusing stratified and cluster sampling

In stratified sampling, you randomly sample from every stratum. In cluster sampling, you randomly select entire clusters and include all members of those clusters. Mixing these up leads to wrong conclusions about which method reduces bias or variability.

Misapplying the 1.5×IQR rule

The outlier fences are Q1 − 1.5×IQR (lower) and Q3 + 1.5×IQR (upper). A common error is subtracting 1.5×IQR from the median or adding it to the mean instead of using Q1 and Q3 as the anchors.

How this unit shows up on the AP exam

Describing and comparing distributions in context

Free-response questions frequently present a graph or summary statistics and ask you to describe or compare distributions. A complete response addresses shape, center, variability, and unusual features for each group and uses explicit comparative language ('Group A has a larger median than Group B'). Omitting context or making vague comparisons are the most common reasons for losing points on these tasks.

Justifying conclusions based on study design

Multiple-choice and free-response questions often describe a study and ask what conclusions are valid. The key reasoning chain is: random selection from a population justifies generalization to that population; random assignment of treatments justifies cause-and-effect conclusions; neither is possible without the corresponding random mechanism. Identifying confounding variables in observational studies is a related skill tested on the exam.

Selecting and interpreting summary statistics

Exam questions ask you to calculate statistics such as the mean, median, IQR, standard deviation, or z-score, and then interpret them in context or explain which measure is more appropriate for a given distribution. Knowing that the median and IQR are resistant to outliers while the mean and standard deviation are not, and being able to explain why, is a recurring reasoning task across both multiple-choice and free-response sections.

Final unit 1 review checklist

  • Classify every variable correctlyFor any data set or study description, identify each variable as categorical or quantitative (discrete or continuous) and label numerical summaries as parameters or statistics.
  • Describe quantitative distributions completelyPractice writing full distribution descriptions that address shape (with correct skew terminology), center, variability, and any outliers, gaps, or clusters, always in the context of the variable and its units.
  • Choose and justify summary statisticsKnow when to use median and IQR (skewed distributions or outliers present) versus mean and standard deviation (roughly symmetric, no extreme outliers), and be able to explain the reasoning.
  • Apply the 1.5×IQR rule and calculate z-scoresPractice computing Q1, Q3, IQR, and the outlier fences; also calculate z = (xᵢ − μ) / σ and interpret the result as a number of standard deviations above or below the mean.
  • Distinguish study types and their conclusionsBe able to identify whether a study is an experiment or observational study, state whether random selection was used, and explain what type of conclusion (generalization, causation, both, or neither) is justified.
  • Identify sampling methods and bias sourcesRecognize SRS, stratified, cluster, and systematic sampling from descriptions; identify voluntary response, undercoverage, nonresponse, and response bias and explain the direction of each bias.
  • Evaluate experimental designsCheck whether a described experiment includes comparison, random assignment, replication, and control; identify whether it uses a completely randomized, randomized block, or matched pairs design.

How to study unit 1

Step 1: Variables and study vocabulary (Topics 1.1-1.2)Read the topic guides for 1.1 and 1.2. Practice identifying the population, sample, observational unit, variable type, and parameter vs. statistic for three or four different study descriptions. Use the key terms resource to check definitions for parameter, statistic, categorical variable, and quantitative variable.
Step 2: Categorical data displays (Topics 1.3-1.4)Build a frequency table and a relative frequency table from a small data set, then sketch a bar chart. Practice reading bar charts and pie charts to make claims in context. Focus on converting between counts and proportions.
Step 3: Quantitative graphs and distribution descriptions (Topics 1.5-1.6)Sketch a histogram and a dotplot from the same data set and compare what each reveals. Write a full distribution description (shape, center, variability, unusual features in context) for at least two different graphs, one symmetric and one skewed.
Step 4: Summary statistics, boxplots, and z-scores (Topics 1.7-1.9)Calculate the five-number summary, IQR, standard deviation, and outlier fences for a data set by hand. Construct a boxplot and identify any outliers. Then calculate z-scores for two values from different distributions and compare their relative positions. Practice choosing between resistant and nonresistant measures for skewed vs. symmetric data.
Step 5: Data collection, sampling, bias, and experimental design (Topics 1.10-1.13)Review the topic guides for 1.10 through 1.13. For each of five study descriptions, identify the study type, sampling method, any bias present, and what conclusions are justified. Practice explaining why random assignment supports causation and why random selection supports generalization. Use available FRQ practice to work through experimental design questions.

More ways to review

Topic study guides

Open the individual guides for Unit 1 when you want a closer review of one topic.

browse guides

Practice questions

Use AP-style practice after you review the notes so you can check what you understand.

start practice

FRQ practice

Practice free-response reasoning and compare your answer with scoring guidance.

practice FRQs

Cram archive videos

Watch past review streams filtered to Unit 1 when you want a video walkthrough.

open videos

Official unit cheatsheet

Open the Fiveable one-page unit review, then explore visual cheatsheets for a quick refresher.

open unit cheatsheet

Score calculator

Estimate your broader AP score goal after you review the course and exam format.

open calculator

Frequently Asked Questions

What topics are covered in AP Stats Unit 1?

AP Stats Unit 1, Exploring One-Variable Data and Collecting Data, contains 13 topics covering investigative questions; populations, samples, variables, parameters, and statistics; categorical and quantitative displays; shape, center, variability, outliers, and z-scores; random sampling; sampling bias; and experimental design. These are the Fall 2026 CED topics used for the May 2027 exam.

How much of the AP Stats exam is Unit 1?

Unit 1 accounts for 20–30% of the multiple-choice section. That range is a section weight, not a promise that every practice form will use the same percentage, so review the full unit rather than trying to predict one exact question count.

What should I be able to do after AP Stats Unit 1?

By the end of Unit 1, you should be able to construct and compare data displays, calculate and interpret summary statistics in context, choose a sampling method, and distinguish conclusions about a population from cause-and-effect conclusions. The exam rewards the method and the interpretation, so practice writing what a result means in the problem's context instead of stopping at a calculator output.

What are common mistakes in AP Stats Unit 1?

Common Unit 1 mistakes include describing a distribution without context, treating the mean as resistant to outliers, confusing random sampling with random assignment, or claiming causation from an observational study. Slow down long enough to identify the variables, population, and requested conclusion before calculating.

How should I study AP Stats Unit 1?

Start with the vocabulary and conditions, then mix short calculations with full-sentence interpretations. For each missed question, label whether the problem was choosing a method, doing the calculation, or interpreting the result. Finish with timed mixed practice so you have to recognize the method without a topic label.

How does AP Stats Unit 1 connect to the rest of the course?

The study-design ideas return whenever you verify inference conditions, while one-variable displays and summaries support later probability and regression work.

Ready to review Unit 1?Start with the notes, check the topic cards, and use the practice or resource links when they are available for this course.