Skip to main content
The new Teacher Workspace is here. Your first 3 assignments are free. Try it →
A colleague sent you $30 off Fiveable. Get the annual plan for $99 →

Test-retest reliability in AP Psychology

In AP Psychology, test-retest reliability is the consistency of a test's scores when the same people take the same test at two different times. If scores stay similar, the test is reliable, which is a psychometric principle every intelligence and achievement test must meet (Topic 2.8).

Verified for the 2027 AP Psychology exam•Last updated October 2026

What is test-retest reliability?

Test-retest reliability asks one question. If you give the same test to the same people on two different occasions, do they get roughly the same scores? If yes, the test has high test-retest reliability. If scores jump around for no good reason, the test isn't measuring anything stable.

Think of it like a bathroom scale. Step on it Monday, step on it Tuesday. If it says 150 one day and 172 the next, you don't trust it, even if one of those numbers happens to be right. In Topic 2.8, the CED states that all psychological assessments, including intelligence tests, should follow sound psychometric principles to be considered useful. Reliability is one of those principles, alongside standardization and validity. Psychologists usually check test-retest reliability by correlating the first set of scores with the second. A strong positive correlation means the test is consistent.

Why test-retest reliability matters in AP® Psychology

Test-retest reliability lives in Unit 2: Cognition, inside Topic 2.8 Intelligence and Achievement. It directly supports AP Psych 2.8.B (explain how intelligence is measured). The essential knowledge there says psychological assessments must adhere to sound psychometric principles to be useful. That's the backbone of any question asking whether a test is any good.

It also feeds into AP Psych 2.8.C, which covers systemic issues in how intelligence scores are used. IQ scores are used to identify students for educational services, so an inconsistent test means real people get sorted on bad data. And because achievement and aptitude tests (AP Psych 2.8.D) are judged by the same standards, the concept applies well beyond IQ tests.

Keep studying AP® Psychology Unit 2

How test-retest reliability connects across the course

Validity (Unit 2)

Reliability asks whether a test is consistent. Validity asks whether it measures what it claims to. A test can be reliable without being valid, like a scale that's always 10 pounds off. But a test can't be valid if it isn't reliable, because random noise can't track a real trait.

Standardization (Unit 2)

The CED defines a standardized test as one given with consistent procedures and environments. That consistency is what makes a retest a fair comparison. If the first sitting is quiet and timed and the second is noisy and untimed, score differences reflect the setting, not the person.

Flynn Effect (Unit 2)

IQ scores have risen across generations due to factors like better nutrition, health care, and socioeconomic status. That's a population shift over decades, not evidence that a test is unreliable. Test-retest reliability is about the same individuals over a short gap, so don't mix the two up.

Stereotype Threat and Bias (Unit 2)

Situational pressures like stereotype threat, or testing someone in a non-native language, can push one person's score up or down between sittings. That's a reminder that a single score is a snapshot, and the conditions around a test shape how consistent and fair it is.

Is test-retest reliability on the AP® Psychology exam?

Expect test-retest reliability in multiple-choice scenario questions that describe what a researcher did and ask you to name the psychometric concept. The giveaway is a same test, same people, two time points setup. A stem might describe students taking a new intelligence test in September and again in October, then ask which property is being evaluated.

The trap is the distractors. Questions in this area often pair reliability with validity and standardization. If a researcher compares test scores with academic performance, that's checking validity, not reliability. If proctors get identical training and follow a script, that's standardization. Only repeated testing of the same group points to test-retest reliability.

No released FRQ has used this term verbatim. It still works well in free-response answers when a prompt asks you to evaluate a measure or a research design. Explaining that a test must give consistent results over time before you can trust its conclusions is the kind of applied reasoning those questions reward.

Test-retest reliability vs Validity

Reliability means consistency. Do you get the same result again? Validity means accuracy. Does the test actually measure the trait it claims to? A test given twice with matching scores is reliable. A test whose scores line up with a real-world outcome, like academic performance, shows validity. Remember that reliability is necessary for validity but doesn't guarantee it. A ruler printed with the wrong spacing gives you the same wrong length every time.

Key things to remember about test-retest reliability

  • Test-retest reliability means the same people get similar scores when they take the same test at two different times.

  • Psychologists usually measure it by correlating the first set of scores with the second, and a strong positive correlation signals a reliable test.

  • A test can be reliable without being valid, but it cannot be valid if it isn't reliable.

  • Standardized procedures and environments help make retest comparisons fair, because changing the conditions can change the scores.

  • On MCQs, look for a scenario with the same test and the same group at two time points, and don't confuse it with validity or standardization.

Frequently asked questions about test-retest reliability

What is test-retest reliability in AP Psychology?

It's the consistency of a test's results when the same people take the same test on two separate occasions. In Topic 2.8, it's one of the psychometric principles an intelligence or achievement test needs to meet to be considered useful.

If a test is reliable, does that mean it's valid?

No. Reliability only shows the test gives consistent results, and a test can be consistently wrong. Validity requires showing the test actually measures what it claims, for example by comparing scores with academic performance.

How is test-retest reliability different from validity?

Test-retest reliability asks whether scores stay the same across two sittings. Validity asks whether those scores reflect the real trait being measured. If a researcher retests the same group, that's reliability. If they compare scores with an outside outcome like grades, that's validity.

Does the Flynn Effect mean IQ tests are unreliable?

No. The Flynn Effect describes IQ scores rising across generations because of societal factors like better nutrition and health care. Test-retest reliability looks at the same individuals over a short time gap, so a generational shift doesn't make a test inconsistent for a given person.

Is test-retest reliability on the AP Psych exam?

Yes, it fits under AP Psych 2.8.B, which says assessments must follow sound psychometric principles. It most often shows up in multiple-choice scenarios where you have to tell reliability apart from validity and standardization.