Skip to main content
The new Teacher Workspace is here. Your first 3 assignments are free. Try it →

Posterior predictive checks

Posterior predictive checks are a Bayesian way to test model fit by simulating new data from the posterior predictive distribution and comparing it to the observed data. In Intro to Probability, you use them to see whether your probability model matches reality.

Last updated July 2026

What are posterior predictive checks?

Posterior predictive checks are a model-checking tool in Intro to Probability and Bayesian inference. You take the probability model you built, use the posterior distribution for its unknown parameter values, and simulate new data from that model. Then you compare those simulated results to the data you actually observed.

The basic idea is simple: if your model is a good description of the random process, the simulated data should look a lot like the real data. If the observed data has a pattern your simulations never produce, that is a sign the model is missing something. You are not asking, "What parameter value is most likely?" You are asking, "Does this model generate data that behaves like my sample?"

This matters because a Bayesian model is not just a formula for one answer. It is a whole data-generating story. Posterior predictive checks test that story by pushing the model forward into new fake samples and seeing whether those samples match features of the observed sample, such as the mean, variance, shape, or tails.

A useful example is a binomial model for coin flips. Suppose your model says the probability of heads is about 0.6, based on your posterior distribution. You can simulate many new coin-flip samples of the same size, then compare the number of heads, the spread, or the runs of heads and tails against the real data. If the real sample has way more streaks than your simulations, your model may be too simple or based on the wrong assumptions.

The comparison can be visual or numerical. You might use histograms, Q-Q plots, or summary statistics. A common move is to calculate a check statistic for the observed data and see where it falls among the same statistic computed from simulated datasets. If the observed value sits far in the tail of the simulated values, that mismatch is a warning sign.

The main thing to remember is that posterior predictive checks are about fit, not just estimation. A posterior can give you a nice-looking parameter estimate and still describe the data poorly. These checks help you catch that before you trust the model too much.

Why posterior predictive checks matter in Intro to Probability

Posterior predictive checks matter in Intro to Probability because they connect probability theory to actual data behavior, not just formulas. A model can have clean probabilities and still miss the real pattern in a sample, and this tool shows you that gap.

They also force you to think like a model builder. Instead of only solving for a probability, expected value, or posterior distribution, you ask whether the process you wrote down could have produced the data in front of you. That shift is a big part of Bayesian reasoning, especially when you are comparing one model to another or deciding whether an assumption like independence or constant probability feels realistic.

These checks are especially useful when the data has shape features that a single summary number misses. Two datasets can have the same mean but very different spread, skew, or tail behavior. Posterior predictive checks let you compare those features directly, which is much more informative than checking one number in isolation.

They also give you a natural place to revise a model. If the observed data keeps failing the check, you may need a better likelihood, a different prior, or a different probability structure altogether. In other words, the check is not the end of the process, it is the feedback loop that tells you whether your probabilistic story is believable.

Keep studying Intro to Probability Unit 12

Official unit cheatsheet

open one-pager

How posterior predictive checks connect across the course

Bayes' theorem

Bayes' theorem is what gets you from a prior and new data to a posterior distribution. Posterior predictive checks come after that update, because you use the posterior to generate simulated data. If Bayes' theorem gives you the updated parameter belief, posterior predictive checks ask whether that updated belief actually produces realistic outcomes.

Posterior distribution

The posterior distribution is the source of the parameter values you sample from during a posterior predictive check. Instead of plugging in one fixed estimate, you carry the uncertainty in the parameter forward into simulated datasets. That is why the check reflects Bayesian uncertainty instead of a single best guess.

Predictive distribution

The predictive distribution describes future or unobserved data under your model. Posterior predictive checks use that idea directly by generating fake samples from the posterior predictive distribution and comparing them to observed data. If the predictive distribution is off, your model may be missing structure in the data.

Conditional Independence

Conditional independence often sits behind the way a probability model is built. If that assumption is wrong, the simulated data from your model may look too smooth, too regular, or too simple compared with the real sample. Posterior predictive checks can reveal those kinds of mismatches.

Are posterior predictive checks on the Intro to Probability exam?

A quiz or problem set may give you a Bayesian model and ask whether simulated data matches an observed sample. Your job is to read the discrepancy, not just compute a number. Look for the feature being checked, such as mean, spread, tails, or unusual clustering, then explain whether the observed data would look typical under the posterior predictive distribution.

You may also be asked to interpret a plot or a table of simulated summaries. If the observed statistic falls far outside the cloud of simulated values, that is evidence of poor fit. The common mistake is treating a posterior predictive check like a test of whether the parameter estimate is "correct." It is really a check on the model's ability to generate data that looks like the sample you saw.

Posterior predictive checks vs Posterior distribution

The posterior distribution is about what you believe for the unknown parameter after seeing data. Posterior predictive checks are about what data the model would generate next, using that posterior. One is a belief about parameters, the other is a fit check for the data-generating process.

Key things to remember about posterior predictive checks

  • Posterior predictive checks compare real data to data simulated from a Bayesian model after the model has been updated with the sample.

  • The goal is to see whether the model can reproduce the shape and features of the observed data, not just estimate a parameter.

  • You can check summaries like the mean, variance, tails, or visual patterns such as histograms and Q-Q plots.

  • If the observed data looks unusual compared with the simulations, the model may be missing a feature like extra spread, skew, or dependence.

  • These checks are part of the feedback loop in Bayesian modeling, so you can revise the model and test again.

Frequently asked questions about posterior predictive checks

What is posterior predictive checks in Intro to Probability?

Posterior predictive checks are a way to compare observed data with new data simulated from a Bayesian model. You use the posterior distribution to generate fake samples, then see whether those samples resemble the real one. In Intro to Probability, this is a direct way to judge model fit.

How are posterior predictive checks different from the posterior distribution?

The posterior distribution tells you how plausible different parameter values are after seeing data. Posterior predictive checks go a step further and ask whether those parameter values produce realistic data. So the posterior is about parameters, while the check is about the model's data output.

What do you compare in a posterior predictive check?

You compare observed data to simulated data using summary statistics or plots. Common checks include the mean, variance, quantiles, tail behavior, histograms, and Q-Q plots. The exact comparison depends on what feature of the data seems most important.

What does it mean if the observed data fails a posterior predictive check?

It usually means the model does not capture some part of the data well. The issue might be the chosen distribution, an independence assumption, or a missing feature like skew or clustering. Failing the check is a sign to revise the model, not just the parameter estimate.

Posterior Predictive Checks | Intro to Probability | Fiveable