Skip to main content
The new Teacher Workspace is here. Your first 3 assignments are free. Try it →

Maximum likelihood estimation

Maximum likelihood estimation, or MLE, is a way to estimate a distribution’s unknown parameters by choosing the values that make the observed data most likely. In Intro to Probability, it shows up when you fit models like binomial, Poisson, or normal distributions to data.

Last updated July 2026

What is maximum likelihood estimation?

Maximum likelihood estimation is the method of choosing parameter values for a probability model that make your observed data as likely as possible. In Intro to Probability, that usually means you already picked a distribution, such as binomial, Poisson, or normal, and now you want the best parameter estimate from the sample you actually saw.

The core idea is the likelihood function. You treat the data as fixed and the parameter as the unknown quantity, then ask: for which parameter value would this data have been most plausible under the model? That is different from asking which parameter makes the probability of the data largest in a casual sense. In MLE, the data are the input and the parameter is what you vary.

A simple example is a Poisson model for counts, like the number of calls arriving per hour. If you observe call counts over many hours, the likelihood is built from the Poisson probability mass function using the candidate rate λ. The MLE is the λ that makes the full observed sample most likely, and for the Poisson model that estimate turns out to be the sample mean.

That is a nice pattern to notice in probability courses: the MLE often matches the natural summary statistic for the distribution. For a binomial proportion, the MLE for p is the sample proportion. For a normal distribution with known variance, the MLE for the mean is the sample mean. The math can be a little different each time, but the setup stays the same, write the likelihood, simplify it if you can, then maximize it.

You will usually maximize the log-likelihood instead of the likelihood itself. Taking logs does not change which parameter gives the maximum, but it turns products into sums and makes the algebra cleaner. In class problems, this is often where the calculation becomes manageable, especially when there are several observations in the sample.

MLE also sits inside statistical inference, because it turns sample data into an estimate about the population or process behind it. The estimate is not guaranteed to be perfect, and small samples can give biased results, but with more data MLE often becomes more accurate and stable.

Why maximum likelihood estimation matters in Intro to Probability

Maximum likelihood estimation matters in Intro to Probability because it is one of the main bridges from probability models to inference. You are not just computing probabilities anymore, you are using a model to estimate an unknown rate, chance, or mean from data.

That shows up any time a problem gives you observations and asks you to identify the parameter that best fits them. If the data look like counts over time, MLE connects naturally to the Poisson distribution. If the data are successes and failures, it connects to binomial probability. Once you know the model, MLE gives you a systematic way to estimate the parameter instead of guessing from the sample.

It also helps you interpret what a good estimate means. MLE does not say the estimated parameter is certain, only that it is the value that made the observed sample most plausible under the assumed model. That distinction matters in probability, because the whole method depends on the model being a reasonable fit in the first place.

This idea also prepares you for later inference tools like confidence intervals and hypothesis tests. Those methods often build on an estimate first, then ask how much uncertainty surrounds it. If you can recognize the likelihood setup, the rest of the inference workflow makes a lot more sense.

Keep studying Intro to Probability Unit 8

Official unit cheatsheet

open one-pager

How maximum likelihood estimation connects across the course

Likelihood Function

MLE is built from the likelihood function. The likelihood takes the data you observed and treats the parameter as the variable, so you can compare which parameter value makes the sample most plausible. If you do not set up the likelihood correctly, the MLE step cannot work.

Statistical Inference

Maximum likelihood estimation is one of the main tools inside statistical inference. Inference is the bigger process of using sample data to say something about a population or random process, and MLE is one way to turn that sample into a parameter estimate before you draw conclusions.

Method of Moments

Method of moments is another way to estimate parameters, but it uses sample moments like the mean or variance instead of maximizing a likelihood. The two methods can give the same answer for some distributions, but MLE is usually tied more directly to the assumed probability model.

Consistency

A consistent estimator gets closer to the true parameter as sample size grows. Many MLEs have this property, which is one reason they are favored in probability and statistics. The idea is that more data should make your estimate settle toward the real value, not wander away from it.

Is maximum likelihood estimation on the Intro to Probability exam?

A problem set question usually asks you to build the likelihood from a named distribution, simplify it, and identify the parameter value that maximizes it. For a Poisson data set, that often means writing the product of Poisson probabilities and then finding the rate parameter that fits the observed counts best. If the course uses a normal or binomial model, you may be asked to recognize the MLE without a full derivation, especially when the answer is the sample mean or sample proportion.

You may also see MLE in interpretation questions. The task is then to explain what the estimate means in context, not just compute it. A good response says something like, this parameter value makes the observed sample most likely under the chosen model, while also noting that the result depends on the model being appropriate.

Maximum likelihood estimation vs Method of Moments

Method of moments and maximum likelihood estimation are both parameter-finding methods, but they get there in different ways. Method of moments matches sample moments to theoretical moments, while MLE chooses the parameter that maximizes the likelihood of the observed data. On many Intro to Probability problems, MLE is the more direct fit-to-data approach.

Key things to remember about maximum likelihood estimation

  • Maximum likelihood estimation finds the parameter value that makes the observed sample most likely under a chosen probability model.

  • In Intro to Probability, MLE commonly appears with binomial, Poisson, and normal distributions.

  • The likelihood function treats the data as fixed and the parameter as the unknown quantity you are trying to optimize.

  • You often work with the log-likelihood because it is easier to simplify and maximize.

  • MLE is a major inference tool, but it depends on choosing a model that actually fits the situation.

Frequently asked questions about maximum likelihood estimation

What is maximum likelihood estimation in Intro to Probability?

Maximum likelihood estimation is a method for choosing distribution parameters that make your observed data most likely. In Intro to Probability, you use it after deciding on a model like binomial, Poisson, or normal, then solve for the parameter value that best fits the sample.

How do you find the MLE of a parameter?

First write the likelihood function from the probability model and the observed data. Then simplify it, often by taking the log, and maximize it with algebra or calculus. For many common distributions, the MLE is a familiar summary like the sample mean or sample proportion.

Is maximum likelihood estimation the same as method of moments?

No. Method of moments matches sample averages or other moments to theoretical moments, while MLE chooses the parameter that makes the sample most probable. They can produce the same estimate in some cases, but they are based on different ideas.

Why does MLE matter for Poisson problems?

Poisson problems often model counts over time or space, like arrivals per hour or defects per page. MLE gives you a clean way to estimate the rate parameter λ from observed counts, and in the Poisson case that estimate is the sample mean.

Maximum Likelihood Estimation | Intro to Probability | Fiveable